LLM Red Team & Benchmark Evaluation Specialist

đź•’ vor 28 Tagen

🗽 New York – Remote

infoinfo

đź’µ $55 - $85 / Stunde

⏰ Vollzeit

🟢 Junior

đź‘» Geisterscore 6%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of 24-MAG

24-MAG

2 - 10 Mitarbeiter

🤝 B2B

đź’Ľ Beratung

B2B • Consulting

24-MAG ist eine Unternehmensberatung fĂĽr kommerzielle Strategie und Umsetzung, die B2B-Unternehmen dabei unterstĂĽtzt, Systeme, Workflows und Betriebsrhythmen fĂĽr Vertrieb, Kundenmanagement und bereichsĂĽbergreifende Projekte zu entwerfen und umzusetzen. Sie konzentrieren sich darauf, verstreute Prozesse in abgestimmte, messbare und skalierbare kommerzielle Funktionen zu verwandeln. Dies umfasst die Struktur des Vertriebskanals, Rahmenwerke fĂĽr das Kundenmanagement und operative Disziplinen fĂĽr Teams, die ein effizientes und zielgerichtetes Wachstum anstreben.

Beschreibung

• Adversarial Model Evaluation • Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks • Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs • Investigate situations where models appear capable while reaching incorrect or unsupported conclusions • Design reproducible experiments to isolate and validate model failure modes • Benchmark & Challenge Design • Convert observed model weaknesses into rigorous benchmark tasks • Develop challenges that are technically demanding while remaining fair and objectively assessable • Define clear task requirements, expected outcomes, and evaluation criteria • Ensure tasks require genuine reasoning rather than allowing shortcuts or superficial pattern matching • Failure Analysis & Documentation • Document findings with clear evidence, methodology, and reproducible steps • Explain why a model failed and which capabilities or assumptions contributed to the error • Produce detailed technical write-ups for researchers and task authors • Track recurring failure patterns across models, prompts, and evaluation environments • Task Strengthening & Research Collaboration • Work with task authors to close loopholes, grading gaps, and unintended shortcuts • Review benchmark tasks for ambiguity, exploitability, and evaluation reliability • Share insights with researchers and other specialists to improve benchmark coverage • Participate in iterative calibration, peer review, and task-refinement workflows

🎯 Anforderungen

• At least 1 year of experience in research, research engineering, security, AI evaluation, or a related technical role • Demonstrated experience identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems • Background in red teaming, adversarial testing, security research, benchmark development, or rigorous model evaluation • Working proficiency in Python and Git • Ability to develop scripts, probes, and analyses independently • Strong familiarity with LLM capabilities, limitations, and evaluation techniques • Excellent written communication and technical documentation skills • Creativity, precision, and persistence when working through ambiguous research problems • Reliable availability for approximately 35 hours per week • A master's degree or PhD in a STEM field is highly relevant

Jetzt Bewerben

Ähnliche Jobs

đź•’ vor 28 Tagen

Books Are Fun

51 - 200

📚 Bildung

🤲 Wohltätigkeit

Outreach Specialist connecting K-6 schools to the Book Blast program through presentations and meetings. Focused on scheduling with decision-makers to promote student home libraries.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $25 / Stunde

⏰ Vollzeit

🟢 Junior

🚫👨‍🎓 Kein Abschluss erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

GE Vernova

10.000+ Mitarbeiter

đź’Ľ Beratung

📦 Logistik

🏭 Fertigung

DERMS Solution Specialist deploying GridOS software for utility customers. Designing distributed-energy planning solutions, testing workflows, training users, and supporting delivery and sales.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $66.000 - $99.000 / Jahr

⏰ Vollzeit

🟢 Junior

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

GE Vernova

10.000+ Mitarbeiter

đź’Ľ Beratung

📦 Logistik

🏭 Fertigung

DERMS solution specialist deploying GridOS for utility distributed-energy planning. Advising utilities, designing customer solutions, and training users for cleaner grid operations.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $66.000 - $99.000 / Jahr

⏰ Vollzeit

🟢 Junior

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

Nouria

1001 - 5000

🍽️ Lebensmittel & Getränke

📦 Logistik

đź›’ Einzelhandel

Visual Merchandising Specialist focused on store layout and product arrangement at Nouria Energy. Collaborating with merchandising and marketing teams to maximize sales and customer experience.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟢 Junior

🟡 Mittelstufe

🚫👨‍🎓 Kein Abschluss erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

Stride, Inc.

5001 - 10000

đź’Ľ Beratung

🏥 Gesundheitswesen

📚 Bildung

Placement Coordinator responsible for transcript evaluations and assisting students and families with placement and graduation planning. Work remotely while collaborating with a team across various locations.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $16 / Stunde

⏰ Vollzeit

🟢 Junior

🗣️🇺🇸🇬🇧 Englisch erforderlich