Search Remote Jobs

LLM Red Team & Benchmark Evaluation Specialist

đź•’ July 27

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

đź’Ľ Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

đź“‹ Description

• Adversarial Model Evaluation • Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks • Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs • Investigate situations where models appear capable while reaching incorrect or unsupported conclusions • Design reproducible experiments to isolate and validate model failure modes • Benchmark & Challenge Design • Convert observed model weaknesses into rigorous benchmark tasks • Develop challenges that are technically demanding while remaining fair and objectively assessable • Define clear task requirements, expected outcomes, and evaluation criteria • Ensure tasks require genuine reasoning rather than allowing shortcuts or superficial pattern matching • Failure Analysis & Documentation • Document findings with clear evidence, methodology, and reproducible steps • Explain why a model failed and which capabilities or assumptions contributed to the error • Produce detailed technical write-ups for researchers and task authors • Track recurring failure patterns across models, prompts, and evaluation environments • Task Strengthening & Research Collaboration • Work with task authors to close loopholes, grading gaps, and unintended shortcuts • Review benchmark tasks for ambiguity, exploitability, and evaluation reliability • Share insights with researchers and other specialists to improve benchmark coverage • Participate in iterative calibration, peer review, and task-refinement workflows

🎯 Requirements

• At least 1 year of experience in research, research engineering, security, AI evaluation, or a related technical role • Demonstrated experience identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems • Background in red teaming, adversarial testing, security research, benchmark development, or rigorous model evaluation • Working proficiency in Python and Git • Ability to develop scripts, probes, and analyses independently • Strong familiarity with LLM capabilities, limitations, and evaluation techniques • Excellent written communication and technical documentation skills • Creativity, precision, and persistence when working through ambiguous research problems • Reliable availability for approximately 35 hours per week • A master's degree or PhD in a STEM field is highly relevant

Apply Now

Similar Jobs

đź•’ July 27

Books Are Fun

51 - 200

📚 Education

🤲 Charity

Outreach Specialist connecting K-6 schools to the Book Blast program through presentations and meetings. Focused on scheduling with decision-makers to promote student home libraries.

🇺🇸 United States – Remote

đź’µ $25 / hour

⏰ Full Time

🟢 Junior

🚫👨‍🎓 No degree required

đź•’ July 27

Horizon Connect @ Wall BCBSNJ

2 - 10

⚕️ Healthcare Insurance

🏥 Healthcare

Clinical role assessing patient needs and coordinating care services at Horizon Blue Cross Blue Shield. Requires NJ RN license and clinical experience.

🇺🇸 United States – Remote

đź’µ $70.5k - $94.4k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🚫👨‍🎓 No degree required

đź•’ July 27

Divided Sky Foundation

11 - 50

🤲 Charity

🏥 Healthcare

🌍 Social Impact

Development & Fundraising Coordinator at Divided Sky Foundation supporting donor relationships and fundraising efforts. Working to enhance recovery through community-driven initiatives in a team environment.

🇺🇸 United States – Remote

đź’µ $28 - $34 / hour

⏰ Full Time

🟢 Junior

🟡 Mid-level

đź•’ July 27

MedPOINT Management

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

Applications Specialist enhancing healthcare solutions for MedPOINT Management. Providing technical support and collaborating with teams to optimize application performance in a remote role.

🇺🇸 United States – Remote

đź’µ $70k - $85k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

đź•’ July 27

Kaukahi, LLC

2 - 10

đź’Ľ Consulting

📦 Logistics

🤝 B2B

Real Property Specialist Jr. managing complex real property functions for government. Responsible for inventory maintenance and regulatory compliance in a remote capacity.

🇺🇸 United States – Remote

đź’µ $75k - $90k / year

⏰ Full Time

🟢 Junior