LLM Red Team & Benchmark Evaluation Specialist

🔥 14 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

💼 Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

📋 Description

• Adversarial Model Evaluation • Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks • Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs • Investigate situations where models appear capable while reaching incorrect or unsupported conclusions • Design reproducible experiments to isolate and validate model failure modes • Benchmark & Challenge Design • Convert observed model weaknesses into rigorous benchmark tasks • Develop challenges that are technically demanding while remaining fair and objectively assessable • Define clear task requirements, expected outcomes, and evaluation criteria • Ensure tasks require genuine reasoning rather than allowing shortcuts or superficial pattern matching • Failure Analysis & Documentation • Document findings with clear evidence, methodology, and reproducible steps • Explain why a model failed and which capabilities or assumptions contributed to the error • Produce detailed technical write-ups for researchers and task authors • Track recurring failure patterns across models, prompts, and evaluation environments • Task Strengthening & Research Collaboration • Work with task authors to close loopholes, grading gaps, and unintended shortcuts • Review benchmark tasks for ambiguity, exploitability, and evaluation reliability • Share insights with researchers and other specialists to improve benchmark coverage • Participate in iterative calibration, peer review, and task-refinement workflows

🎯 Requirements

• At least 1 year of experience in research, research engineering, security, AI evaluation, or a related technical role • Demonstrated experience identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems • Background in red teaming, adversarial testing, security research, benchmark development, or rigorous model evaluation • Working proficiency in Python and Git • Ability to develop scripts, probes, and analyses independently • Strong familiarity with LLM capabilities, limitations, and evaluation techniques • Excellent written communication and technical documentation skills • Creativity, precision, and persistence when working through ambiguous research problems • Reliable availability for approximately 35 hours per week • A master's degree or PhD in a STEM field is highly relevant

Apply Now

Similar Jobs

🔥 15 minutes ago

Books Are Fun

51 - 200

📚 Education

🤲 Charity

Outreach Specialist connecting K-6 schools to the Book Blast program through presentations and meetings. Focused on scheduling with decision-makers to promote student home libraries.

🔥 17 minutes ago

Divided Sky Foundation

11 - 50

🤲 Charity

🏥 Healthcare

🌍 Social Impact

Development & Fundraising Coordinator at Divided Sky Foundation supporting donor relationships and fundraising efforts. Working to enhance recovery through community-driven initiatives in a team environment.

🇺🇸 United States – Remote

💵 $28 - $34 / hour

⏰ Full Time

🟢 Junior

🟡 Mid-level

🔥 27 minutes ago

MedPOINT Management

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

Applications Specialist enhancing healthcare solutions for MedPOINT Management. Providing technical support and collaborating with teams to optimize application performance in a remote role.

🇺🇸 United States – Remote

💵 $70k - $85k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🔥 37 minutes ago

Kaukahi, LLC

2 - 10

💼 Consulting

📦 Logistics

🤝 B2B

Real Property Specialist Jr. managing complex real property functions for government. Responsible for inventory maintenance and regulatory compliance in a remote capacity.

🇺🇸 United States – Remote

💵 $75k - $90k / year

⏰ Full Time

🟢 Junior

🔥 2 hours ago

Agiliti

5001 - 10000

🏥 Healthcare

📦 Logistics

🤝 B2B

Surgical Services Specialist driving strategic sales initiatives in assigned territory for Agiliti. Identifying and closing new business bookings opportunities while managing customer relationships.