
11 - 50 employees
Founded 2022
🤖 Artificial Intelligence
☁️ SaaS
Artificial Intelligence • Research • SaaS
METR is a research organization focused on evaluating and understanding the capabilities and risks of frontier AI systems. They develop and run tests to measure AI performance, particularly in autonomous capabilities and AI R&D. By collaborating with leading AI companies like Anthropic and OpenAI, METR aims to enhance third-party evaluations to manage the emerging risks of advanced AI technologies comprehensively.
🔥 0 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2022
🤖 Artificial Intelligence
☁️ SaaS
Artificial Intelligence • Research • SaaS
METR is a research organization focused on evaluating and understanding the capabilities and risks of frontier AI systems. They develop and run tests to measure AI performance, particularly in autonomous capabilities and AI R&D. By collaborating with leading AI companies like Anthropic and OpenAI, METR aims to enhance third-party evaluations to manage the emerging risks of advanced AI technologies comprehensively.
• Develop difficult, novel tasks for AI models that remain challenging as model time horizons grow • Perform quality assurance for existing tasks, verifying solvability and that models receive only the necessary information • Baseline tasks within areas of expertise when helpful • Score task completions from AIs or human baseliners • Improve task development infrastructure and workflows • Contribute to evaluations of frontier AI systems supporting METR’s Time Horizons methodology
• Several years of experience working on complex software engineering projects and codebases • Experience building hard, ideally agent-based, AI evaluations • Experience with evaluations such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA • Ideally experience using the Inspect framework • High attention to detail, including spotting misspecifications and ambiguity • Familiarity with METR infrastructure, including Hawk, is a plus • Familiarity with the methodology behind METR’s Time Horizons work is a plus • Availability for 20–40 hours per week • Minimum of 1 hour of overlap with the Pacific Coast Time workday
• Flexible schedule determined by you • Remote work worldwide • Flexible working hours of 20–40 hours per week • Individuals who contribute >80 hours will be acknowledged in the final research output (if desired)
Apply Now🕒 June 26
Forward Deployed Engineer at Vozy deploying and optimizing AI voice agents for customer resolutions. Collaborate with clients and manage AI integrations in a dynamic environment.
🌏 Anywhere in the World
💰 $1M Debt Financing on 2022-10
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
👷🏻♀️ Engineer
🗣️🇪🇸 Spanish Required
Java
JavaScript
Python
🕒 November 12, 2025
Seafarer for Columbia Shipmanagement's offshore expanding fleet as a 2nd Engineer on Cable Layer vessels. Collaborate with multicultural teams and adhere to safety standards.