Machine Learning Engineer – Model Evaluation, Experimentation

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mercor

Mercor

51 - 200 employees

Founded 2023

🔥 Funding within the last year

💰 $350M Series C - Mercor on 2025-10

Mercor is a company for which no descriptive text was provided in the input. Additional information (products, services, target customers, or industry specifics) is needed to create an accurate summary and select appropriate industries.

📋 Description

• Design well-defined, multi-step machine-learning tasks from real ML research ideas • Implement changes, run training experiments, and analyze results to establish correct solutions • Build tasks around reinforcement-learning concepts such as reward functions and training behavior • Evaluate frontier models and identify where and why they fall short • Compare notes with researchers and fellow experts to maintain consistency, rigor, and fairness • Work in a tight feedback loop with the lab’s researchers

🎯 Requirements

• MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain • 1+ years of experience in a research or research-engineering role • Hands-on experience training and evaluating ML models and running experiments end-to-end, including setup, execution, and analysis • Strong familiarity with large language models, including their capabilities, limitations, and evaluation techniques • Working proficiency in Python and Git • Comfort working in scripting and notebook environments • Basic understanding of reinforcement learning, including reward functions and policy training, preferred • Past experience in AI training, model evaluation, or benchmark/task authoring preferred • High attention to detail and creativity in task design • Strong written communication skills • Ability to work independently through ambiguous, open-ended problems • Ability to engage reliably for approximately 35 hours per week

🏖️ Benefits

• W-2 employment • Payroll, benefits, and compliance administered by Cincinnatus LLC • Fully remote work within the United States • Approximately 35 hours per week • Opportunity to be placed at a leading AI lab as part of its extended workforce • Reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process

Apply Now

Similar Jobs

🕒 2 days ago

Torc Robotics

501 - 1000

🚘 Automotive

📦 Logistics

🚗 Transport

ML Engineer scaling Torc Robotics’ simulation platform for autonomous trucks. Embedding with Autonomy teams to operationalize replay, recompute, metrics, visualization, and model integrations.

🕒 2 days ago

GLOBO

1001 - 5000

🏥 Healthcare

⚖️ Legal

📣 Marketing

AI/ML Engineer building data pipelines and intelligent features for GLOBO, a B2B translation and interpretation technology platform. Maintaining reliable, safe, cost-efficient production AI systems.

🕒 2 days ago

The Home Depot

10,000+ employees

🏗️ Construction

📦 Logistics

🛒 Retail

Machine Learning Engineer II developing and deploying production ML solutions for Home Depot, the world’s largest home improvement specialty retailer. Building scalable models, APIs, monitoring, and data pipelines.

🇺🇸 United States – Remote

💵 $100k - $150k / year

💰 Debt Financing on 2007-07

⏰ Full Time

🟢 Junior

🤖 Machine Learning Engineer

🚫👨‍🎓 No degree required

🕒 2 days ago

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Machine Learning Engineer developing real-time underwriting models for Affirm’s buy-now-pay-later platform. Building feature pipelines, production ML systems, and monitoring workflows for repayment-risk decisions.

🕒 3 days ago

Zillow

5001 - 10000

🏠 Real Estate

🛍️ eCommerce

👥 B2C

Machine Learning Engineer building production agentic AI systems for Zillow, a U.S. real estate platform. Advancing reasoning, orchestration, and scalable ML infrastructure for home-buying experiences.