Machine Learning Engineer, Evals

Job not on LinkedIn

🕒 July 28

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nous Research

Nous Research

11 - 50 employees

Founded 2023

🤖 Artificial Intelligence

🔬 Science

💰 $50M Series A on 2025-05

Artificial Intelligence • Science

Nous Research is a US-based leader in the open-source AI movement that trains large open-source language models and builds infrastructure to coordinate distributed, unbiased training. Its stated mission is to advance human rights and freedoms by creating and proliferating open-source language models, supporting their unrestricted availability and use, and furthering scientific and popular understanding of these models. The company focuses on applied AI research areas including model architecture, data synthesis, fine-tuning, and reasoning, and provides tooling and platforms (e. g. , Hermes, Nous Portal, Psyche, Nous Chat, simulators) to accelerate model development and deployment.

📋 Description

• Run the full eval pipeline end to end and reproduce known results during onboarding, pairing with a senior engineer on the first task • Build a judge calibration protocol using human-labeled decisions, agreement metrics, drift-zone identification, and reproducible documentation • Extend benchmarks such as GAIA, τ-Bench, and SWE-bench with tasks targeting capability gaps, including prompts, environments, rubrics, automated graders, and QA • Analyze model-output failures, categorize failure modes, quantify prevalence, and recommend changes to training data, judge prompts, or benchmarks • Own recurring evaluation workflows such as regression suites, judge-drift dashboards, and red-team evaluations • Ship evaluation tooling used by researchers

🎯 Requirements

• 3+ years in software engineering, ML engineering, data science, or a research-adjacent role • Concrete evaluation experience from coursework, an internship, a side project, open source work, or a job • Experience with at least one LLM evaluation framework, such as Harbor or Nemo Evaluator • Hands-on LLM experience with prompting and few-shot design; ideally fine-tuning or RAG; regular use of coding agents • Solid Python and ability to write clean, tested, version-controlled code • Comfort with Git, CI/CD basics, Docker, and the Linux command line, including SSH, tmux, and debugging remote jobs • Understanding of basic evaluation statistics, including accuracy on imbalanced judges, Cohen's κ, and confidence intervals • At least 3 specified evaluation competencies involving judge calibration, failure analysis, agent benchmarks, eval dataset design, or non-determinism and variance reporting • Clear communication with researchers and engineers • Comfort with ambiguity, planning from half-formed requests, and knowing when to ask for help • Preferred: RLVR/RLHF pipeline experience • Preferred: training data curation experience • Preferred: distributed eval orchestration experience • Preferred: benchmark design from scratch • Preferred: red teaming and adversarial evaluation experience • Preferred: familiarity with psychometrics or measurement theory

Apply Now

Similar Jobs

🕒 July 28

The Voleon Group

201 - 500

💸 Finance

🤖 Artificial Intelligence

Senior Machine Learning Engineer at Voleon utilizing AI and ML for quantitative trading. Collaborating with researchers to build and maintain data pipelines and machine learning models.

🇺🇸 United States – Remote

💵 $290k - $395k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

Linux

Numpy

Pandas

Python

PyTorch

Scikit-Learn

Tensorflow

🕒 July 28

Paradigm Health

51 - 200

🏥 Healthcare

💊 Pharmaceuticals

☁️ SaaS

Senior Machine Learning Engineer developing sophisticated ML models at Paradigm Health. Collaborating with cross-functional teams to enhance clinical trial workflows and patient engagement.

🇺🇸 United States – Remote

💰 $203M Series A on 2023-02

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

🕒 July 28

Altis Labs

11 - 50

🤖 Artificial Intelligence

🏥 Healthcare

🧬 Biotechnology

Senior Machine Learning Scientist responsible for designing 3D volumetric imaging architectures in AI. Collaborates on innovative models and contributes to impactful research in oncology trials.

🇺🇸 United States – Remote

💵 $175k - $300k / year

💰 $6M Seed Round - Altis Labs on 2023-06

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

🕒 July 28

TRM Labs

201 - 500

₿ Crypto

📋 Compliance

🤝 B2B

Senior ML Systems Engineer building scalable AI and ML systems infrastructure for TRM Labs, focusing on Large Language Models and agentic systems development.

🇺🇸 United States – Remote

💵 $200k - $275k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

🕒 July 28

TRM Labs

201 - 500

₿ Crypto

📋 Compliance

🤝 B2B

Senior MLOps Engineer at TRM Labs focusing on building and scaling AI/ML infrastructure. Responsible for deploying models and ensuring reliability in a fast-paced environment.

🇺🇸 United States – Remote

💵 $200k - $275k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer