Senior Software Engineer – Model Training, AI Evals

Job not on LinkedIn

🕒 June 29

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 44%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Chegg Inc.

Chegg Inc.

1001 - 5000 employees

🛍️ eCommerce

📚 Education

☁️ SaaS

eCommerce • Education • SaaS

Chegg Inc. is a software development company that provides individualized learning support to students pursuing their educational journeys. Available on demand 24/7, the Chegg platform is powered by over a decade of learning insights and offers AI-powered academic support. It connects students with a vast network of subject matter experts to ensure quality learning and helps millions of students build essential academic, life, and job skills for success.

📋 Description

• Own evaluation solutions end-to-end, including methodology design, pipeline and tooling development, deployment into training and production workflows, and ongoing maintenance • Design and build evaluation frameworks and harnesses for LLMs and multi-step AI agents • Define rubrics and scoring methodologies for open-ended tasks such as tutoring quality, reasoning, and citation accuracy • Build automated model-graded and rule-based evaluators and validate them against human judgment • Develop agent-specific evaluations for tool use, multi-turn completion, planning, failure recovery, and cost/latency tradeoffs • Create golden datasets, adversarial test sets, and regression suites • Translate evaluation results into training signals for SFT/RLHF/RLAIF data curation, reward modeling, and fine-tuning priorities • Instrument production systems to collect interaction data for evaluation and training • Build dashboards and reporting on model quality trends across releases • Ensure methodological rigor through statistical significance, inter-rater reliability, sampling strategy, and resistance to metric gaming • Collaborate with Trust & Safety and Legal/Compliance on responsible-AI evaluations • Mentor engineers and establish evaluation as a core part of model development

🎯 Requirements

• 5+ years of software/ML engineering experience, including hands-on work building or maintaining evaluation, testing, or measurement infrastructure for ML systems • Direct experience with foundation model evaluation and benchmarking beyond using APIs; experience at a foundation model lab or similarly frontier research environment strongly preferred • Demonstrated ability to own an evaluation solution end-to-end, from design and dataset/methodology creation through pipeline build, deployment, and production monitoring • Direct experience designing evals for LLMs and/or AI agents • Strong programming skills in Python, PyTorch / TensorFlow, and production-grade data/ML pipelines • Hands-on knowledge of AWS, including SageMaker and Bedrock, and Databricks, including MLflow, Unity Catalog, and Delta Lake • Applied statistics knowledge, including sample size, variance, significance testing, and aggregate metric limitations • Working knowledge of LLM training and adaptation, including pretraining, SFT, RLHF/RLAIF, and DPO • Experience with agentic architectures, including tool calling, multi-step planning, memory, and orchestration frameworks • Familiarity with evaluation and observability tooling, tracing, dataset versioning, and experiment tracking • Excellent cross-functional collaboration skills • Bias toward rigor and skepticism when evaluating metrics • Bonus: post-training techniques and evaluation methodologies, model benchmarking, and data quality practices such as curation, filtering, deduplication, and quality scoring

🏖️ Benefits

• Salary and benefits per Chegg's compensation bands and the candidate's location, in accordance with applicable pay transparency requirements • Inclusive work environment • Equal opportunity employment

Apply Now

Similar Jobs

🕒 June 29

Playpower Labs

11 - 50

📚 Education

🤖 Artificial Intelligence

Senior Software Engineer developing full-stack features using Angular/React and Java/Node.js for EdTech companies. Aiming to enhance learner experiences using advanced technology and engineering.

Angular

Cloud

Docker

Java

JavaScript

Node.js

React

TypeScript

🕒 June 29

Exavalu

201 - 500

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Azure Tech Lead designing Azure API gateways and middleware integrations for an IT services company. Leading project delivery, resource allocation, and enterprise system connectivity.

Azure

🕒 June 25

HighLevel

201 - 500

💼 Consulting

📦 Logistics

☁️ SaaS

Senior FullStack Engineer building authentication, authorization, and permission systems for HighLevel’s AI-powered business operating system. Scaling secure multi-tenant SaaS infrastructure for millions of users.

Cloud

ElasticSearch

Google Cloud Platform

JavaScript

Microservices

MongoDB

Node.js

SQL

TypeScript

🕒 June 24

Gururo

11 - 50

💼 Consulting

📣 Marketing

📚 Education

Software Developer collaborating with teams to write efficient algorithms and software code. Testing and upgrading existing software, creating technical documentation.

Angular

Java

🕒 June 24

Twilio

5001 - 10000

🔌 API

🤝 B2B

Tech Lead developing AI prototypes and production-ready solutions at Twilio, which powers personalized communications for businesses and developers. Driving frontier-technology R&D, architecture, experimentation, and engineering standards for autonomous communication products.

Angular

AWS

Azure

Java

JavaScript

Node.js

NoSQL

Python

React

Spring Boot

SpringBoot

SQL

Go