Machine Learning Engineer – Model Evaluation, Experimentation

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Weekday (YC W21)

Weekday (YC W21)

11 - 50 employees

Founded 2021

💼 Consulting

👥 HR Tech

☁️ SaaS

Consulting • HR Tech • SaaS

Weekday is a modern recruitment platform that combines AI technologies with a vast database of potential candidates, aiming to streamline the hiring process for companies in India. They offer various services, including a proactive outreach approach that helps employers connect with top talent, as well as tools for candidates to easily apply for jobs. Weekday's emphasis on candidate engagement through multiple channels, including email, WhatsApp, and phone calls, sets it apart in the competitive landscape of recruitment agencies.

📋 Description

• Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis. • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria. • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes. • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable. • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions. • Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency.

🎯 Requirements

• Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline. • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role. • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows. • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis. • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies. • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments. • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred. • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable. • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently. • Strong written communication skills for documenting experimental methodologies and technical findings. • Ability to commit approximately 35 hours per week on a consistent basis.

Apply Now