AI/ML Research Engineer, LLM Post-Training – Evaluation

Job not on LinkedIn

🔥 2 hours ago

🇺🇸 United States – Remote

💵 $80k - $175k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🗣️ LLM Engineer

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of InnoData

InnoData

2 - 10 employees

Founded 2019

🤝 B2B

💼 Consulting

🌍 Social Impact

B2B • Consulting • Social Impact

InnoData is INNOvation DATA SCS, an Italian social cooperative based in Foggia that identifies itself as a provider of technological solutions. The company website (currently under maintenance) highlights "Soluzioni tecnologiche" (technological solutions) and emphasizes social impact ("Impatto sociale"). Contact details listed include Via Francesco Crispi 65, 71121 Foggia, Italy. Based on the available information, InnoData appears to operate at the intersection of technology and social impact, likely offering tech-focused services to other organizations.

📋 Description

• design and implement the pipelines and tooling that connect data, evaluation, and post-training • help customers and internal teams move from evaluation findings to measurable model improvements • build fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization) • integrate evaluation harnesses into model development loops • improve experiment reliability and throughput • support advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions • contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation. • Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery • Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking • Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses • Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows • Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring • Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift • Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems • Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions • Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements • Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects • Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team.

🎯 Requirements

• BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field (MS/PhD preferred) • 2-3 years of relevant industry or research engineering experience in ML/AI systems • Hands-on experience with LLM training / fine-tuning / post-training, including at least one of: • supervised fine-tuning (SFT) • preference optimization (e.g., DPO or related methods) • RLHF / RLAIF-style workflows • task- or domain-adaptation of foundation models • Strong programming skills in Python and experience building production-quality ML code • Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks) • Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons • Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging • Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads (GPU/accelerator environments preferred) • Experience with large-scale data processing and workflow orchestration in support of model training/evaluation • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads • Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences.

Apply Now

Similar Jobs

🔥 6 hours ago

EWOR

201 - 500

📚 Education

💸 Finance

💼 Consulting

Co-Founder / CEO responsible for launching and scaling AI Infrastructure startups in a supportive environment. Engage with experienced entrepreneurs while developing a high-growth venture.

🇺🇸 United States – Remote

💰 $34.2M Series A - EWOR on 2025-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

🗣️ LLM Engineer

🔥 6 hours ago

EWOR

201 - 500

📚 Education

💸 Finance

💼 Consulting

Co-Founder / CCO leading an AI Infrastructure startup backed by experienced entrepreneurs. Building startup with coaching and funding support to reach significant revenues.

🇺🇸 United States – Remote

💰 $34.2M Series A - EWOR on 2025-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

🗣️ LLM Engineer

🕒 July 14

Tiger Analytics

1001 - 5000

🏥 Healthcare

📦 Logistics

📣 Marketing

Forward Deployed Engineer at Tiger Analytics driving deployment and scaling of Generative AI solutions. Collaborate with teams to operationalize AI research into production-grade infrastructure.

🕒 June 29

Vultr

201 - 500

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

Account Executive at Vultr responsible for driving growth in AI Infrastructure Sales. Own strategic customer relationships and guide clients through AI cloud infrastructure adoption.

🇺🇸 United States – Remote

💵 $90k - $110k / year

💰 $329M Debt Financing - Vultr on 2025-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

🗣️ LLM Engineer

🕒 June 27

Netflix

10,000+ employees

📱 Media

👥 B2C

Research Engineer at Netflix developing LLM prototypes to enhance member understanding. Collaborating with cross-functional teams in a dynamic and innovative environment.