Applied Research Scientist, LLM Evaluation – Post-Training

Job not on LinkedIn

🕒 July 27

🇺🇸 United States – Remote

💵 $175k - $225k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 28%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of InnoData

InnoData

2 - 10 employees

Founded 2019

🤝 B2B

💼 Consulting

🌍 Social Impact

B2B • Consulting • Social Impact

InnoData is INNOvation DATA SCS, an Italian social cooperative based in Foggia that identifies itself as a provider of technological solutions. The company website (currently under maintenance) highlights "Soluzioni tecnologiche" (technological solutions) and emphasizes social impact ("Impatto sociale"). Contact details listed include Via Francesco Crispi 65, 71121 Foggia, Italy. Based on the available information, InnoData appears to operate at the intersection of technology and social impact, likely offering tech-focused services to other organizations.

📋 Description

• Lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement. • Help define the next generation of evaluation-driven model improvement workflows. • Study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes. • Design experiments that produce credible, actionable conclusions. • Support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations. • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement. • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes. • Develop and validate evaluation frameworks for LLM and multimodal systems. • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign. • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines.

🎯 Requirements

• MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred) • 5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models • Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization) • Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly.

Apply Now

Similar Jobs

🕒 July 27

Arizona

201 - 500

📣 Marketing

📱 Media

Research Scientist IV developing and evaluating AI/NLP systems for scientific feasibility assessment. Leading design and implementation in a collaborative research environment with Python expertise.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

🕒 July 27

Predactiv

51 - 200

🤖 Artificial Intelligence

📣 Marketing

☁️ SaaS

Research Scientist at Predactiv developing next-generation AI technologies focused on user representation learning and generative AI applications. Conducting applied research and collaborating with engineering teams to translate research into production systems.

🇺🇸 United States – Remote

💵 $130k - $160k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

🕒 July 27

JHU EEHPC

11 - 50

📚 Education

🤖 Artificial Intelligence

🏥 Healthcare

Research Assistant overseeing data collection and management for research studies in educational technology. Collaborating with faculty and research teams to support data accuracy and reporting.

🇺🇸 United States – Remote

💵 $17 - $30 / hour

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

🕒 July 22

Thomson Reuters

10,000+ employees

💼 Consulting

⚖️ Legal

🛡️ Insurance

Senior Applied Scientist developing neural search systems for legal and professional content. Collaborating on retrieval architecture and contributing to major research publications.

🇺🇸 United States – Remote

💵 $137.1k - $254.7k / year

⏰ Full Time

🟠 Senior

🧬 Research Scientist

🦅 H1B Visa Sponsor

infoinfo

🕒 July 21

Samsara

1001 - 5000

📦 Logistics

🏗️ Construction

🏥 Healthcare

Applied Scientist role generating insights from customer interactions using ML / CV for Samsara. Opportunity to work with large volumes of sensor data for impactful solutions.

🇺🇸 United States – Remote

💵 $170.2k - $286k / year

💰 Seed Round on 2014-08

⏰ Full Time

🟠 Senior

🧬 Research Scientist

🦅 H1B Visa Sponsor

infoinfo