Research Intern – Applied Reinforcement Learning

Job not on LinkedIn

🕒 July 3

🏄 California, Washington – Remote

info

💵 $35 - $45 / hour

👨‍🎓 Internship

⚪️ Entry-level

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Thermo Fisher Scientific

Thermo Fisher Scientific

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Healthcare • Consulting • Logistics

Thermo Fisher Scientific is a leading global supplier of scientific instrumentation, reagents and consumables, and software services. They support the life sciences, healthcare, and analytical chemistry sectors by providing robust solutions for laboratory research and production processes. Their innovative products and services encompass a range of applications, including diagnostics, lab workflow automation, and drug discovery.

📋 Description

• Design and evaluate reinforcement learning (RL) systems for agentic AI workflows • Develop RL environments, reward models, and post-training pipelines for LLM-based agents • Create end-to-end RL pipelines for agentic systems (simulation → training → evaluation) • Align LLM-based agents using RLHF, DPO, PPO, and emerging methods • Design reward functions, verifiers, and evaluation frameworks • Build simulation environments (digital twins) for enterprise workflows • Ensure scalable training and inference for RL-based systems • Document experiments, ablations, and findings for research and productionization

🎯 Requirements

• PhD candidate in CS, ML, or related field with research in reinforcement learning or agentic AI • Strong Python and PyTorch skills with GPU-based training experience • Solid understanding of RL fundamentals (MDPs, policy gradients, value methods) • Experience with LLMs and post-training techniques (RLHF, DPO, PPO, etc.) • Strong experimentation practices (ablation, reproducibility, clear reporting) • Experience with RL environments (Gymnasium, RLlib, Stable Baselines) (preferred) • Research in offline RL, model-based RL, or hierarchical RL (preferred) • Publications at top ML conferences (NeurIPS, ICML, ICLR, ACL) (preferred) • Experience with simulation, synthetic data, or multi-agent systems (preferred) • Distributed training and large-scale experimentation (preferred)

🏖️ Benefits

• Competitive stipend • Mentorship from researchers and engineers • Access to modern GPU infrastructure • Opportunities to publish and present research

Apply Now

Similar Jobs

🕒 June 25

Americans United for Separation of Church and State

11 - 50

⚖️ Legal

💼 Consulting

🤝 Non-profit

Legal intern role supporting civil rights litigation at Americans United. Involves research and drafting legal documents while working remotely from anywhere in the United States.

🇺🇸 United States – Remote

💵 $7.4k / year

👨‍🎓 Internship

⚪️ Entry-level

🚫👨‍🎓 No degree required

🕒 June 25

Oddin.gg

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Applied Science Intern at Valka exploring World Models in AI video generation. Collaborate with diverse innovators to redefine digital content creation.

🕒 June 25

Encephalo

1 - 10

🏗️ Construction

🏥 Healthcare

📦 Logistics

Equity Research Intern at Encephalo supporting finance research initiatives in a remote role. Engaging with industry groups and collaborating closely with analysts on quantitative research.

🕒 June 25

BLACKCLOAK

11 - 50

🔒 Cybersecurity

☁️ SaaS

SkillBridge Internship for active duty service-members at BlackCloak. Gain cybersecurity experience while maintaining active duty service status with hands-on training.

🇺🇸 United States – Remote

💰 $11M Series A on 2021-07

👨‍🎓 Internship

⚪️ Entry-level

🚫👨‍🎓 No degree required

🕒 June 25

ProSidian Consulting

11 - 50

📦 Logistics

🏭 Manufacturing

🛡️ Insurance

Market Development Intern supporting marketing and business development initiatives for ProSidian. Interns gain hands-on experience with various market research aspects, helping expand company market presence.