Research Engineer – Post-Training

Job not on LinkedIn

🔥 15 hours ago

🌐 United States, Australia – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

📚 Research Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Pluralis Research

Pluralis Research

1 - 10 employees

🤖 Artificial Intelligence

🌐 Web 3

Artificial Intelligence • Web 3

Pluralis Research is a foundational AI research lab focused on Protocol Learning — decentralized, multi‑participant training of foundation models where no single participant holds a full copy of the model. The group develops methods to enable communication‑efficient model and pipeline parallelism, unextractable collaborative models, and high‑compression context parallelism so community‑trained, community‑owned frontier models can scale over low‑bandwidth, internet‑connected devices. Their work targets practical systems and algorithms that make decentralized training competitive with centralized training while enabling new ownership and economic models.

📋 Description

• Build the RL post-training stack end-to-end, including rollout ingestion, reward computation, policy updates, and distributing updated weights across the network • Set the direction for the post-training stack and execute its development • Adapt RL algorithms to asynchronous, high-latency, partially trusted generation • Address staleness tolerance, off-policy corrections, and communication-efficient policy updates • Build evaluations demonstrating model improvement • Deliver the first decentralized post-trained model release as a public artifact

🎯 Requirements

• Hands-on experience running RL post-training on large language models, including RLHF, RLVR, or reasoning-focused RL • Experience with rollout generation, asynchronous training loops, and weight synchronization • Production-quality Python and PyTorch skills • Experience with concurrency, failure handling, and profiling before optimizing • Publications in RL post-training, asynchronous or distributed RL, or nearby fields, or unpublished work that can be defended in detail • Belief in Protocol Learning as a viable path for collective, trustless, and sovereign AI • Professional-level English proficiency, written and spoken • Comfort working across timezones • Nice to have: experience training over slow networks or with decentralized/federated setups • Nice to have: familiarity with vLLM or SGLang serving internals • Nice to have: experience with reward modeling or verifiable-reward datasets • Nice to have: experience with P2P networking and NAT traversal • Nice to have: experience at proprietary, open-weight, and open-source AI labs

🏖️ Benefits

• Significant equity ownership for key technical contributors in addition to a high base salary • Flexible work environment with team members distributed globally • Optional full visa sponsorship and relocation support to either Australia or the US

Apply Now

Similar Jobs

🕒 4 days ago

GuidePoint Security

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Innovation Engineer building secure AWS generative AI capabilities for GuidePoint Security, a cybersecurity services provider. Designing AI integrations, security controls, and operational processes for internal teams.

🕒 August 24

Clarity Innovations, Inc.

11 - 50

💼 Consulting

📣 Marketing

📚 Education

Senior principal engineer conducting reverse engineering, vulnerability research, and exploit development. Supporting Clarity Innovations’ Intelligence Community and Department of Defense national security missions.

🕒 August 20

SecurityScorecard

501 - 1000

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior research engineer turning threat intelligence research into production detections, feeds, and APIs. SecurityScorecard provides cybersecurity ratings and risk intelligence for organizations worldwide.

🕒 August 19

LiveKit

11 - 50

🔌 API

🤖 Artificial Intelligence

📡 Telecommunications

Research Engineer building reinforcement-learning models for LiveKit’s voice AI infrastructure. Developing training data, evaluations, and production models for voice, SMS, and chat agents.

🇺🇸 United States – Remote

💵 $135k - $300k / year

💰 Venture Round on 2022-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

📚 Research Engineer

🕒 August 12

Netflix

10,000+ employees

📱 Media

👥 B2C

Research Engineer developing and deploying machine-learning models for Netflix’s global streaming business. Improving member lifecycle and monetization through experimentation, algorithms, and scalable real-time systems.