Research Engineer – RL Infrastructure

🕒 vor 4 Monaten

🏄 California – Remote

info

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

📚 Forschungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Prime Intellect

Prime Intellect

1 - 10 Mitarbeiter

🤖 Künstliche Intelligenz

☁️ SaaS

Artificial Intelligence • SaaS • Cloud Computing

Prime Intellect ist ein Unternehmen, das sich darauf konzentriert, die KI-Entwicklung zu demokratisieren, indem es skalierbare und dezentrale Computerressourcen für das Training von Modellen bereitstellt. Ihre Plattform ermöglicht es Nutzern, weltweit verfügbare Computerressourcen zu finden und zu teilen, um den Training von hochmodernen Modellen durch verteilte Cluster zu ermöglichen. Sie fördern das kollektive Eigentum an KI-Innovationen, einschließlich Sprach- und wissenschaftlicher Modelle. Prime Intellect bietet zudem eine Reihe von GPU-Optionen, um kostengünstiges und effizientes Modelltraining zu erleichtern. Sie streben danach, die Forschung im Bereich des dezentralen Trainings und die Entwicklung von Open-Source-KI weltweit voranzutreiben.

Beschreibung

• Build and optimize the systems infrastructure behind large-scale RL and distributed training workloads. • Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers. • Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements. • Work on distributed training systems spanning data, tensor, and pipeline parallel workloads. • Help shape the architecture of our RL training stack, including async rollout and post-training systems. • Contribute to open-source libraries and internal infrastructure used for frontier-scale model training. • Collaborate closely with researchers and infrastructure engineers to translate bottlenecks into concrete systems improvements. • Stay at the frontier of training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.

🎯 Anforderungen

• Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference. • Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling. • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy. • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism. • Strong understanding of GPU architecture, profiling, and performance debugging. • Ability to identify bottlenecks across the stack and drive improvements from first principles. • Comfort working in a fast-moving environment with ambiguous problems and high ownership.

🏖️ Vorteile

• Competitive compensation, including equity. • Flexible work arrangements, with the option to work remotely or in person from our San Francisco office. • Visa sponsorship and relocation support for international candidates. • Quarterly team offsites, hackathons, conferences, and learning opportunities. • A deeply technical, high-agency team working on infrastructure for open superintelligence.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 4 Monaten

Deepgram

51 - 200

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Machine Learning Engineer at Deepgram prototyping novel modeling ideas to build scalable speech technologies. Collaborate with researchers to develop world-class speech recognition and synthesis systems.

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $250.000 / Jahr

💰 €47.000.000 Series B im 2022-11

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

📚 Forschungsingenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Rwazi

11 - 50

🤖 Künstliche Intelligenz

☁️ SaaS

🤝 B2B

Research Engineer building technical infrastructure for decision research. Developing prototypes and evaluation tooling in a remote flexible environment.

🇺🇸 Vereinigte Staaten – Remote

💰 €12.000.000 Series A - Rwazi im 2025-07

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

📚 Forschungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Censys

51 - 200

🔒 Cybersecurity

🏢 Unternehmen

Senior Security Research Engineer conducting Internet measurement research at Censys. Partnering with teams to deliver new security prototypes and analyze large datasets.

🇺🇸 Vereinigte Staaten – Remote

💵 $153.000 - $212.000 / Jahr

⏰ Vollzeit

🟠 Senior

📚 Forschungsingenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

BigQuery

🕒 vor 6 Monaten

Bugcrowd

201 - 500

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Exploit Development Specialist creating novel vulnerability discovery for threat actors. Focused on real-world exploit development utilizing expert reverse engineering and complex system vulnerabilities.

🇺🇸 Vereinigte Staaten – Remote

💵 $154.800 - $193.500 / Jahr

💰 €30.000.000 Series D im 2020-04

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

📚 Forschungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Anthropic

11 - 50

🤖 Künstliche Intelligenz

☁️ SaaS

🏢 Unternehmen

Research Engineer at Anthropic developing training environments for agentic AI. Collaborating on research and engineering responsibilities aiming for safe and beneficial AI systems.

🇺🇸 Vereinigte Staaten – Remote

💵 $500.000 - $800.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

📚 Forschungsingenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich