Senior ML Systems Engineer, Frameworks & Tooling

Stelle nicht auf LinkedIn

🕒 vor 8 Monaten

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🇬🇧 UK-Skilled-Worker-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Cohere

Cohere

11 - 50 Mitarbeiter

🤖 Künstliche Intelligenz

🏢 Unternehmen

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Cohere ist eine führende KI-Plattform, die Unternehmen fortschrittliche Sprachmodelle und einen integrierten Arbeitsbereich bietet, der auf Effizienz und Sicherheit ausgelegt ist. Mit einer Reihe von leistungsstarken generativen und Retrieval-Modellen ermöglicht Cohere Organisationen die Optimierung von Arbeitsabläufen, die Verbesserung der Datensicherheit und das Erschließen von Erkenntnissen über verschiedene Branchen hinweg durch mehrsprachige Fähigkeiten. Ihr Fokus auf maßgeschneiderte KI-Lösungen gewährleistet den Schutz kritischer Daten und erleichtert die nahtlose Integration in bestehende organisatorische Prozesse.

Beschreibung

• Build and own the training framework responsible for large-scale LLM training. • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. • Collaborate closely with infra teams to ensure Slurm setups, container environments, and hardware configurations support high-performance training. • Investigate and resolve performance bottlenecks across the ML systems stack. • Build robust systems that ensure reproducible, debuggable, large-scale runs.

🎯 Anforderungen

• Strong engineering experience in large-scale distributed training or HPC systems. • Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. • Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. • Experience working with containerized environments (Docker, Singularity/Apptainer). • A track record of building tools that increase developer velocity for ML teams. • Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability. • Strong collaboration skills — you’ll work closely with infra, research, and deployment teams.

🏖️ Vorteile

• An open and inclusive culture and work environment • Work closely with a team on the cutting edge of AI research • Weekly lunch stipend, in-office lunches & snacks • Full health and dental benefits, including a separate budget to take care of your mental health • 100% Parental Leave top-up for up to 6 months • Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement • Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend • 6 weeks of vacation (30 working days!)

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 10 Monaten

BPM LLP

501 - 1000

💸 Finanzen

💼 Beratung

☁️ SaaS

AI/ML Engineer at BPM developing AI-powered applications for operational efficiency and client collaboration. Join a global team in a dynamic advisory firm focusing on innovative solutions.

🇬🇧 Vereinigtes Königreich – Remote

💵 $150.000 - $180.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich