Senior Principal AI Engineer

🕒 August 6

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🤖 AI Engineer

👻 Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cerence Inc.

Cerence Inc.

1001 - 5000 employees

Founded 2019

💼 Consulting

📦 Logistics

🏭 Manufacturing

💰 Grant on 2020-12

Consulting • Logistics • Manufacturing

Cerence Inc. is a global company focused on providing AI-powered solutions, particularly in the automotive industry. They specialize in conversational and generative AI technologies that create intelligent, natural, and personalized interactions between humans and vehicles. With innovations like their proprietary automotive large language models, Cerence enhances user experiences across various forms of transport including cars, two-wheelers, and trucks. The company has over 500 million vehicles shipped with its AI technology, serving more than 80 OEMs and Tier 1 customers worldwide. Cerence is dedicated to continuous advancements in AI, aiming to revolutionize in-car user experiences through fast delivery and seamless integration of their solutions.

📋 Description

• Design and operate distributed training systems for large neural networks across GPU clusters • Optimize multi-node, multi-GPU execution for throughput and utilization • Diagnose and resolve compute, memory, and networking bottlenecks • Improve training stability and fault tolerance at scale • Partner with research and applied ML teams to productionize large-model training pipelines • Build and optimize GPU cluster orchestration with Slurm, Kubernetes, Ray, and RunAI • Ensure efficient scheduling, isolation, and fairness across training workloads • Optimize and debug distributed communication using NCCL, RDMA, InfiniBand, and NVLink • Scale large-model training with PyTorch Distributed, Megatron-LM, and DeepSpeed • Own multi-node launch configurations, failure recovery, and performance tuning • Apply activation checkpointing, ZeRO Stage 1–3, and offload strategies • Balance compute, memory, and communication to increase model size and batch scale • Eliminate GPU utilization, networking, communication, instability, and large-scale training failure modes • Enable larger models, faster iteration cycles, and more reliable research-to-production pipelines

🎯 Requirements

• Deep hands-on experience with distributed systems or ML systems • Experience running large-scale workloads on GPU clusters • Production experience with PyTorch distributed training • Strong understanding of data, tensor, and pipeline parallelism • Low-level understanding of GPU communication and networking • Experience with Slurm, Kubernetes, Ray, and RunAI • Experience with NCCL, RDMA, InfiniBand, and NVLink • Experience with PyTorch Distributed, Megatron-LM, and DeepSpeed • Experience with activation checkpointing, ZeRO Stage 1–3, and offload techniques • Experience working with large language models or foundation models is a strong plus • Deep systems expertise valued over pure model architecture experience

Apply Now

Similar Jobs

🕒 August 6

GE Aerospace

10,000+ employees

🏭 Manufacturing

🎖️ Defense

💼 Consulting

AI/ML Engineer building production LLM applications, models, APIs, and MLOps capabilities for GE Aerospace operational decision-making. Driving AI solutions for Commercial Engine Services.

🕒 August 6

GE Aerospace

10,000+ employees

🚀 Aerospace

🏭 Manufacturing

🎖️ Defense

Remote AI/ML Engineer building production LLM applications, ML models, APIs, and MLOps capabilities. Transforming GE Aerospace operational data into decision-support solutions.

🇺🇸 United States – Remote

💵 $112k - $150k / year

💰 $2G Post-IPO Debt - GE Aerospace on 2025-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 AI Engineer

🕒 August 5

Dataminr

501 - 1000

🤖 Artificial Intelligence

🔐 Security

📱 Media

Senior Director leading AI engineering for Dataminr’s real-time intelligence platform. Translating advanced AI research into scalable solutions for detecting global events and risks.

🕒 August 5

Tealium

501 - 1000

💼 Consulting

📣 Marketing

🏥 Healthcare

Senior engineer building AI-powered developer tools for Tealium’s real-time customer data platform. Driving modernization, automated testing, code quality, and engineering productivity across R&D.

🕒 August 5

Agilent Technologies

10,000+ employees

🍽️ Food & Beverage

🏥 Healthcare

💼 Consulting

AI engineering lead shaping agentic tools, copilots, and SDLC automation at Agilent Technologies. Improving developer productivity, software quality, security, and delivery across PSD.