Senior Principal AI Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cerence Inc.

Cerence Inc.

1001 - 5000 employees

Founded 2019

💼 Consulting

📦 Logistics

🏭 Manufacturing

💰 Grant on 2020-12

Consulting • Logistics • Manufacturing

Cerence Inc. is a global company focused on providing AI-powered solutions, particularly in the automotive industry. They specialize in conversational and generative AI technologies that create intelligent, natural, and personalized interactions between humans and vehicles. With innovations like their proprietary automotive large language models, Cerence enhances user experiences across various forms of transport including cars, two-wheelers, and trucks. The company has over 500 million vehicles shipped with its AI technology, serving more than 80 OEMs and Tier 1 customers worldwide. Cerence is dedicated to continuous advancements in AI, aiming to revolutionize in-car user experiences through fast delivery and seamless integration of their solutions.

📋 Description

• Design and operate distributed training systems for large neural networks across GPU clusters • Optimize multi-node, multi-GPU execution for throughput and utilization • Diagnose and resolve compute, memory, and networking bottlenecks • Improve training stability and fault tolerance at scale • Partner with research and applied ML teams to productionize large-model training pipelines • Build and optimize GPU cluster orchestration with Slurm, Kubernetes, Ray, and RunAI • Ensure efficient scheduling, isolation, and fairness across training workloads • Optimize and debug distributed communication using NCCL, RDMA, InfiniBand, and NVLink • Scale large-model training with PyTorch Distributed, Megatron-LM, and DeepSpeed • Own multi-node launch configurations, failure recovery, and performance tuning • Apply activation checkpointing, ZeRO Stage 1–3, and offload strategies • Balance compute, memory, and communication to increase model size and batch scale • Eliminate GPU utilization, networking, communication, instability, and large-scale training failure modes • Enable larger models, faster iteration cycles, and more reliable research-to-production pipelines

🎯 Requirements

• Deep hands-on experience with distributed systems or ML systems • Experience running large-scale workloads on GPU clusters • Production experience with PyTorch distributed training • Strong understanding of data, tensor, and pipeline parallelism • Low-level understanding of GPU communication and networking • Experience with Slurm, Kubernetes, Ray, and RunAI • Experience with NCCL, RDMA, InfiniBand, and NVLink • Experience with PyTorch Distributed, Megatron-LM, and DeepSpeed • Experience with activation checkpointing, ZeRO Stage 1–3, and offload techniques • Experience working with large language models or foundation models is a strong plus • Deep systems expertise valued over pure model architecture experience

Apply Now

Similar Jobs

🔥 31 minutes ago

Navitus Health Solutions

1001 - 5000

💼 Consulting

🛡️ Insurance

🏥 Healthcare

Enterprise AI Engineer building governed generative AI, machine learning, and agentic workflows for Navitus, a pharmacy benefit manager. Developing cloud integrations, reusable AI tooling, and MLOps/LLMOps infrastructure.

🔥 32 minutes ago

Cotiviti

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior AI Engineer developing clinical NLP, agentic AI, and golden datasets for Cotiviti’s healthcare technology platform. Building explainable coding solutions and evaluation workflows remotely in the United States.

🔥 1 hour ago

Rapid7

1001 - 5000

🔒 Cybersecurity

Lead Product Manager defining Rapid7’s autonomous AI-SOC investigation platform strategy. Owning investigation content, threat intelligence workflows, quality metrics, and MDR triage automation.

🔥 1 hour ago

HiddenLayer

51 - 200

🤖 Artificial Intelligence

🔒 Cybersecurity

💳 Fintech

Lead engineer owning AI attack simulation architecture and execution at HiddenLayer, which protects technologies from adversarial AI attacks. Building agentic security-testing systems and guiding product-focused engineering teams.

🔥 1 hour ago

Navitus Health Solutions

1001 - 5000

💼 Consulting

🛡️ Insurance

🏥 Healthcare

Enterprise AI architect advancing secure, governed AI capabilities for Navitus, a pharmacy benefit manager. Defining AI platforms, agentic workflows, integrations, and operational standards across the organization.