Senior Principal Software Engineer

🕒 il y a 1 mois

🗣️🇺🇸🇬🇧 Anglais requis

C++

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Cerence Inc.

Cerence Inc.

1001 - 5000 employés

Fondée en 2019

💼 Conseil

📦 Logistique

🏭 Fabrication

💰 Grant en 2020-12

Consulting • Logistics • Manufacturing

Cerence Inc. est une entreprise mondiale spécialisée dans les solutions basées sur l'intelligence artificielle, en particulier dans l'industrie automobile. Elle se spécialise dans les technologies d'IA conversationnelle et générative qui créent des interactions intelligentes, naturelles et personnalisées entre les humains et les véhicules. Avec des innovations telles que leurs modèles de langage de grande taille propriétaires pour l'automobile, Cerence améliore les expériences utilisateur à travers diverses formes de transport, y compris les voitures, les deux-roues et les camions. L'entreprise a plus de 500 millions de véhicules équipés de sa technologie d'IA, servant plus de 80 fabricants d'équipement d'origine (OEM) et clients de rang 1 dans le monde entier. Cerence s'engage dans des avancées continues en intelligence artificielle, visant à révolutionner l'expérience utilisateur en voiture grâce à une livraison rapide et une intégration fluide de ses solutions.

Description

• Optimize and deploy high ‑ performance LLM inference pipelines • Own inference runtimes across data center, edge, and embedded platforms • Push model performance through quantization, kernel fusion, and cache optimization • Drive latency and throughput improvements that directly impact production products • Enable efficient, reliable deployment without external vendor dependency • Build deep expertise and ownership of: vLLM TensorRT‑LLM llama.cpp QAIRT • Extend and tune inference engines using custom CUDA kernels • Adapt runtimes for constrained and embedded deployment environments • Implement and evaluate quantization strategies: INT8, INT4, FP4, FP8, mixed precision AWQ GPTQ • Balance accuracy, latency, memory footprint, and throughput • Optimize key–value cache performance through: Paging Prefix caching Cache ‑ aware memory layout design • Design and tune: Batching strategies Continuous batching Speculative decoding

🎯 Exigences

• Proven experience optimizing ML inference performance in production • Deep understanding of GPU architecture and memory hierarchies • Hands ‑ on experience with CUDA and low ‑ level performance tuning • Experience deploying models beyond research environments • Critical Technical Skills • Inference engines: vLLM, TensorRT ‑ LLM, llama.cpp, QAIRT • CUDA kernel development and profiling • Quantization techniques: INT8/INT4/FP4/FP8, AWQ, GPTQ • KV cache optimisation and memory layout design • Latency optimisation: batching, speculative decoding, continuous batching

🏖️ Avantages

• Annual bonus opportunity • Insurance coverage (medical, dental, vision, life, and disability) • Paid time off • Paid holidays • Company contribution to the RRSP (Registered Retirement Savings Plan) • Equity awards for certain positions and levels • Remote and/or hybrid work available depending on the position

Postuler Maintenant

Emplois Similaires

🕒 il y a 1 mois

Albelissa Engineering, IT & Digital Solutions

51 - 200

🎯 Recrutement

💼 Conseil

Senior Full Stack Developer integrating new features and working with clients in retail electronics. Requires 5+ years experience and knowledge of Microsoft tech stack.

🇺🇸 États-Unis – Télétravail

💵 €35 000 / an

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇮🇹 Italien requis

Entity Framework

GraphQL

JavaScript

SQL

.NET

🕒 il y a 1 mois

Dyson

10 000+ employés

🔧 Matériel

🏭 Fabrication

🛒 Commerce de détail

Senior Software Engineer developing scalable platform components and supporting cloud infrastructure at Robert Half. Leading design and implementation with a focus on CI/CD and platform reliability.

🇺🇸 États-Unis – Télétravail

💵 $104 000 - $153 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

🧑‍💻 Développeur Full-Stack

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Render

11 - 50

💼 Conseil

📦 Logistique

☁️ SaaS

Software Engineer responsible for developing and operating compute infrastructure across multiple cloud providers. Focused on Kubernetes, container orchestration, and performance optimization.

🇺🇸 États-Unis – Télétravail

💵 $170 000 - $290 000 / an

⏰ Temps Plein

🟠 Senior

🔴 Expert

🧑‍💻 Développeur Full-Stack

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Stitch Fix

5001 - 10000

📣 Marketing

📦 Logistique

💼 Conseil

Lead Engineer on Product Catalog Team for Stitch Fix redefining retail with technology and data. Responsible for evolving catalog systems and improving product data quality.

🇺🇸 États-Unis – Télétravail

💵 $111 800 - $186 000 / an

💰 €36 900 000 Venture Round en 2017-11

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

KoBold Metals

51 - 200

🤖 Intelligence artificielle

🔬 Science

⚡ Énergie

Senior Software Engineer applying software engineering and machine learning for mineral exploration at KoBold Metals. Collaborating with data scientists and geologists to shape the future of energy transition metal discovery.

🇺🇸 États-Unis – Télétravail

💵 $170 000 - $215 000 / an

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇺🇸🇬🇧 Anglais requis

Numpy

Python