
1001 - 5000 Mitarbeiter
Gegründet 2019
💼 Beratung
📦 Logistik
🏭 Fertigung
💰 Grant im 2020-12
Consulting • Logistics • Manufacturing
Cerence Inc. ist ein global tätiges Unternehmen mit Fokus auf AI‑gestützte Lösungen, insbesondere für die Automobilindustrie. Das Unternehmen ist auf Conversational AI und Generative AI spezialisiert, die intelligente, natürliche und personalisierte Interaktionen zwischen Menschen und Fahrzeugen ermöglichen. Mit Innovationen wie proprietären Automotive Large Language Models (LLMs) verbessert Cerence das Nutzererlebnis über verschiedene Verkehrsmittel hinweg – von Pkw über Zweiräder bis hin zu Lkw. Weltweit sind über 500 Millionen Fahrzeuge mit Cerence Technologie ausgeliefert; das Unternehmen bedient mehr als 80 OEMs und Tier‑1‑Zulieferer. Cerence treibt die Weiterentwicklung von AI kontinuierlich voran, mit dem Ziel, das Nutzererlebnis im Fahrzeug durch schnelle Bereitstellung und nahtlose Integration seiner Lösungen zu revolutionieren.
🕒 vor 1 Monat
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

1001 - 5000 Mitarbeiter
Gegründet 2019
💼 Beratung
📦 Logistik
🏭 Fertigung
💰 Grant im 2020-12
Consulting • Logistics • Manufacturing
Cerence Inc. ist ein global tätiges Unternehmen mit Fokus auf AI‑gestützte Lösungen, insbesondere für die Automobilindustrie. Das Unternehmen ist auf Conversational AI und Generative AI spezialisiert, die intelligente, natürliche und personalisierte Interaktionen zwischen Menschen und Fahrzeugen ermöglichen. Mit Innovationen wie proprietären Automotive Large Language Models (LLMs) verbessert Cerence das Nutzererlebnis über verschiedene Verkehrsmittel hinweg – von Pkw über Zweiräder bis hin zu Lkw. Weltweit sind über 500 Millionen Fahrzeuge mit Cerence Technologie ausgeliefert; das Unternehmen bedient mehr als 80 OEMs und Tier‑1‑Zulieferer. Cerence treibt die Weiterentwicklung von AI kontinuierlich voran, mit dem Ziel, das Nutzererlebnis im Fahrzeug durch schnelle Bereitstellung und nahtlose Integration seiner Lösungen zu revolutionieren.
• Optimize and deploy high ‑ performance LLM inference pipelines • Own inference runtimes across data center, edge, and embedded platforms • Push model performance through quantization, kernel fusion, and cache optimization • Drive latency and throughput improvements that directly impact production products • Enable efficient, reliable deployment without external vendor dependency • Build deep expertise and ownership of: vLLM TensorRT‑LLM llama.cpp QAIRT • Extend and tune inference engines using custom CUDA kernels • Adapt runtimes for constrained and embedded deployment environments • Implement and evaluate quantization strategies: INT8, INT4, FP4, FP8, mixed precision AWQ GPTQ • Balance accuracy, latency, memory footprint, and throughput • Optimize key–value cache performance through: Paging Prefix caching Cache ‑ aware memory layout design • Design and tune: Batching strategies Continuous batching Speculative decoding
• Proven experience optimizing ML inference performance in production • Deep understanding of GPU architecture and memory hierarchies • Hands ‑ on experience with CUDA and low ‑ level performance tuning • Experience deploying models beyond research environments • Critical Technical Skills • Inference engines: vLLM, TensorRT ‑ LLM, llama.cpp, QAIRT • CUDA kernel development and profiling • Quantization techniques: INT8/INT4/FP4/FP8, AWQ, GPTQ • KV cache optimisation and memory layout design • Latency optimisation: batching, speculative decoding, continuous batching
• Annual bonus opportunity • Insurance coverage (medical, dental, vision, life, and disability) • Paid time off • Paid holidays • Company contribution to the RRSP (Registered Retirement Savings Plan) • Equity awards for certain positions and levels • Remote and/or hybrid work available depending on the position
Jetzt Bewerben🕒 vor 1 Monat
Senior Full Stack Developer integrating new features and working with clients in retail electronics. Requires 5+ years experience and knowledge of Microsoft tech stack.
🗣️🇮🇹 Italienisch erforderlich
🕒 vor 1 Monat
Senior Software Engineer developing scalable platform components and supporting cloud infrastructure at Robert Half. Leading design and implementation with a focus on CI/CD and platform reliability.
🇺🇸 Vereinigte Staaten – Remote
💵 $104.000 - $153.000 / Jahr
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Tech Lead for Consumer Team to drive technical direction and execution for consumer web experience at Koalafi. Leading a team of engineers in modernizing systems and delivering tools for financial needs.
🇺🇸 Vereinigte Staaten – Remote
💵 $167.723 - $217.053 / Jahr
💰 Debt Financing im 2022-08
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Lead Engineer on Product Catalog Team for Stitch Fix redefining retail with technology and data. Responsible for evolving catalog systems and improving product data quality.
🇺🇸 Vereinigte Staaten – Remote
💵 $111.800 - $186.000 / Jahr
💰 €36.900.000 Venture Round im 2017-11
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Senior Software Engineer applying software engineering and machine learning for mineral exploration at KoBold Metals. Collaborating with data scientists and geologists to shape the future of energy transition metal discovery.
🇺🇸 Vereinigte Staaten – Remote
💵 $170.000 - $215.000 / Jahr
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🗣️🇺🇸🇬🇧 Englisch erforderlich
Numpy
Python