Embedded AI Engineer, On-Device Models

🕒 July 7

🇺🇸 United States – Remote

💵 $219.3k - $274.1k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 AI Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 24%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Deepgram

Deepgram

51 - 200 employees

Founded 2015

💼 Consulting

🏥 Healthcare

📦 Logistics

💰 $47M Series B on 2022-11

Consulting • Healthcare • Logistics

Deepgram is a leading voice AI company that provides powerful APIs for speech-to-text, text-to-speech, and language understanding applications. Their platform enables developers to build sophisticated voice AI solutions for use cases such as contact centers, medical transcription, conversational AI, and more. Known for unmatched accuracy, speed, and cost-effectiveness, Deepgram's technology is trusted by top enterprises and startups worldwide. By offering real-time and highly accurate transcription capabilities, Deepgram helps businesses gain insights from voice data, making it an essential tool for transforming voice interactions.

📋 Description

• Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators. • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets. • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr. • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity. • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity. • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship. • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms. • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

🎯 Requirements

• Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices. • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments. • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation. • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains. • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance. • Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development. • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

🏖️ Benefits

• Offers Equity • Offers Bonus • 10% Annual Bonus

Apply Now

Similar Jobs

🕒 July 7

Danaher Corporation

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior AI Engineer responsible for developing AI-powered solutions for legal workflows. Collaborating with legal stakeholders and employing advanced AI techniques for contract and compliance tasks.

🕒 July 1

Quilter

1 - 10

💼 Consulting

🏭 Manufacturing

🔧 Hardware

Senior/Staff Backend Engineer at Quilter designing integrations with leading CAD tools. Shape technical direction and mentor team members in an innovative startup environment.

🕒 June 30

Crogl, Inc.

11 - 50

🔒 Cybersecurity

🤖 Artificial Intelligence

☁️ SaaS

AI Engineer developing LLM-powered features and workflows for security automation at Crogl. Collaborating with teams to build and evaluate AI systems for security investigations.

🕒 June 29

Otsuka Pharmaceutical Companies (U.S.)

1001 - 5000

💊 Pharmaceuticals

🏥 Healthcare

🧬 Biotechnology

Senior Director overseeing data engineering and AI initiatives to drive R&D strategy in pharmaceuticals. Leading a high-performing team, delivering advanced AI solutions and ensuring regulatory compliance.

🕒 June 27

Netflix

10,000+ employees

📱 Media

👥 B2C

Staff AI Engineer at Netflix designing AI infrastructure and workflows for innovative ad delivery. Engaging in a greenfield opportunity to create impactful AI systems from scratch.