Senior Machine Learning Engineer – Inference Platform

🕒 vor 2 Monaten

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Wizard

Wizard

51 - 200 Mitarbeiter

🤖 Künstliche Intelligenz

🛍️ eCommerce

🛒 Einzelhandel

💰 €50.000.000 Series A im 2021-10

Artificial Intelligence • eCommerce • Retail

Wizard AI ist ein Unternehmen, das ein maßgeschneidertes Einkaufserlebnis durch künstliche Intelligenz bietet. Es bietet einen einzigartigen textbasierten Service, der Produktempfehlungen aus dem Internet kuratiert und dabei KI-Modelle nutzt, die Benutzerpräferenzen verstehen, um Bedürfnisse vorherzusagen und personalisierte Vorschläge zu machen. Wizard AI vereinfacht das Einkaufen, indem es Produkte findet, bestellt und verfolgt sowie Rückgaben über SMS-Kommunikation verwaltet. Dieser Service ist darauf ausgelegt, Zeit zu sparen und die Annehmlichkeit des Online-Shoppings zu verbessern, indem Produktdaten, Kundenbewertungen und andere digitale Inhalte effizient jongliert werden, um fundierte Empfehlungen für Kunden zu geben.

Beschreibung

• Own and evolve our multi-engine inference platform, supporting a variety of model types and serving requirements. • Build and improve production ML pipelines — taking models from experimentation to reliable, high-throughput serving. • Define and implement model versioning, rollout, rollback, and lifecycle management strategies that ensure reproducibility and operational reliability. • Define and enforce serving-layer SLAs, including latency, availability, GPU utilization, Time-to-First-Token (TTFT), and Inter-Token Latency (ITL). • Build observability, monitoring, alerting, and operational tooling for production inference systems. • Apply software engineering best practices, including testing, CI/CD integration, and reproducibility across ML workflows. • Optimize inference performance through efficient resource utilization, hardware-aware serving strategies, and cost-conscious infrastructure design. • Ensure ML serving systems are secure, scalable, and operationally resilient. • Partner with ML, Data, Product, and DevOps teams to turn ideas into production systems, driving the technical decisions on serving and scale.

🎯 Anforderungen

• Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience. • 5–8+ years of experience in Software Engineering, ML Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of production ML serving systems. • Hands-on experience running an LLM serving engine (vLLM, TGI, TensorRT-LLM, or SGLang) in production under real load — not just managed or hosted endpoints. • Strong Python skills and software engineering fundamentals, combined with deep systems and infrastructure knowledge. • Experience with cloud platforms such as AWS, GCP, or Azure, and familiarity with ML lifecycle tooling, experimentation platforms, and model registries. • Strong grasp of inference performance — continuous batching, KV-cache and GPU-memory behavior, quantization, and CPU-versus-GPU bottlenecks — with the instinct to profile before tuning. • Experience serving heterogeneous workloads, including LLMs, embedding models, and extraction models, each with distinct latency, throughput, and scaling requirements. • Demonstrated ability to balance latency, throughput, reliability, and infrastructure cost while operating production-scale ML systems. • Experience in high-growth startup environments and comfort operating in fast-moving, evolving technical landscapes.

🏖️ Vorteile

• Health insurance • Flexible work arrangements • Professional development opportunities

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 2 Monaten

Hugging Face

51 - 200

🤖 Künstliche Intelligenz

☁️ SaaS

🏢 Unternehmen

Senior Python Software Engineer developing and maintaining Gradio and Trackio for Hugging Face. Collaborating with the ML community to enhance developer tools for machine learning.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 2 Monaten

Multi Media, LLC

51 - 200

💼 Beratung

📣 Marketing

📱 Medien

Machine Learning Engineer at Multi Media, LLC enhancing AI and ML systems for global-scale consumer platforms. Collaborating with cross-functional teams to improve recommendations and user engagement.

🇺🇸 Vereinigte Staaten – Remote

💵 $180.000 - $200.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

Numpy

Pandas

Python

PyTorch

Scikit-Learn

Tensorflow

🕒 vor 2 Monaten

Hugging Face

51 - 200

🤖 Künstliche Intelligenz

☁️ SaaS

🏢 Unternehmen

Open-Source Machine Learning Engineer improving open-source ML ecosystem at Hugging Face. Collaborating with users and contributors on libraries like Transformers and Pytorch.

🗣️🇺🇸🇬🇧 Englisch erforderlich

Python

PyTorch

Tensorflow

🕒 vor 2 Monaten

BeOne Medicines

10.000+ Mitarbeiter

🧬 Biotechnologie

🏥 Gesundheitswesen

💊 Pharmazie

Associate Director leading AI and machine learning initiatives to improve oncology research productivity at BeOne. Partnering with global teams to implement advanced AI solutions while driving R&D digital transformation.

🇺🇸 Vereinigte Staaten – Remote

💵 $158.400 - $208.400 / Jahr

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Azure

Cloud

Python

PyTorch

Scikit-Learn

Tensorflow

🕒 vor 2 Monaten

Root Inc.

1001 - 5000

🚘 Automobilindustrie

💼 Beratung

🛡️ Versicherung

Lead Machine Learning Engineer at Root enhancing customer lifetime value modeling systems using machine learning. Collaborating with data scientists and business teams to optimize ML systems for insurance solutions.

🇺🇸 Vereinigte Staaten – Remote

💵 $164.000 - $205.000 / Jahr

💰 €200.000.000 Post IPO debt im 2024-11

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich