Lead Machine Learning Operations Engineer

🕒 vor 1 Monat

🏄 California, New York – Remote

info

💵 $157.000 - $235.000 / Jahr

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Paramount

Paramount

10.000+ Mitarbeiter

Gegründet 1912

💼 Beratung

📣 Marketing

📱 Medien

Consulting • Marketing • Media

Paramount ist ein globales Multimedia-Unterhaltungs- und Nachrichtenunternehmen, das eine Vielzahl von Dienstleistungen anbietet, darunter direkte digitale Abonnement-Video-on-Demand- und Live-Streaming-Dienste über Paramount+. Es besitzt auch Pluto TV, einen führenden kostenlosen Streaming-Fernsehdienst, MTV, die weltweit führende Jugendunterhaltungsmarke, und CBS Sports, ein führender Anbieter von Sportübertragungen im Fernsehen. Paramount Pictures ist seit 1912 ein legendärer Produzent und Vertreiber von Filmen und verfügt über eine Bibliothek mit über 1. 000 Titeln. Das Unternehmen engagiert sich stark für Inklusion und Wirkung und konzentriert sich auf Vielfalt, globale Nachhaltigkeit und Inhalte, die Veränderung bewirken. Als bedeutender Akteur im Bereich von Live- und On-Demand-Streaming-Diensten bietet Paramount eine breite Palette an Inhalten, von Sport bis zur Kinderunterhaltung, Comedy und wegweisenden Dokumentationen, und beeinflusst sowohl lineare als auch Streaming-Plattformen weltweit.

Beschreibung

• Own ML production reliability strategy • Define and lead the operational strategy for production ML systems, including monitoring, traceability, deployment safety, incident response, and post-deployment validation. • Set the standards ML teams use to assess model health, performance, and trustworthiness in production. • Own model traceability and governance • Ensure every production model has clear lineage (data, features, code, artifacts, validation, deployment history) and drive adoption of model registry and metadata tooling across ML teams. • Build end-to-end ML observability • Design and implement monitoring across the full ML signal path: data arrival, feature freshness, distribution stability, candidate generation, ranking behavior, model metrics, serving latency, and SLA performance. • Define production health metrics • Partner with ML, data, product, and business stakeholders to define post-deployment metrics covering model quality, system reliability, business guardrails, and degradation indicators. • Detect drift and degradation proactively • Detect data drift, feature drift, model behavior changes, and silent failures before they impact customers via thresholding, alerting, anomaly detection, and release-over-release monitoring. • Lead diagnostic tooling and root-cause analysis • Build dashboards, logs, and diagnostic workflows that progress quickly from “recommendations look off” to root cause, with context captured across candidates, features, scores, ranking decisions, and downstream outcomes. • Own ML deployment safety • Define and operate automated gates that prevent bad models or bad data from being promoted to production. • Partner with MLEs to establish validation checks, rollback criteria, canary strategies, shadow testing, and release health reviews. • Lead ML incident response • Own incident response practices for ML systems, including rollback playbooks, hotfix strategies, severity definitions, tradeoff frameworks, communications, and post-mortems. • Drive closure of systemic gaps after incidents rather than only resolving the immediate issue. • Partner across ML Platform, Data, and ML • Partner with DevOps/Platform on infrastructure and observability needs; with Data Engineering on data quality, drift, and freshness; and with ML Engineering to embed operational requirements into development and deployment workflows. • Set standards and mentor others • Act as the technical lead for ML operations: establish reusable patterns, playbooks, and standards, and mentor engineers on reliability, observability, and operational rigor.

🎯 Anforderungen

• 5+ years of experience in machine learning engineering, ML platform, applied ML, MLOps, data platform, reliability engineering, or a related technical role. • Demonstrated experience operating production ML systems, including monitoring, deployment, incident response, model validation, data quality, or reliability ownership. • Experience leading technical initiatives across multiple engineering teams, especially where success required influencing architecture, tooling, standards, or adoption. • Hands-on experience with model registries, feature stores, ML metadata systems, production monitoring, model deployment pipelines, or ML observability platforms. • Solid knowledge of end-to-end ML systems, including training data, features, model artifacts, offline validation, online serving, post-deployment metrics, and business outcome measurement. • Ability to reason about ML operational failure modes: stale features, distribution shift, training-serving skew, delayed labels, and offline-online metric gaps. • Solid SQL skills and comfort investigating data quality, feature distributions, model outputs, pipeline behavior, and production anomalies. • Track record of cross-functional collaboration with Platform, Data, and ML Engineering to deliver production-grade operational capabilities. • Solid written and verbal communication skills, including the ability to explain ML system health, risks, incidents, and tradeoffs to both technical and non-technical stakeholders.

🏖️ Vorteile

• medical • dental • vision • 401(k) plan • life insurance coverage • disability benefits • tuition assistance program • PTO

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Docker, Inc

51 - 200

💼 Beratung

☁️ SaaS

ML Engineer developing intelligence-driven product capabilities for Docker's platform. Collaborating with founding engineers to shape technical direction and build ML systems that enhance security and governance.

🇺🇸 Vereinigte Staaten – Remote

💵 $138.500 - $225.500 / Jahr

💰 €105.000.000 Series C im 2022-03

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Workiva

1001 - 5000

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Senior Staff Machine Learning Engineer defining how AI is architected and deployed across Workiva’s platform. Leading design and implementation of enterprise AI systems for mission-critical workflows.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Local Infusion

1 - 10

🏥 Gesundheitswesen

Machine Learning Engineer at Local Infusion building AI-driven technology to enhance specialty infusion care. Develop models for operational efficiency and patient treatment acceleration.

🇺🇸 Vereinigte Staaten – Remote

💰 €4.000.000 Seed Round im 2022-11

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Cloud

Pandas

Python

PyTorch

Scikit-Learn

SQL

Tensorflow

🕒 vor 1 Monat

GitKraken

51 - 200

☁️ SaaS

🤝 B2B

Senior Machine Learning Engineer or Applied Data Scientist at GitKraken. Developing practical AI solutions to improve developer workflows with high ownership and speed.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🤖 Machine-Learning-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Torc Robotics

501 - 1000

🚘 Automobilindustrie

📦 Logistik

🚗 Transport

Machine Learning Engineer developing AI initiatives for autonomous vehicle software. Leading data science projects while mentoring team members in a collaborative environment.

🇺🇸 Vereinigte Staaten – Remote

💵 $153.200 - $183.800 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Machine-Learning-Entwickler

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich