Senior Machine Learning Operations Engineer

🕒 July 10

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of BetMGM

BetMGM

501 - 1000 employees

Founded 2018

💼 Consulting

📣 Marketing

✈️ Travel

💰 $25.1k Seed Round on 2022-10

Consulting • Marketing • Travel

BetMGM is the exclusive sports betting partner to MGM Resorts nationwide, offering both online and physical casino gaming experiences. As a leader in online casino and poker gaming, BetMGM also collaborates with sister brands such as Borgata Online and PartyCasino. The company is committed to fostering an inclusive workplace and emphasizes responsible gambling practices, ensuring a sustainable and enjoyable gaming environment for its customers.

📋 Description

• Stand up and operate BetMGM's ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex), with Terraform-managed infrastructure. • Build self-service scaffolds that let data scientists ship a model end-to-end without a ticket queue — cookie-cutter project templates with CI, drift monitoring, alerting, IaC, and Snowflake connectivity pre-baked. • Design and operate batch scoring pipelines — SageMaker Batch Transform, dbt-orchestrated scoring against Snowflake, Snowpark ML — with explicit freshness and cost SLAs. • Design and operate real-time inference paths — SageMaker real-time endpoints, Lambda + Bedrock for GenAI, API Gateway — with stated latency budgets (typically sub-100ms) and graceful degradation under load. • Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity — training-serving skew is treated as an incident, not a tradeoff. • Build CI/CD for ML — model registry, automated retraining triggers, model versioning, lineage from feature → training run → deployed model → live prediction. • Implement champion/challenger, shadow deployments, and canary releases as platform primitives so individual model teams do not reinvent them per project. • Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor — pick one and standardize) with paging that routes to humans who can fix it. • Own MLOps incident response — production model failures are SEV events with postmortems. • Right-size endpoints, batch caching, request batching, and autoscaling. State cost-per-prediction targets up front and meet them. • Integrate LLM APIs (Bedrock, Anthropic, OpenAI) into production paths — RAG pipelines, agent eval frameworks, prompt versioning, cost and latency observability.

🎯 Requirements

• BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field — or equivalent practical experience. • 5+ years shipping software in production — Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging — including time on-call. • 3+ years operating ML in production — you have owned a model in prod that served real traffic, with stated latency and cost budgets and a runbook you wrote. • AWS depth across the SageMaker surface (Training, Endpoints, Batch Transform, Model Registry, Pipelines). • Snowflake fluency — Snowpark ML, Cortex, dbt-orchestrated batch scoring, RBAC for ML workloads. • IaC for ML — Terraform + SageMaker Pipelines or equivalent. No manual console deployments to production. • Feature store experience — SageMaker Feature Store, Tecton, or Feast — with explicit ownership of online/offline parity. • Champion/challenger, shadow, and canary deployment patterns as production muscle, not blog-post familiarity. • Drift and model monitoring — Evidently, Arize, WhyLabs, or SageMaker Model Monitor — wired to a paging path. • Software-engineering-first mindset — you treat ML systems as systems, not notebooks.

🏖️ Benefits

• Medical, Dental, Vision, Life, and Disability Insurance • 401(k) with company match • Pre-tax spending accounts including health care FSA and commuter savings • Flexible paid time off • Professional development reimbursement and ongoing skills training opportunities • Employee resource groups • Swag, ticket giveaways, and more!

Apply Now

Similar Jobs

🕒 July 10

Unity

5001 - 10000

🏭 Manufacturing

💼 Consulting

📣 Marketing

Senior Machine Learning Engineer at Unity architecting AI-powered bidding systems to maximize ad performance. Collaborating across teams to drive innovation in bidding algorithms.

🕒 July 9

Credit Acceptance

1001 - 5000

💸 Finance

💳 Fintech

🚘 Automotive

Manager of Predictive Modeling & Machine Learning leading statistical and ML model development at Credit Acceptance. Driving strategic decision-making across credit risk, collections, and business operations.

🇺🇸 United States – Remote

💵 $158k - $170k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🦅 H1B Visa Sponsor

info

🕒 July 9

hatch I.T.

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

AI/ML Engineer optimizing machine learning capabilities for Expression's edge computing platforms. Collaborating with software engineers to develop AI pipelines in constrained environments.

🇺🇸 United States – Remote

💵 $130k - $170k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🕒 July 9

Deepgram

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Applied ML Engineer developing robust models for Deepgram's leading voice AI platform. Streamlining the research-to-production pipeline while collaborating with research scientists.

🇺🇸 United States – Remote

💵 $150k - $220k / year

💰 $47M Series B on 2022-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🦅 H1B Visa Sponsor

info

🕒 July 8

Torc Robotics

501 - 1000

🚘 Automotive

📦 Logistics

🚗 Transport

Senior ML Engineer developing VLM datasets at Torc Robotics. Designing pipeline, advancing autonomous vehicle tech, and collaborating with cross-functional teams.