Senior Machine Learning Engineer, ML Infrastructure

🔥 1 minute ago

☕ Washington – Remote

infoinfo

💵 $165.6k - $273.4k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Unity

Unity

5001 - 10000 employees

Founded 2005

🏭 Manufacturing

💼 Consulting

📣 Marketing

Manufacturing • Consulting • Marketing

Unity is a leading platform providing tools and services for creators to build games, apps, and immersive experiences. With high-quality graphics, multi-platform support, and AI enhancements, Unity empowers developers in video gaming and other industries like manufacturing. Unity offers services such as Unity Engine, Unity Cloud, and Unity Grow, which include real-time 3D creation, user acquisition, and game monetization. The company also provides extensive resources, community support, and professional training for creators at all levels. Through its success plans and developer resources, Unity aids in solving complex problems and bringing ideas to life for gaming and enterprise professionals.

📋 Description

• Design and operate large-scale online inference infrastructure serving production ML models with low latency and high reliability • Develop infrastructure supporting distributed training workflows using PyTorch, Ray Data, and Ray Train • Integrate ML pipelines with workflow orchestration systems such as Flyte or Airflow • Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning • Improve ML systems observability through latency, throughput, error-rate, cost, saturation, and model-health monitoring • Partner with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency • Improve reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation • Lead architectural improvements to make the online ML platform more robust, user-friendly, scalable, and cost-efficient

🎯 Requirements

• Experience building and operating production-grade online ML inference systems • Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems • Experience optimizing inference workloads using dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability • Strong programming skills in Python, with practical experience working on production ML systems and high-scale services • Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management • Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback • Strong systems thinking and ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs • Proven ability to lead technical direction and influence architectural decisions across teams without formal authority • Sufficient knowledge of English for professional verbal and written exchanges

🏖️ Benefits

• Equity awards • Participation in company incentive plans, such as annual discretionary bonuses or sales commissions • Comprehensive health, life, and disability insurance • Commute subsidy • Employee stock ownership • Competitive retirement/pension plans • Generous vacation and personal days • Leave and family-care programs for new parents • Office food snacks • Mental Health and Wellbeing programs and support • Employee Resource Groups • Global Employee Assistance Program • Training and development programs • Volunteering and donation matching program

Apply Now

Similar Jobs

🔥 15 hours ago

iRhythm Technologies, Inc.

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Machine Learning Scientist developing AI and signal-processing algorithms for iRhythm’s wearable cardiac diagnostics. Translating ECG research into scalable medical-device solutions and clinical insights.

🔥 18 hours ago

NBCUniversal

10,000+ employees

📱 Media

Deep Learning Engineer developing computer vision and graphics algorithms for NBCUniversal's film, television, streaming, and theme park content. Deploying models on large-scale geospatial datasets to generate 3D customer content.

🔥 20 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Machine learning engineer developing LiDAR, camera perception, and sensor-fusion models for NVIDIA DRIVE AV. Productizing efficient C++ and Python ML systems for autonomous vehicles.

🔥 21 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Deep learning engineer building and deploying LLM, VLM, and VLA systems for NVIDIA autonomous vehicles. Integrating production-grade AI with vehicle firmware and safety-critical systems.

🔥 23 hours ago

Torc Robotics

501 - 1000

🚘 Automotive

📦 Logistics

🚗 Transport

Senior ML Engineer developing ML tracking models and sensor-fusion software for Torc’s autonomous trucks. Building production systems with PyTorch, C++, Python, and vehicle sensor data.