Senior Machine Learning Engineer, ML Infrastructure – Online

Job not on LinkedIn

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of UNITYTECH CONSULTING

UNITYTECH CONSULTING

11 - 50 employees

Founded 2022

💼 Consulting

Consulting

UNITYTECH CONSULTING is a technology-focused consulting firm (inferred from the name) that appears to provide professional consulting services. There is no company-specific information in the supplied data (the provided text contains unrelated website demo entries for "Outdoor Adventure" themes), so this description is conservative and based primarily on the company name. No details about offerings, customers, or size were provided.

📋 Description

• Design and operate large-scale online inference infrastructure serving production ML models with low latency and high reliability • Develop infrastructure supporting distributed training workflows • Integrate ML pipelines with workflow orchestration systems for reliable multi-stage training workflows • Optimize model performance through compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime tuning • Improve observability through latency, throughput, error-rate, cost, saturation, and model-health monitoring • Partner with ML engineers to accelerate model iteration while maintaining production safety, scalability, and cost efficiency • Improve reliability and reproducibility of model serving workflows, including packaging, artifact validation, compatibility testing, and deployment automation • Lead architectural improvements to make the online ML platform more robust, user-friendly, scalable, and cost-efficient

🎯 Requirements

• Experience building and operating production-grade online ML inference systems • Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems • Experience optimizing inference workloads using dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability • Strong programming skills in Python, with practical experience working on production ML systems and high-scale services • Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management • Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback • Strong systems thinking and ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs • Proven ability to lead technical direction and influence architectural decisions across teams without formal authority • Sufficient knowledge of English for professional verbal and written exchanges • Work visa/immigration sponsorship is not available for this position • Relocation support is not available for this position

🏖️ Benefits

• Comprehensive health, life, and disability insurance • Commute subsidy • Employee stock ownership • Competitive retirement/pension plans • Generous vacation and personal days • Support for new parents through leave and family-care programs • Office food snacks • Mental Health and Wellbeing programs and support • Employee Resource Groups • Global Employee Assistance Program • Training and development programs • Volunteering and donation matching program • Equity awards • Participation in company incentive plans, such as annual discretionary bonuses or sales commissions

Apply Now

Similar Jobs

🔥 5 hours ago

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior backend engineer building scalable batch infrastructure for Affirm, a buy-now-pay-later company. Designing distributed systems and workflow platforms supporting ML, product, and financial engineering.

🔥 6 hours ago

Netflix

10,000+ employees

📱 Media

👥 B2C

Full stack engineer building Netflix’s messaging and communications platforms for content production. Architecting scalable frontend and backend systems that power global entertainment workflows.

🔥 15 hours ago

Mercola

201 - 500

🛍️ eCommerce

👥 B2C

🧘 Wellness

Systems Engineer modernizing Mercola’s production infrastructure across development, cloud, networking, identity, databases, and automation. Improving reliability through CI/CD, scripting, deployments, troubleshooting, and operational documentation.

🔥 18 hours ago

SmartTech

2 - 10

🤝 B2B

🏢 Enterprise

💼 Consulting

Lead SCADA Engineer architecting supervisory control systems for mission-critical data centers and infrastructure. Setting standards, reviewing implementations, and supporting commissioning across concurrent projects.

🔥 20 hours ago

HavocAI

11 - 50

📦 Logistics

🏭 Manufacturing

🎖️ Defense

AI Infrastructure Engineer building secure LLM agents, RAG pipelines, and ML tooling. Supporting defense autonomy teams with reliable internal AI systems and workflows.

🇺🇸 United States – Remote

💵 $175k - $200k / year

💰 Seed Round on 2024-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer