Search Remote Jobs

Senior ML Systems Engineer, Inference

Job not on LinkedIn

🔥 2 minutes ago

🇺🇸 United States – Remote

💵 $150k - $220k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token • Build tooling that makes performance measurements rigorous and repeatable • Profile and diagnose performance problems across the serving stack, from scheduling and memory management to kernels and interconnect • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments • Turn findings into production-ready runtimes, configurations, and defaults • Collaborate with product and infrastructure teams to shape how inference is offered on Runpod • Monitor the fast-moving inference ecosystem and evaluate what to adopt, build, or contribute back • Trace serving engine/runtime bottlenecks and implement fixes when configuration tuning is insufficient • Own LLM serving performance end to end across models, hardware generations, and workloads

🎯 Requirements

• 5+ years of professional system engineering experience • Deep, hands-on experience with vLLM, SGLang, or a comparable serving engine in production or at serious benchmark scale • Strong software engineering skills in Python • Comfortable working in large, performance-critical codebases • Solid understanding of LLM inference performance, including batching, memory, parallelism, and latency-throughput trade-offs • Experience with inference optimization techniques such as quantization, speculative decoding, or distributed serving • Rigor in benchmarking and performance analysis • Comfort with GPU profiling tools • Ability to explain results clearly in writing and turn them into decisions • Eligible to work in the United States • Must not require employment visa sponsorship • Preferred: experience writing or tuning GPU kernels in CUDA or Triton • Preferred: contributions to inference or ML systems projects • Preferred: experience with multi-node GPU systems and high-speed networking • Preferred: experience at a company where inference cost and latency were core business metrics

🏖️ Benefits

• Meaningful equity in a fast-growing company; everyone on the team receives stock options • Generous medical, dental & vision plans • Flexible PTO • Remote work-first arrangement • $1,200 Home Office & Equipment Stipend • Passionate team on the cutting edge of AI infrastructure, with culture, learning, and ownership at the heart of how the company scales

Apply Now

Similar Jobs

🕒 Yesterday

Slate Auto

201 - 500

🚘 Automotive

🏭 Manufacturing

🚗 Transport

Senior AI/ML Engineer building production GenAI and ML systems for Slate’s affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.

🕒 Yesterday

MNTN

201 - 500

📣 Marketing

🤝 B2B

🛍️ eCommerce

Machine learning engineer operationalizing scalable production models for MNTN’s Connected TV advertising platform. Building reliable ML systems that optimize campaigns for brands.

🕒 2 days ago

Slate Auto

201 - 500

🚘 Automotive

🏭 Manufacturing

🚗 Transport

AI/ML Engineer building production GenAI, computer vision, and data systems for Slate’s affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.

🕒 3 days ago

ZoomInfo

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

Machine Learning Engineer extending ZoomInfo’s B2B data graph and AI agents. Deploying machine learning, language models, and entity resolution for company intelligence.

🕒 3 days ago

Torc Robotics

501 - 1000

🚘 Automotive

📦 Logistics

🚗 Transport

Senior ML Engineer developing camera-based perception models for Torc’s autonomous trucking software. Building production-ready vision systems for object detection, segmentation, depth, and scene understanding.