Senior ML Systems Engineer, Inference

Job not on LinkedIn

🔥 4 hours ago

🇺🇸 United States – Remote

💵 $150k - $220k / year

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token • Build tooling that makes performance measurements rigorous and repeatable • Profile and diagnose performance problems across the serving stack, from scheduling and memory management to kernels and interconnect • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments • Turn findings into production-ready runtimes, configurations, and defaults • Collaborate with product and infrastructure teams to shape how inference is offered on Runpod • Monitor the fast-moving inference ecosystem and evaluate what to adopt, build, or contribute back • Trace serving engine/runtime bottlenecks and implement fixes when configuration tuning is insufficient • Own LLM serving performance end to end across models, hardware generations, and workloads

🎯 Requirements

• 5+ years of professional system engineering experience • Deep, hands-on experience with vLLM, SGLang, or a comparable serving engine in production or at serious benchmark scale • Strong software engineering skills in Python • Comfortable working in large, performance-critical codebases • Solid understanding of LLM inference performance, including batching, memory, parallelism, and latency-throughput trade-offs • Experience with inference optimization techniques such as quantization, speculative decoding, or distributed serving • Rigor in benchmarking and performance analysis • Comfort with GPU profiling tools • Ability to explain results clearly in writing and turn them into decisions • Eligible to work in the United States • Must not require employment visa sponsorship • Preferred: experience writing or tuning GPU kernels in CUDA or Triton • Preferred: contributions to inference or ML systems projects • Preferred: experience with multi-node GPU systems and high-speed networking • Preferred: experience at a company where inference cost and latency were core business metrics

🏖️ Benefits

• Meaningful equity in a fast-growing company; everyone on the team receives stock options • Generous medical, dental & vision plans • Flexible PTO • Remote work-first arrangement • $1,200 Home Office & Equipment Stipend • Passionate team on the cutting edge of AI infrastructure, with culture, learning, and ownership at the heart of how the company scales

Apply Now

Similar Jobs

🕒 Yesterday

Slate Auto

201 - 500

🚘 Automotive

🏭 Manufacturing

🚗 Transport

Senior AI/ML Engineer building production GenAI and ML systems for Slate’s affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.

AWS

Azure

Google Cloud Platform

Pandas

Python

PyTorch

Scikit-Learn

🕒 Yesterday

MNTN

201 - 500

📣 Marketing

🤝 B2B

🛍️ eCommerce

Machine learning engineer operationalizing scalable production models for MNTN’s Connected TV advertising platform. Building reliable ML systems that optimize campaigns for brands.

Airflow

BigQuery

Google Cloud Platform

Python

PyTorch

SQL

🕒 2 days ago

Slate Auto

201 - 500

🚘 Automotive

🏭 Manufacturing

🚗 Transport

AI/ML Engineer building production GenAI, computer vision, and data systems for Slate’s affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.

AWS

Azure

Google Cloud Platform

Pandas

Python

PyTorch

Scikit-Learn

🕒 4 days ago

ZoomInfo

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

Machine Learning Engineer extending ZoomInfo’s B2B data graph and AI agents. Deploying machine learning, language models, and entity resolution for company intelligence.

Python

PyTorch

SQL

🕒 4 days ago

Torc Robotics

501 - 1000

🚘 Automotive

📦 Logistics

🚗 Transport

Senior ML Engineer developing camera-based perception models for Torc’s autonomous trucking software. Building production-ready vision systems for object detection, segmentation, depth, and scene understanding.

Python

PyTorch

Ray