
51 - 200 employees
Founded 2022
đ¤ Artificial Intelligence
âď¸ SaaS
đ¤ B2B
đ° $20M Seed on 2024-06
Artificial Intelligence ⢠SaaS ⢠B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
đĽ 2 minutes ago
đşđ¸ United States â Remote
đľ $150k - $220k / year
â° Full Time
đ Senior
đ¤ Machine Learning Engineer
đť Ghost score 7%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2022
đ¤ Artificial Intelligence
âď¸ SaaS
đ¤ B2B
đ° $20M Seed on 2024-06
Artificial Intelligence ⢠SaaS ⢠B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
⢠Define inference performance measurements, including throughput, time to first token, inter-token latency, and cost per token ⢠Build tooling that makes performance measurements rigorous and repeatable ⢠Profile and diagnose performance problems across the serving stack, from scheduling and memory management to kernels and interconnect ⢠Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments ⢠Turn findings into production-ready runtimes, configurations, and defaults ⢠Collaborate with product and infrastructure teams to shape how inference is offered on Runpod ⢠Monitor the fast-moving inference ecosystem and evaluate what to adopt, build, or contribute back ⢠Trace serving engine/runtime bottlenecks and implement fixes when configuration tuning is insufficient ⢠Own LLM serving performance end to end across models, hardware generations, and workloads
⢠5+ years of professional system engineering experience ⢠Deep, hands-on experience with vLLM, SGLang, or a comparable serving engine in production or at serious benchmark scale ⢠Strong software engineering skills in Python ⢠Comfortable working in large, performance-critical codebases ⢠Solid understanding of LLM inference performance, including batching, memory, parallelism, and latency-throughput trade-offs ⢠Experience with inference optimization techniques such as quantization, speculative decoding, or distributed serving ⢠Rigor in benchmarking and performance analysis ⢠Comfort with GPU profiling tools ⢠Ability to explain results clearly in writing and turn them into decisions ⢠Eligible to work in the United States ⢠Must not require employment visa sponsorship ⢠Preferred: experience writing or tuning GPU kernels in CUDA or Triton ⢠Preferred: contributions to inference or ML systems projects ⢠Preferred: experience with multi-node GPU systems and high-speed networking ⢠Preferred: experience at a company where inference cost and latency were core business metrics
⢠Meaningful equity in a fast-growing company; everyone on the team receives stock options ⢠Generous medical, dental & vision plans ⢠Flexible PTO ⢠Remote work-first arrangement ⢠$1,200 Home Office & Equipment Stipend ⢠Passionate team on the cutting edge of AI infrastructure, with culture, learning, and ownership at the heart of how the company scales
Apply Nowđ Yesterday
Senior AI/ML Engineer building production GenAI and ML systems for Slateâs affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.
đşđ¸ United States â Remote
đľ $141.8k - $212.7k / year
â° Full Time
đ Senior
đ¤ Machine Learning Engineer
đŚ H1B Visa Sponsor
đ Yesterday
Machine learning engineer operationalizing scalable production models for MNTNâs Connected TV advertising platform. Building reliable ML systems that optimize campaigns for brands.
đ 2 days ago
AI/ML Engineer building production GenAI, computer vision, and data systems for Slateâs affordable customizable vehicles. Applying AI across manufacturing, supply chain, and physical operations.
đşđ¸ United States â Remote
đľ $123.3k - $185k / year
â° Full Time
đĄ Mid-level
đ Senior
đ¤ Machine Learning Engineer
đŚ H1B Visa Sponsor
đ 3 days ago
Machine Learning Engineer extending ZoomInfoâs B2B data graph and AI agents. Deploying machine learning, language models, and entity resolution for company intelligence.
đşđ¸ United States â Remote
đľ $128.1k - $201.3k / year
đ° Private Equity Round on 2014-04
â° Full Time
đĄ Mid-level
đ Senior
đ¤ Machine Learning Engineer
đŚ H1B Visa Sponsor
đ 3 days ago
Senior ML Engineer developing camera-based perception models for Torcâs autonomous trucking software. Building production-ready vision systems for object detection, segmentation, depth, and scene understanding.
đşđ¸ United States â Remote
đľ $177.3k - $212.8k / year
â° Full Time
đ Senior
đ¤ Machine Learning Engineer
đŚ H1B Visa Sponsor