Director of Infrastructure Engineering

πŸ”₯ 1 hour ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

πŸ€– Artificial Intelligence

☁️ SaaS

🀝 B2B

πŸ’° $20M Seed on 2024-06

Artificial Intelligence β€’ SaaS β€’ B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

πŸ“‹ Description

β€’ Lead and scale Runpod’s core cloud and bare-metal environments β€’ Establish rigorous Site Reliability Engineering practices β€’ Oversee the design, scaling, and operation of global network backbone β€’ Direct architecture and performance tuning of distributed storage systems β€’ Hire, mentor, and grow engineering managers and senior ICs β€’ Partner with product teams to forecast capacity requirements β€’ Drive measurable improvements in infrastructure reliability β€’ Provide architectural oversight for provisioning and network fabrics β€’ Ensure infrastructure primitives are robust and available.

🎯 Requirements

β€’ 7+ years leading software, infrastructure, SRE, or networking teams β€’ 8+ years building and operating large-scale distributed systems β€’ Strong architectural understanding of ultra-low latency networking β€’ Experience building high-performance distributed storage systems β€’ Strong foundation in reliability engineering and infrastructure-as-code β€’ Experience building culture across distributed technical teams β€’ Clear written and verbal communication skills β€’ Successful completion of a background check.

πŸ–οΈ Benefits

β€’ Meaningful equity in a fast-growing company β€’ Generous medical, dental & vision plans. β€’ Flexible PTO- take the time you need to recharge. β€’ $1,200 Home Office & Equipment Stipend

Apply Now

Similar Jobs

πŸ”₯ 2 hours ago

Sequen

11 - 50

☁️ SaaS

πŸ€– Artificial Intelligence

🀝 B2B

Core Systems Engineer designing and optimizing Sequen’s model serving infrastructure with Rust. Responsible for low-latency APIs and ensuring system reliability at enterprise scale.

Cloud

Microservices

Rust

πŸ”₯ 2 hours ago

Danaher

10,000+ employees

🧬 Biotechnology

πŸ₯ Healthcare

πŸ”¬ Science

Staff Platform Infrastructure Engineer responsible for designing, deploying, and maintaining cloud infrastructure. Joining a leading science and technology company focused on innovation and impact.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

Go

πŸ”₯ 10 hours ago

Northrop Grumman

10,000+ employees

🏭 Manufacturing

πŸ“¦ Logistics

πŸŽ–οΈ Defense

Principal Software Engineer leading a team in System Infrastructure Engineering for Northrop Grumman. Handling server administration, automation, and cybersecurity compliance in a remote position with required travel.

Ansible

Cyber Security

Linux

Python

VMware

πŸ”₯ 13 hours ago

Empower Pharmacy

1001 - 5000

πŸ’Š Pharmaceuticals

πŸ₯ Healthcare

🏭 Manufacturing

Principal Cloud Infrastructure Architect leading cloud strategy and AI integration for Empower Pharmacy. Architecting cloud-native solutions while building a high-performing engineering team.

Cloud

πŸ”₯ 14 hours ago

HubSpot

1001 - 5000

🀝 B2B

☁️ SaaS

πŸ“£ Marketing

VP, Infrastructure Engineering leading HubSpot's cloud platforms and AI transformation in engineering. Enabling teams to build and deploy at scale for enhanced productivity and customer experience.

Cloud

Distributed Systems