Datacenter Infrastructure Specialist

🔥 15 hours ago

🇺🇸 United States – Remote

💵 €105.3k - €140.4k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Validate new hardware and ensure partner deployments meet Runpod specifications for distributed AI/ML workloads • Monitor fleet health and identify performance degradation • Audit downtime and provide technical data needed to protect customer SLAs • Use LLMs and AI agents to automate network triage and generate dynamic fleet runbooks • Coordinate technical incident communications and translate outages into actionable resolutions • Support the growth of infrastructure partners • Own the technical lifecycle and operational health of Runpod’s high-density GPU fleet • Serve as infrastructure advisor, technical translator, adopter, and incident commander for hardware partners • Bridge hardware partners and internal engineering teams

🎯 Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering • Strong proficiency in standard datacenter networking and performance troubleshooting • Exposure to RDMA, InfiniBand, or RoCE highly preferred • Hands-on experience with the NVIDIA Software Stack, including driver installation and performance utilities • Understanding of multi-node performance tuning • Solid Linux system administration skills • Experience with containerization using Docker • System-level troubleshooting and performance tuning at kernel and hardware interface layers • Clear written and verbal communication skills • Ability to explain hardware or networking issues to technical partners and internal leadership • Willingness to participate in a future on-call rotation • Experience in a fast-paced startup environment preferred • Experience managing or optimizing bare-metal HPC environments at massive scale preferred • Experience with Grafana, Prometheus, or Datadog preferred • Proficiency in Python, Go (Golang), or Bash preferred • Must be eligible to work in the United States; Runpod is currently unable to sponsor employment visas

🏖️ Benefits

• Meaningful equity; everyone on the team receives stock options • Generous medical, dental & vision plans; 100% coverage for employees and partial coverage for dependents • Flexible PTO • Remote-first work with inclusive, collaborative teams • $1,200 Home Office & Equipment Stipend • Culture, learning, and ownership opportunities

Apply Now

Similar Jobs

🔥 16 hours ago

WorkOS

51 - 200

🔌 API

🏢 Enterprise

🤝 B2B

Infrastructure Security Engineer building workload identity, authentication, and authorization infrastructure. WorkOS provides developer tools and APIs for enterprise-ready software.

🔥 18 hours ago

SYNCREON

10,000+ employees

🚘 Automotive

📦 Logistics

🚗 Transport

Senior Mainframe Systems Programmer supporting IBM z/OS infrastructure for a recruitment and staffing services company. Managing mainframe software, systems architecture, and operational tools.

🕒 Yesterday

hud (YC W25)

1 - 10

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Senior Infrastructure Engineer scaling HUD’s RL training-data and evaluation infrastructure. Improving AWS reliability, backend performance, CI/CD, and developer experience for frontier AI labs.

🕒 Yesterday

Magpie Literacy

11 - 50

📚 Education

👥 B2C

Senior Software Engineer building Magpie’s backend, frontend, and AWS infrastructure. Scaling a nonprofit platform that helps children become confident readers.

🕒 Yesterday

Hanwha Energy USA

201 - 500

⚡ Energy

🤝 B2B

🏠 Real Estate

Cloud & AI Infrastructure Engineer building secure, governed Azure and AI platforms. Supporting production applications, automation, observability, and cost optimization for Hanwha’s energy business.