Datacenter Infrastructure Specialist

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $120k - $160k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Validate new hardware and ensure partner deployments meet Runpod specifications for distributed AI/ML workloads • Monitor fleet health and identify performance degradation • Audit downtime and provide technical data to protect customer SLAs • Use LLMs and AI agents to automate network triage and generate dynamic fleet runbooks • Coordinate technical incident communications and translate outages into actionable resolutions • Support the growth of infrastructure partners • Own the technical lifecycle and operational health of Runpod’s high-density GPU fleet • Advise, translate for, onboard, and act as incident commander for hardware partners • Bridge hardware partners and internal engineering teams • Contribute to the resilience, scalability, and revenue velocity of Runpod’s global physical infrastructure

🎯 Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering • Strong proficiency in standard datacenter networking and performance troubleshooting • Exposure to RDMA, InfiniBand, or RoCE highly preferred • Hands-on experience with the NVIDIA Software Stack, including driver installation and performance utilities • Understanding of multi-node performance tuning • Solid Linux system administration skills • Experience with containerization, including Docker • System-level troubleshooting and performance tuning at kernel and hardware interface layers • Clear written and verbal communication skills • Ability to participate in a future on-call rotation • Detail-oriented and proactive approach to identifying potential failures • Preferred: startup experience, HPC bare-metal environments at massive scale, Grafana, Prometheus, Datadog, Python, Go, Bash, and internal APIs • Eligible to work in the United States; Runpod is currently unable to sponsor employment visas

🏖️ Benefits

• Meaningful equity; everyone on the team receives stock options • Medical, dental & vision plans; 100% coverage for employees and partial coverage for dependents • Flexible PTO • Remote-first work environment with inclusive, collaborative teams • $1,200 Home Office & Equipment Stipend • Culture, learning, and ownership opportunities

Apply Now

Similar Jobs

🔥 15 minutes ago

Kyndryl

10,000+ employees

💼 Consulting

📦 Logistics

🏥 Healthcare

VMware Infrastructure Specialist supporting Kyndryl’s managed enterprise technology services. Designing, integrating, automating, and troubleshooting hybrid cloud and on-premises infrastructure.

🔥 5 hours ago

The Bancorp

501 - 1000

💼 Consulting

📦 Logistics

🏦 Banking

Senior Systems Engineer operating Windows, VMware, AWS, and Azure infrastructure for The Bancorp’s fintech banking services. Strengthening identity, security, resiliency, and automation.

🔥 6 hours ago

Federal City Recovery Services - DC

51 - 200

🏥 Healthcare

🤝 Non-profit

🌍 Social Impact

Infrastructure Engineer leading Paramount’s migration from on-premises systems to AWS hybrid architecture. Owning cloud strategy, secure landing zones, PCI-DSS compliance, and workload modernization.

🔥 17 hours ago

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

ICT Infrastructure Engineer supporting SAIC’s Army Reserve enterprise network and telecommunications construction projects. Reviewing RFIs, submittals, and quality assurance inspections across U.S. locations.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $500M Post-IPO Debt - SAIC on 2025-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

🔥 22 hours ago

Bishop Fox

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Infrastructure Engineer operating Bishop Fox’s AWS, Azure, and Microsoft 365 infrastructure. Building Terraform, Go, Python, and CI/CD systems for offensive-security services.