Datacenter Infrastructure Specialist

Job not on LinkedIn

🔥 21 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads. • Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs. • We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet. • Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions. • Support the growth of our infrastructure partners.

🎯 Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering. • Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred. • Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning. • Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers. • Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership. • As our global fleet scales, this role may require participating in an on-call rotation in the future. • You are detail-oriented and proactive when it comes to identifying potential failures before they impact customers. • Experience working in a fast-paced environment where you have contributed to building operational workflows. • Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale. • Experience with Grafana, Prometheus, or Datadog to monitor system health. • Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.

🏖️ Benefits

• Meaningful equity in a fast-growing AI infra company — everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans — we cover 100% for all employees and partial for dependents. • Flexible PTO — take the time you need to recharge. • Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication. • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. • $1,200 Home Office & Equipment Stipend — We set you up for success from day one with gear and support to create your ideal workspace.

Apply Now

Similar Jobs

🔥 1 hour ago

iT1

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Technical Architect for Azure Infrastructure at iT1, focusing on complex projects using Microsoft technologies. Responsible for design, implementation, and delivery of cloud-based solutions for enterprise clients.

🔥 5 hours ago

AdvanSix

1001 - 5000

🏭 Manufacturing

🤝 B2B

IT/OT Infrastructure & Cyber Architect defining and governing secure, scalable architectures. Collaborating across IT, OT, and cybersecurity to support digital transformation in manufacturing environments.

🇺🇸 United States – Remote

💵 $130.7k - $196.1k / year

💰 $12M Grant - AdvanSix on 2024-09

⏰ Full Time

🟠 Senior

🔴 Lead

👷 Infrastructure Engineer

🔥 5 hours ago

Nava

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Infrastructure Engineer building scalable infrastructure for government services. Leading technical projects for cloud-based applications that millions rely on.

🔥 20 hours ago

WashU IT

501 - 1000

📚 Education

🏢 Enterprise

☁️ SaaS

Systems Engineer II at WashU managing technology solutions like IPv4/IPV6 Address Management and DHCP. Responsible for capacity planning, documentation, and technical validation of services.