Datacenter Infrastructure Specialist

🕒 July 29

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads. • Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs. • We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet. • Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions. • Support the growth of our infrastructure partners.

🎯 Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering. • Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred. • Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning. • Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers. • Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership. • As our global fleet scales, this role may require participating in an on-call rotation in the future. • You are detail-oriented and proactive when it comes to identifying potential failures before they impact customers. • Experience working in a fast-paced environment where you have contributed to building operational workflows. • Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale. • Experience with Grafana, Prometheus, or Datadog to monitor system health. • Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.

🏖️ Benefits

• Meaningful equity in a fast-growing AI infra company — everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans — we cover 100% for all employees and partial for dependents. • Flexible PTO — take the time you need to recharge. • Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication. • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. • $1,200 Home Office & Equipment Stipend — We set you up for success from day one with gear and support to create your ideal workspace.

Apply Now

Similar Jobs

🕒 July 28

RefinedScience

11 - 50

🏥 Healthcare

💼 Consulting

📦 Logistics

Cloud Infrastructure Engineer responsible for cloud infrastructure and automation on Google Cloud Platform. Collaborating with teams to deliver scalable and secure infrastructure in a regulated healthcare environment.

🇺🇸 United States – Remote

💵 $110k - $130k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

🕒 July 28

Harrington Process Solutions

501 - 1000

🤝 B2B

🏭 Manufacturing

🍽️ Food & Beverage

Hands-on Infrastructure Engineer supporting IT infrastructure across on-premise and cloud environments. Key role in managing operations, leading initiatives, and collaborating across teams.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

🕒 July 28

Centralize

2 - 10

☁️ SaaS

🤝 B2B

🏢 Enterprise

Infrastructure Engineer focusing on scalability for Centralize's backend systems and integration pipelines handling millions of events daily. Working with Postgres, AWS, and Redis in a fast-growing startup environment.

🇺🇸 United States – Remote

💵 $190k - $260k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

🕒 July 28

The Voleon Group

201 - 500

💸 Finance

🤖 Artificial Intelligence

Senior Linux Infrastructure Engineer at Voleon, building scalable and secure Linux infrastructure for finance. Emphasizing modern tooling and robust security practices.

🇺🇸 United States – Remote

💵 $215k - $250k / year

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

🕒 July 28

Material Group

2 - 10

🎯 Recruiter

🚀 Aerospace

⚡ Energy

Infrastructure Engineer responsible for ownership and operation of cloud accelerator platform. Building control plane, managing provisioning, and ensuring reliability across compute fleet.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer