Technical Staff Member – Datacenter Operations

Job not on LinkedIn

🔥 1 minute ago

🏄 California – Remote

infoinfo

💵 $150k - $300k / year

⏰ Full Time

🔴 Lead

⚙️ Operations

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Prime Intellect

Prime Intellect

1 - 10 employees

🤖 Artificial Intelligence

☁️ SaaS

Artificial Intelligence • SaaS • Cloud Computing

Prime Intellect is a company focused on democratizing AI development by providing scalable and decentralized computing resources for training models. Their platform allows users to find and share global compute resources, enabling the training of state-of-the-art models through distributed clusters. They promote the collective ownership of AI innovations, including language and scientific models. Prime Intellect also offers a range of GPU options to facilitate affordable and efficient model training. They aim to advance decentralized training research and open-source AI development on a global scale.

📋 Description

• Own the operational readiness of the physical infrastructure behind the GPU cloud • Coordinate rack deployment, cabling, inventory, and acceptance testing for new GPU capacity with datacenter partners and engineering teams • Maintain asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories • Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows • Establish maintenance plans and change procedures that minimize customer disruption and protect equipment and data • Track capacity readiness, hardware failure trends, repair times, and operational risks • Automate repetitive reporting and operational workflows • Partner with facility teams on power, cooling, environmental monitoring, and readiness for high-density GPU deployments • Create runbooks and escalation procedures • Support incident response across datacenter and infrastructure teams • Coordinate deployments, hardware maintenance, and incident response with datacenter partners • Turn new capacity into dependable production infrastructure and reduce time to repair

🎯 Requirements

• 3+ years in datacenter operations, hardware infrastructure, or production systems operations • Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling • Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents • Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools • Strong operational judgment, documentation habits, and ownership of issues through resolution • Knowledge of GPU server components, PCIe devices, memory, storage, and hardware diagnostics • Knowledge of rack power budgeting, redundant power paths, airflow, and high-density cooling fundamentals • Knowledge of fiber and copper cabling, optics, labeling, and physical network troubleshooting • Knowledge of asset tracking, spares management, change control, and incident management • Basic scripting for inventory, health checks, and operational automation • Familiarity with safe datacenter working practices • Experience with NVIDIA DGX/HGX systems or large GPU cluster deployments is a plus • Experience with liquid-cooled infrastructure and facility engineering coordination is a plus • Experience bringing up new datacenter sites or expanding multi-site capacity is a plus • Experience with hardware qualification, burn-in testing, and reliability analysis is a plus • Experience integrating physical operations with automated fleet provisioning is a plus

🏖️ Benefits

• Equity incentives • Direct work with customers and world-class engineering team • Opportunity to impact systems powering next-generation AI breakthroughs

Apply Now

Similar Jobs

🔥 2 hours ago

Thermo Fisher Scientific

10,000+ employees

🏥 Healthcare

📦 Logistics

🏭 Manufacturing

Operations Strategy Consultant leading clinical trial proposal strategy for Thermo Fisher Scientific’s clinical development programs. Managing RFPs, bid defenses, operational deliverables, budgets, and cross-functional teams.

🔥 2 hours ago

AutoNation

10,000+ employees

🛡️ Insurance

📦 Logistics

🚘 Automotive

Used-vehicle operations director improving sales, inventory, acquisition, and reconditioning across AutoNation dealerships. Coaching store leaders and driving performance through analytics and standardized retail processes.

🔥 4 hours ago

At-Bay

201 - 500

🛡️ Insurance

🔒 Cybersecurity

💳 Fintech

Director leading GTM strategy, revenue operations, and AI initiatives for At-Bay, an InsurSec company protecting small businesses from digital risks.

🇺🇸 United States – Remote

💰 $3.7M Venture Round on 2022-09

⏰ Full Time

🔴 Lead

⚙️ Operations

🔥 6 hours ago

Aligned Data Centers

501 - 1000

🏗️ Construction

💼 Consulting

📦 Logistics

Electrical operations SME optimizing Aligned Data Centers’ critical infrastructure reliability and efficiency. Driving design standards, training, incident recovery, audits, and sustainable technology evaluation.

🔥 7 hours ago

CNG Holdings, Inc.

1001 - 5000

💸 Finance

👥 B2C

District Director overseeing multi-unit retail operations, sales, compliance, and teams for CNG Holdings in Wisconsin. Managing store performance across a remote district with up to 50% travel.