Support Engineer, GPU Infrastructure

🔥 14 hours ago

🐊 Florida – Remote

infoinfo

💵 $95k - $130k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

📞 Support Engineer

👻 Ghost score 22%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Hydra Host

Hydra Host

11 - 50 employees

🔧 Hardware

🏢 Enterprise

🤖 Artificial Intelligence

💰 $10M Seed Round on 2022-04

Hardware • Enterprise • Artificial Intelligence

Hydra Host is a provider of high-performance computing solutions, offering dedicated bare metal GPU server access optimized for AI and HPC workloads. Their platform allows users to access and rent top-tier GPUs globally, providing unparalleled performance, security, and customization. Hydra Host's infrastructure includes a marketplace, known as Brokkr, that offers a wide array of GPU configurations and solutions tailored for mission-critical applications such as AI, big data, and machine learning. Through their robust, secure, and scalable solutions, Hydra Host ensures customers enjoy full control over their server environments, with options for scalability and future-readiness. The company's offerings are trusted by leading firms seeking efficient and innovative computing solutions.

📋 Description

• Diagnose Linux server issues end to end, including boot and network boot failures, kernel and driver problems, filesystems, storage pressure, services, memory and CPU behavior, and instability • Diagnose hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics • Isolate server-side network problems involving NICs, drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures • Troubleshoot NVIDIA GPU servers, including GPU availability, thermal throttling, driver and VBIOS mismatch, PCIe, XID errors, and host-level conditions • Separate hardware, OS, network, application, and configuration issues before escalation • Own incidents through resolution or clean handoff, set severity by blast radius, and identify related tickets • Escalate to engineering with evidence packs and participate in root cause analysis and post-incident reviews • Coordinate with data center partners on remote hands, reboots, cabling, optics checks, component replacement, and physical inspection • Open and track hardware RMAs with OEMs through replacement and validate repaired or replaced equipment • Maintain accurate asset records and support server turn-ups, migrations, and decommissions • Communicate clear cause and timeline updates to technically sophisticated customers • Write and improve runbooks, operational procedures, categories, and closure reasons • Build Python, Bash, or similar scripts and tools for health checks, data collection, and routine operations • Contribute to infrastructure-as-code, configuration management, monitoring, alerting, and support workflow improvements • Work a defined shift, participate in an escalation rotation, and provide written shift handoffs

🎯 Requirements

• Three or more years supporting production servers, data center infrastructure, or bare metal and cloud environments • Strong hands-on Linux troubleshooting, including logs, dmesg, systemd, storage tooling, and network utilities • Experience with server hardware, including CPU and memory, storage and filesystems, RAID, PCIe, NICs, power, BIOS and UEFI, firmware, and drivers • Out-of-band management experience with IPMI, Redfish, iDRAC, iLO, or similar • Working TCP/IP knowledge and ability to determine whether a problem is on the host or network • Fault-domain reasoning and production judgment • Experience with ticketing, monitoring, incident management, or infrastructure management systems • Clear written English • Helpful but not required: NVIDIA GPU servers at scale; CUDA, NCCL, NVLink, or DCGM; HPC or AI training environments; InfiniBand or high-performance Ethernet; Dell, HPE, Supermicro, or Lenovo platforms; NVMe, ZFS, Ceph, or distributed storage; Prometheus, Grafana, or similar observability tooling; NetBox or another infrastructure and asset register; Ansible, Terraform, or configuration management; Git-based infrastructure workflows; optics, transceivers, DAC or AOC cabling; geographically distributed third-party facilities

🏖️ Benefits

• Defined shift schedule agreed before starting • Escalation rotation for high-severity issues outside shift hours • Written handoff at the end of every shift • Opportunity to influence the structure of a newly built support function

Apply Now

Similar Jobs

🔥 14 hours ago

Velera

1001 - 5000

💼 Consulting

📣 Marketing

💳 Fintech

IT Support Analyst providing first-level remote support for Velera’s fintech and payments clients. Troubleshooting hardware, software, networks, and ServiceNow incidents across phone, chat, and email.

🇺🇸 United States – Remote

💵 $19 - $23 / hour

⏰ Full Time

🟢 Junior

🟡 Mid-level

📞 Support Engineer

🚫👨‍🎓 No degree required

🔥 16 hours ago

Highland Electric Fleets

51 - 200

🚘 Automotive

📦 Logistics

🚗 Transport

Field Support Technician supporting EV charging-site launches, vehicle inspections, and charger troubleshooting. Highland Electric Fleets enables affordable electric fleets for schools, municipalities, and fleet operators.

🔥 17 hours ago

Jones Lang LaSalle Americas, Inc.

10,000+ employees

🏠 Real Estate

🤝 B2B

💼 Consulting

Insider Risk Technical Analyst supporting JLL, a global real estate and investment management company. Triaging insider-threat escalations and analyzing DLP, SIEM, EDR, and other security data.

🔥 18 hours ago

Impact Advisors

501 - 1000

💼 Consulting

🏥 Healthcare

🤖 Artificial Intelligence

Oracle Fusion HCM analyst supporting healthcare consulting clients’ HR systems. Managing modules, reporting, security, releases, troubleshooting, and continuous optimization.

🔥 21 hours ago

Gravitate

51 - 200

⚡ Energy

☁️ SaaS

📦 Logistics

L3 Support Engineer owning complex SaaS customer escalations for Gravitate Energy. Troubleshooting databases, APIs, integrations, and data layers while partnering with Product and Engineering.