Senior Platform Support Engineer

Job not on LinkedIn

🔥 19 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Submer

Submer

51 - 200 employees

Founded 2015

🔧 Hardware

⚡ Energy

🤝 B2B

Hardware • Energy • B2B

Submer is a provider of connected intelligence for AI infrastructure, specializing in advanced liquid/immersion cooling and modular data-center solutions. They design, build and operate AI-ready environments (from power and land to cloud and edge), manufacture immersion cooling pods (SmartPod EVO/EXO), and offer GPUaaS/AIaaS and modular deployment services. Submer emphasizes energy and water savings, high-density thermal architectures for demanding AI workloads, and sovereign-ready, globally deployable solutions.

📋 Description

• Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure • Diagnose and resolve complex issues affecting GPU clusters, compute nodes, networking, and storage • Investigate incidents across firmware, drivers, operating systems, and platform services • Perform root cause analysis for major incidents and contribute to long-term remediation • Serve as a technical escalation point for complex or high-priority support cases • Troubleshoot GPU compute nodes, Kubernetes clusters, networking infrastructure, and storage systems • Analyze logs, telemetry, and monitoring signals to identify platform instability • Monitor and investigate security alerts; validate, triage, and escalate potential security incidents • Support distributed GPU training and inference workloads, including multi-GPU and multi-node jobs • Troubleshoot GPU workload scheduling, job queues, scheduling constraints, and resource fragmentation • Diagnose high-performance networking issues involving RDMA and RoCE • Coordinate with customer data center technicians for remote diagnostics and hardware interventions • Validate on-premise GPU infrastructure installations and deployments • Coordinate hardware replacements and RMA processes, and validate hardware health after replacements • Participate in 24/7 on-call rotations and resolve incidents according to SLAs • Improve runbooks, troubleshooting guides, support documentation, incident response processes, and operational tooling • Collaborate with platform, infrastructure, networking, deployment, and engineering teams • Develop automation scripts and tools, and improve observability dashboards and alerts • Mentor medior support engineers and lead knowledge-base and training initiatives

🎯 Requirements

• 5+ years of experience in cloud support, infrastructure operations, or systems administration • Experience supporting large-scale infrastructure environments or GPU clusters • Strong Linux systems administration skills • Experience troubleshooting compute, networking, and storage layers • Familiarity with Kubernetes platforms and containerized workloads • Experience with GPU hardware platforms or HPC environments • Familiarity with GPU monitoring tools and debugging GPU-related issues • Solid understanding of L2/L3 networking, routing, and load balancing • Ability to diagnose connectivity issues affecting distributed workloads • Strong troubleshooting and incident response skills • Experience participating in on-call rotations and handling production incidents • Ability to perform root cause analysis and drive operational improvements • Excellent written and verbal communication skills • Ability to explain complex technical concepts to technical and non-technical stakeholders • Proven ability to collaborate with cross-functional engineering teams • Technical stack includes Linux (Ubuntu), NVIDIA GPU platforms, CUDA drivers, nvidia-smi, Kubernetes, container runtimes, KubeVirt, NVIDIA Cumulus, TCP/IP, VLAN, VXLAN, OVS/OVN, BGP, VRFs, DNS, DHCP, RDMA, NVLink, NCCL, Grafana, Zabbix, Wazuh, TheHive, Cortex, Python, Bash, Ansible, Terraform, Jira, Confluence, Zendesk, PagerDuty, and Slack • Malaysia or comparable time zone • Average 40 hours per week with 9x5 business-hour support and after-hours on-call response/resolution for category 1 incidents

🏖️ Benefits

• Attractive compensation package reflecting your expertise and experience • A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach • Exciting career evolution in a fast-growing scale-up • Remote work modality • Flexible work environment • Equal opportunity employment

Apply Now