
51 - 200 employees
Founded 2015
đ§ Hardware
⥠Energy
đ¤ B2B
Hardware ⢠Energy ⢠B2B
Submer is a provider of connected intelligence for AI infrastructure, specializing in advanced liquid/immersion cooling and modular data-center solutions. They design, build and operate AI-ready environments (from power and land to cloud and edge), manufacture immersion cooling pods (SmartPod EVO/EXO), and offer GPUaaS/AIaaS and modular deployment services. Submer emphasizes energy and water savings, high-density thermal architectures for demanding AI workloads, and sovereign-ready, globally deployable solutions.
đĽ 19 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2015
đ§ Hardware
⥠Energy
đ¤ B2B
Hardware ⢠Energy ⢠B2B
Submer is a provider of connected intelligence for AI infrastructure, specializing in advanced liquid/immersion cooling and modular data-center solutions. They design, build and operate AI-ready environments (from power and land to cloud and edge), manufacture immersion cooling pods (SmartPod EVO/EXO), and offer GPUaaS/AIaaS and modular deployment services. Submer emphasizes energy and water savings, high-density thermal architectures for demanding AI workloads, and sovereign-ready, globally deployable solutions.
⢠Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure ⢠Diagnose and resolve complex issues affecting GPU clusters, compute nodes, networking, and storage ⢠Investigate incidents across firmware, drivers, operating systems, and platform services ⢠Perform root cause analysis for major incidents and contribute to long-term remediation ⢠Serve as a technical escalation point for complex or high-priority support cases ⢠Troubleshoot GPU compute nodes, Kubernetes clusters, networking infrastructure, and storage systems ⢠Analyze logs, telemetry, and monitoring signals to identify platform instability ⢠Monitor and investigate security alerts; validate, triage, and escalate potential security incidents ⢠Support distributed GPU training and inference workloads, including multi-GPU and multi-node jobs ⢠Troubleshoot GPU workload scheduling, job queues, scheduling constraints, and resource fragmentation ⢠Diagnose high-performance networking issues involving RDMA and RoCE ⢠Coordinate with customer data center technicians for remote diagnostics and hardware interventions ⢠Validate on-premise GPU infrastructure installations and deployments ⢠Coordinate hardware replacements and RMA processes, and validate hardware health after replacements ⢠Participate in 24/7 on-call rotations and resolve incidents according to SLAs ⢠Improve runbooks, troubleshooting guides, support documentation, incident response processes, and operational tooling ⢠Collaborate with platform, infrastructure, networking, deployment, and engineering teams ⢠Develop automation scripts and tools, and improve observability dashboards and alerts ⢠Mentor medior support engineers and lead knowledge-base and training initiatives
⢠5+ years of experience in cloud support, infrastructure operations, or systems administration ⢠Experience supporting large-scale infrastructure environments or GPU clusters ⢠Strong Linux systems administration skills ⢠Experience troubleshooting compute, networking, and storage layers ⢠Familiarity with Kubernetes platforms and containerized workloads ⢠Experience with GPU hardware platforms or HPC environments ⢠Familiarity with GPU monitoring tools and debugging GPU-related issues ⢠Solid understanding of L2/L3 networking, routing, and load balancing ⢠Ability to diagnose connectivity issues affecting distributed workloads ⢠Strong troubleshooting and incident response skills ⢠Experience participating in on-call rotations and handling production incidents ⢠Ability to perform root cause analysis and drive operational improvements ⢠Excellent written and verbal communication skills ⢠Ability to explain complex technical concepts to technical and non-technical stakeholders ⢠Proven ability to collaborate with cross-functional engineering teams ⢠Technical stack includes Linux (Ubuntu), NVIDIA GPU platforms, CUDA drivers, nvidia-smi, Kubernetes, container runtimes, KubeVirt, NVIDIA Cumulus, TCP/IP, VLAN, VXLAN, OVS/OVN, BGP, VRFs, DNS, DHCP, RDMA, NVLink, NCCL, Grafana, Zabbix, Wazuh, TheHive, Cortex, Python, Bash, Ansible, Terraform, Jira, Confluence, Zendesk, PagerDuty, and Slack ⢠Malaysia or comparable time zone ⢠Average 40 hours per week with 9x5 business-hour support and after-hours on-call response/resolution for category 1 incidents
⢠Attractive compensation package reflecting your expertise and experience ⢠A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach ⢠Exciting career evolution in a fast-growing scale-up ⢠Remote work modality ⢠Flexible work environment ⢠Equal opportunity employment
Apply Now