Linux Infrastructure Engineer, Bare Metal, Storage, AI Factory Infrastructure

Job not on LinkedIn

🔥 0 minutes ago

🇸🇬 Singapore – Remote

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Uvation

Uvation

11 - 50 employees

💼 Consulting

🏥 Healthcare

🤖 Artificial Intelligence

Consulting • Healthcare • Artificial Intelligence

Uvation is a comprehensive technology services company providing innovative solutions across the IT infrastructure space. They specialize in artificial intelligence and cutting-edge cybersecurity solutions, offering products powered by Dell and Nvidia AI. Uvation delivers various managed services, including IT operations, security operations, and network operations, optimizing performance and enhancing security for businesses. Their services extend into public clouds with partnerships involving major platforms like Oracle Cloud, Google Cloud, and Amazon AWS. Uvation also operates a marketplace offering competitive pricing on a wide range of hardware and software products and provides a rewards program to incentivize customer engagement.

📋 Description

• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure • Build and manage infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms • Deploy and manage bare metal servers and BMaaS platforms • Design, implement, and support enterprise Linux infrastructure at scale • Deploy and manage GPU-accelerated infrastructure for AI/ML workloads • Support GPU clusters, AI training environments, and HPC workloads • Provision, monitor, optimize, and manage GPU infrastructure lifecycle • Administer enterprise storage and Ceph platforms, including capacity planning, performance tuning, and failure recovery • Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure • Support high-availability, clustering, disaster recovery, and mission-critical production environments • Troubleshoot operating system, hardware, GPU, networking, and storage challenges • Automate operational work using Bash and Python • Create operational documentation, runbooks, and infrastructure standards

🎯 Requirements

• Senior-level, highly experienced Linux infrastructure engineering expertise • Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments • Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale • Experience deploying and managing GPU-accelerated AI/ML infrastructure • Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms • Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance optimization • Knowledge of AI Factory architecture and infrastructure requirements • Experience supporting GPU clusters, AI training environments, and HPC workloads • Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design • Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command preferred • Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O • Strong hands-on Ceph experience, including cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery • Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, or NetApp • Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines • Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP • Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting • Experience with high availability, clustering, and disaster recovery • Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage • Bash and Python scripting for automation and operational efficiency • Experience creating operational documentation, runbooks, and infrastructure standards • Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC workload schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure/hybrid cloud exposure listed as nice to have • Not primarily a CI/CD, Terraform, GitOps, application delivery, cloud-only, or software development candidate

Apply Now