Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure

Job not on LinkedIn

🔥 0 minutes ago

🇷🇴 Romania – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Uvation

Uvation

11 - 50 employees

💼 Consulting

🏥 Healthcare

🤖 Artificial Intelligence

Consulting • Healthcare • Artificial Intelligence

Uvation is a comprehensive technology services company providing innovative solutions across the IT infrastructure space. They specialize in artificial intelligence and cutting-edge cybersecurity solutions, offering products powered by Dell and Nvidia AI. Uvation delivers various managed services, including IT operations, security operations, and network operations, optimizing performance and enhancing security for businesses. Their services extend into public clouds with partnerships involving major platforms like Oracle Cloud, Google Cloud, and Amazon AWS. Uvation also operates a marketplace offering competitive pricing on a wide range of hardware and software products and provides a rewards program to incentivize customer engagement.

📋 Description

• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure for enterprise workloads and AI/ML environments • Build and manage infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms • Deploy, provision, operate, and lifecycle-manage bare metal servers and BMaaS platforms • Deploy and manage GPU-accelerated infrastructure, GPU clusters, and AI training environments • Support high-performance computing (HPC) workloads and AI Factory environments • Manage enterprise Linux storage, Ceph clusters, and high-performance AI storage platforms • Design and troubleshoot high-bandwidth, low-latency data center networking • Support mission-critical production environments with high availability, clustering, and disaster recovery • Troubleshoot operating system, hardware, GPU, networking, and storage issues • Automate operational tasks using Bash and Python • Create operational documentation, runbooks, and infrastructure standards • Collaborate with the dedicated DevOps team while focusing on infrastructure rather than CI/CD or application delivery

🎯 Requirements

• Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments • Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads • Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms • Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance optimization • Knowledge of AI Factory architecture, GPU clusters, AI training environments, and HPC workloads • Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design • Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command • Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O • Strong hands-on experience with Ceph, including cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery • Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp • Understanding of NVMe-over-Fabrics (NVMe-oF), RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines • Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP • Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting • Experience with high availability, clustering, and disaster recovery • Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage • Bash and Python scripting for automation and operational efficiency • Experience creating operational documentation, runbooks, and infrastructure standards • Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure/hybrid cloud exposure are nice to have

Apply Now