Senior/Staff Kubernetes Infrastructure Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of fal

fal

51 - 200 employees

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Artificial Intelligence • API • SaaS

fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.

📋 Description

• Design, automate, validate, and deliver the complete lifecycle of customer compute environments, from provisioning through upgrades, recovery, and decommissioning • Use AI to automate and accelerate infrastructure delivery and operations • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads • Build and maintain Linux images and automated OS-provisioning workflows • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP • Configure distributed and shared storage for high-performance workloads • Build monitoring, alerting, diagnostics, and automated recovery for customer environments • Develop reusable tooling, standards, documentation, and runbooks • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs

🎯 Requirements

• 5+ years of experience building and operating production Linux infrastructure • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, and troubleshooting • Experience with Linux virtualization: KVM/QEMU, libvirt, and VFIO device passthrough • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry • Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark • Practical scripting experience • Experience with configuration-management tools such as Ansible • Ability to diagnose complex, cross-layer infrastructure issues • Strong communication and ability to drive technical decisions across teams • Track record of moving quickly, taking ownership, and continuously improving systems • Legally authorized to work in the United States • Nice-to-have: Production Slurm experience • Nice-to-have: High-performance networking experience with NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, or IMEX • Nice-to-have: Hugepages, NUMA, CPU pinning, SR-IOV, DPDK, Ceph, Lustre, Weka, KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, Nornir, AI training/inference/distributed GPU workload infrastructure, or Python/Go proficiency

🏖️ Benefits

• Equity • Salary range of $180K–$250K

Apply Now

Similar Jobs

🔥 4 hours ago

LatamCent

11 - 50

💼 Consulting

📣 Marketing

🎯 Recruiter

Lead security, infrastructure, reliability, and Web3 operations for a global payments and payroll platform. Own GCP, Cloudflare, CI/CD, disaster recovery, and Ethereum settlement security.

🔥 7 hours ago

AIP Publishing

51 - 200

📱 Media

🔬 Science

Cloud Infrastructure Engineer supporting Azure operations for AIP Publishing, a physical sciences publisher. Managing infrastructure, security, DevSecOps, incident response, and automation.

🕒 Yesterday

International Justice Mission

1001 - 5000

🤝 Non-profit

🤲 Charity

🌍 Social Impact

IT Infrastructure Specialist operating Microsoft Azure environments for International Justice Mission, a global organization protecting vulnerable people from violence. Automating, securing, and improving cloud platform operations.

🕒 Yesterday

Temporal Technologies

51 - 200

💼 Consulting

🏭 Manufacturing

📣 Marketing

Senior infrastructure engineer scaling Temporal’s open-source programming platform across cloud, compute, networking, and observability. Driving architecture, roadmaps, reliability, and cost optimization for infrastructure at scale.

🕒 Yesterday

Reddit, Inc.

501 - 1000

💼 Consulting

📣 Marketing

📱 Media

Senior ML infrastructure engineer building Reddit recommendation and personalization systems. Designing scalable training, evaluation, serving, and monitoring pipelines for high-traffic production ML.