Software Engineer, Infrastructure

🔥 20 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of fal

fal

51 - 200 employees

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Artificial Intelligence • API • SaaS

fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.

📋 Description

• Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) • Leverage AI to an extreme level to build tools and automate alerting and recovery • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) • Develop a suite of automated error detection and recovery processes • Work with partners to solve technical issues

🎯 Requirements

• 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes) • Strong software engineering skills in Python; you write production tooling, not scripts • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory) • Experience building internal tools or dashboards for infrastructure visibility • Excellent communication and ability to drive technical decisions across teams • Self-starter who executes quickly, takes ownership, and constantly seeks improvement • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump) (Nice to have) • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2 (Nice to have) • Experience with AMD GPUs (Nice to have) • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM) (Nice to have) • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001) (Nice to have)

🏖️ Benefits

• Interesting and challenging work • A lot of learning and growth opportunities • Regular team events and offsites

Apply Now