
51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
🔥 0 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
• Design, automate, validate, and deliver the complete lifecycle of customer compute environments, from provisioning through upgrades, recovery, and decommissioning • Use AI to automate and accelerate infrastructure delivery and operations • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads • Build and maintain Linux images and automated OS-provisioning workflows • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP • Configure distributed and shared storage for high-performance workloads • Build monitoring, alerting, diagnostics, and automated recovery for customer environments • Develop reusable tooling, standards, documentation, and runbooks • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs
• 5+ years of experience building and operating production Linux infrastructure • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, and troubleshooting • Experience with Linux virtualization: KVM/QEMU, libvirt, and VFIO device passthrough • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry • Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark • Practical scripting experience • Experience with configuration-management tools such as Ansible • Ability to diagnose complex, cross-layer infrastructure issues • Strong communication and ability to drive technical decisions across teams • Track record of moving quickly, taking ownership, and continuously improving systems • Legally authorized to work in the United States • Nice-to-have: Production Slurm experience • Nice-to-have: High-performance networking experience with NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, or IMEX • Nice-to-have: Hugepages, NUMA, CPU pinning, SR-IOV, DPDK, Ceph, Lustre, Weka, KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, Nornir, AI training/inference/distributed GPU workload infrastructure, or Python/Go proficiency
• Equity • Salary range of $180K–$250K
Apply Now🔥 4 hours ago
Lead security, infrastructure, reliability, and Web3 operations for a global payments and payroll platform. Own GCP, Cloudflare, CI/CD, disaster recovery, and Ethereum settlement security.
🔥 7 hours ago
Cloud Infrastructure Engineer supporting Azure operations for AIP Publishing, a physical sciences publisher. Managing infrastructure, security, DevSecOps, incident response, and automation.
🇺🇸 United States – Remote
💵 $125k - $135k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
👷 Infrastructure Engineer
🕒 Yesterday
IT Infrastructure Specialist operating Microsoft Azure environments for International Justice Mission, a global organization protecting vulnerable people from violence. Automating, securing, and improving cloud platform operations.
🕒 Yesterday
Senior infrastructure engineer scaling Temporal’s open-source programming platform across cloud, compute, networking, and observability. Driving architecture, roadmaps, reliability, and cost optimization for infrastructure at scale.
🇺🇸 United States – Remote
💵 $176k - $237.6k / year
💰 $75M Series B on 2023-02
⏰ Full Time
🟠 Senior
👷 Infrastructure Engineer
🦅 H1B Visa Sponsor
🕒 Yesterday
Senior ML infrastructure engineer building Reddit recommendation and personalization systems. Designing scalable training, evaluation, serving, and monitoring pipelines for high-traffic production ML.