
51 - 200 Mitarbeiter
🤖 Künstliche Intelligenz
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal ist eine generative Medienplattform für Entwickler, die Zugang zu einer umfangreichen Galerie produktionsreifer Bild-, Video-, Audio- und 3D-Generativmodelle bietet, sowie serverlose GPU-Inferenz und bedarfsabhängige Rechencluster zum Training und zur Feinabstimmung. Die Plattform bietet einheitliche APIs und SDKs, um Hunderte offener Modelle oder privater Gewichtungen abzurufen, eine leistungsstarke Inferenzengine, verwaltete serverlose GPU-Bereitstellungen und dedizierte Cluster mit moderner NVIDIA-Hardware für groß angelegte Trainings. fal richtet sich an Entwickler und Unternehmenskunden mit Features wie SOC 2 Compliance, privaten Endpunkten, Nutzungsanalysen und Unternehmenssupport und eignet sich zum Bauen, Implementieren und Skalieren generativer KI-gestützter Produkte.
🕒 vor 15 Tagen
🇺🇸 Vereinigte Staaten – Remote
💵 $180.000 - $250.000 / Jahr
⏰ Vollzeit
🟠 Senior
👷 IT-Infrastrukturingenieur
👻 Geisterscore 0%
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

51 - 200 Mitarbeiter
🤖 Künstliche Intelligenz
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal ist eine generative Medienplattform für Entwickler, die Zugang zu einer umfangreichen Galerie produktionsreifer Bild-, Video-, Audio- und 3D-Generativmodelle bietet, sowie serverlose GPU-Inferenz und bedarfsabhängige Rechencluster zum Training und zur Feinabstimmung. Die Plattform bietet einheitliche APIs und SDKs, um Hunderte offener Modelle oder privater Gewichtungen abzurufen, eine leistungsstarke Inferenzengine, verwaltete serverlose GPU-Bereitstellungen und dedizierte Cluster mit moderner NVIDIA-Hardware für groß angelegte Trainings. fal richtet sich an Entwickler und Unternehmenskunden mit Features wie SOC 2 Compliance, privaten Endpunkten, Nutzungsanalysen und Unternehmenssupport und eignet sich zum Bauen, Implementieren und Skalieren generativer KI-gestützter Produkte.
• Design, automate, validate, and deliver the complete lifecycle of customer compute environments, from provisioning through upgrades, recovery, and decommissioning • Use AI to automate and accelerate infrastructure delivery and operations • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads • Build and maintain Linux images and automated OS-provisioning workflows • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP • Configure distributed and shared storage for high-performance workloads • Build monitoring, alerting, diagnostics, and automated recovery for customer environments • Develop reusable tooling, standards, documentation, and runbooks • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs
• 5+ years of experience building and operating production Linux infrastructure • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, and troubleshooting • Experience with Linux virtualization: KVM/QEMU, libvirt, and VFIO device passthrough • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry • Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark • Practical scripting experience • Experience with configuration-management tools such as Ansible • Ability to diagnose complex, cross-layer infrastructure issues • Strong communication and ability to drive technical decisions across teams • Track record of moving quickly, taking ownership, and continuously improving systems • Legally authorized to work in the United States • Nice-to-have: Production Slurm experience • Nice-to-have: High-performance networking experience with NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, or IMEX • Nice-to-have: Hugepages, NUMA, CPU pinning, SR-IOV, DPDK, Ceph, Lustre, Weka, KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, Nornir, AI training/inference/distributed GPU workload infrastructure, or Python/Go proficiency
• Equity • Salary range of $180K–$250K
Jetzt Bewerben🕒 vor 15 Tagen
Lead security, infrastructure, reliability, and Web3 operations for a global payments and payroll platform. Own GCP, Cloudflare, CI/CD, disaster recovery, and Ethereum settlement security.
🇺🇸 Vereinigte Staaten – Remote
💵 $170.000 - $210.000 / Jahr
⏰ Vollzeit
🟠 Senior
👷 IT-Infrastrukturingenieur
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 15 Tagen
Cloud Infrastructure Engineer supporting Azure operations for AIP Publishing, a physical sciences publisher. Managing infrastructure, security, DevSecOps, incident response, and automation.
🇺🇸 Vereinigte Staaten – Remote
💵 $125.000 - $135.000 / Jahr
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
👷 IT-Infrastrukturingenieur
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 16 Tagen
IT Infrastructure Specialist operating Microsoft Azure environments for International Justice Mission, a global organization protecting vulnerable people from violence. Automating, securing, and improving cloud platform operations.
🇺🇸 Vereinigte Staaten – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
👷 IT-Infrastrukturingenieur
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 17 Tagen
Senior ML infrastructure engineer building Reddit recommendation and personalization systems. Designing scalable training, evaluation, serving, and monitoring pipelines for high-traffic production ML.
🇺🇸 Vereinigte Staaten – Remote
💵 $190.800 - $267.100 / Jahr
⏰ Vollzeit
🟠 Senior
👷 IT-Infrastrukturingenieur
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 17 Tagen
PaaS database automation consultant designing and managing Azure cloud databases for HCSC, a major U.S. health insurer. Automating provisioning, monitoring, standards, and security across healthcare infrastructure.
🇺🇸 Vereinigte Staaten – Remote
💵 $112.200 - $202.600 / Jahr
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
👷 IT-Infrastrukturingenieur
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich