Software Engineer, Site Reliability

🔥 22 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of fal

fal

51 - 200 employees

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Artificial Intelligence • API • SaaS

fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.

📋 Description

• Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads • Build and maintain CI/CD pipelines and deployment infrastructure • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability • Build dashboards, alerting, and anomaly detection across our systems • Define and enforce SLOs and build out incident response processes • Manage and improve our networking, load balancing, and service mesh configurations • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering

🎯 Requirements

• 5+ years experience in managing critical production systems and software development workflows • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible) • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD) • Proficiency in Python and either Go or Bash for tooling and automation • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog) • Excellent communication and ability to drive technical decisions across teams • Self-starter who executes quickly, takes ownership, and constantly seeks improvement • Nice to have: Experience with managing GPU and AI/ML workloads • Nice to have: Experience with kernel-based monitoring and routing (eBPF, XDP) • Nice to have: Experience with security tooling (Falco, Coroot, SIEM) • Nice to have: Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB) • Nice to have: Experience with distributed storage systems (Ceph, Longhorn, etc.)

🏖️ Benefits

• Interesting and challenging work • A lot of learning and growth opportunities • Regular team events and offsites

Apply Now

Similar Jobs

🕒 6 days ago

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

Join Fundraise Up as a Senior DevOps Engineer. Lead DevOps initiatives managing monitoring and CI/CD for a global fundraising platform team.

🗣️🇷🇺 Russian Required

Ansible

Docker

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python