
51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
🔥 22 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
• Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads • Build and maintain CI/CD pipelines and deployment infrastructure • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability • Build dashboards, alerting, and anomaly detection across our systems • Define and enforce SLOs and build out incident response processes • Manage and improve our networking, load balancing, and service mesh configurations • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
• 5+ years experience in managing critical production systems and software development workflows • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible) • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD) • Proficiency in Python and either Go or Bash for tooling and automation • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog) • Excellent communication and ability to drive technical decisions across teams • Self-starter who executes quickly, takes ownership, and constantly seeks improvement • Nice to have: Experience with managing GPU and AI/ML workloads • Nice to have: Experience with kernel-based monitoring and routing (eBPF, XDP) • Nice to have: Experience with security tooling (Falco, Coroot, SIEM) • Nice to have: Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB) • Nice to have: Experience with distributed storage systems (Ceph, Longhorn, etc.)
• Interesting and challenging work • A lot of learning and growth opportunities • Regular team events and offsites
Apply Now🕒 6 days ago
Join Fundraise Up as a Senior DevOps Engineer. Lead DevOps initiatives managing monitoring and CI/CD for a global fundraising platform team.
🇹🇷 Turkey – Remote
💵 $5.2k - $5.9k / month
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🗣️🇷🇺 Russian Required
Ansible
Docker
Grafana
Jenkins
Kubernetes
Linux
Prometheus
Python