
51 - 200 employees
đ¤ Artificial Intelligence
đ API
âď¸ SaaS
Artificial Intelligence ⢠API ⢠SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
đ July 28
đšđˇ Turkey â Remote
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 18%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
đ¤ Artificial Intelligence
đ API
âď¸ SaaS
Artificial Intelligence ⢠API ⢠SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
⢠Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads ⢠Build and maintain CI/CD pipelines and deployment infrastructure ⢠Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability ⢠Build dashboards, alerting, and anomaly detection across our systems ⢠Define and enforce SLOs and build out incident response processes ⢠Manage and improve our networking, load balancing, and service mesh configurations ⢠Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
⢠5+ years experience in managing critical production systems and software development workflows ⢠Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible) ⢠Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS ⢠Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD) ⢠Proficiency in Python and either Go or Bash for tooling and automation ⢠Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog) ⢠Excellent communication and ability to drive technical decisions across teams ⢠Self-starter who executes quickly, takes ownership, and constantly seeks improvement ⢠Nice to have: Experience with managing GPU and AI/ML workloads ⢠Nice to have: Experience with kernel-based monitoring and routing (eBPF, XDP) ⢠Nice to have: Experience with security tooling (Falco, Coroot, SIEM) ⢠Nice to have: Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB) ⢠Nice to have: Experience with distributed storage systems (Ceph, Longhorn, etc.)
⢠Interesting and challenging work ⢠A lot of learning and growth opportunities ⢠Regular team events and offsites
Apply Nowđ July 21
Join Fundraise Up as a Senior DevOps Engineer. Lead DevOps initiatives managing monitoring and CI/CD for a global fundraising platform team.
đšđˇ Turkey â Remote
đľ $5.2k - $5.9k / month
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŁď¸đˇđş Russian Required
Ansible
Docker
Grafana
Jenkins
Kubernetes
Linux
Prometheus
Python