
51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
🕒 July 28
🌐 India, Australia, +1 more countries – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 18%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
🤖 Artificial Intelligence
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.
• Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform
• 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership • Strong systems fundamentals: distributed systems, networking, observability, and incident management • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production • Familiarity with security and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus • A bias toward automation, measurement, and blameless postmortems
• You will have access to our massive GPU cluster for inference and evaluation • Some core technologies we use include Python, torch, diffusers, Kubernetes, and the fal Python SDK
Apply Now🕒 July 28
Customer Deployment Engineer at Harrison.ai ensuring best deployment experience for healthcare AI solutions. Collaborating with customers and partners for successful integrations and support.
🇮🇳 India – Remote
💰 Series B on 2021-12
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
DNS
Linux
Python
TCP/IP
🕒 July 28
DevOps Engineer responsible for designing, building, and enhancing automation solutions using Microsoft Power Automate. Collaborating with global teams to improve workflows and operational efficiency.
Cloud
RPA
🕒 July 27
Senior Azure DevOps Engineer responsible for managing and optimising Azure cloud infrastructure. Join QuantumLoopAI, a healthtech company scaling its AI-platform solutions.
🇮🇳 India – Remote
💰 $2M Convertible note on 2025-02
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Azure
Cloud
Docker
Kubernetes
Terraform
🕒 July 27
DevOps Engineer responsible for improving customer experience through deployments and integrations. Collaborate on technical support and back-end system integration for enhanced operational efficiency.
Jenkins
Python
Ruby
🕒 July 24
Senior DevOps Engineer driving cloud cost optimization strategies across GCP, AWS, and Firebase. Collaborating with teams to improve resource efficiency and organizational impact.
AWS
BigQuery
Cloud
EC2
Firebase
Google Cloud Platform
Kubernetes
MongoDB
Node.js
Python
Terraform