Machine Learning Engineer, Reliability

🕒 July 28

🌐 India, Australia, +1 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 18%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of fal

fal

51 - 200 employees

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Artificial Intelligence • API • SaaS

fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.

📋 Description

• Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform

🎯 Requirements

• 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership • Strong systems fundamentals: distributed systems, networking, observability, and incident management • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production • Familiarity with security and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus • A bias toward automation, measurement, and blameless postmortems

🏖️ Benefits

• You will have access to our massive GPU cluster for inference and evaluation • Some core technologies we use include Python, torch, diffusers, Kubernetes, and the fal Python SDK

Apply Now

Similar Jobs

🕒 July 28

harrison.ai

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Customer Deployment Engineer at Harrison.ai ensuring best deployment experience for healthcare AI solutions. Collaborating with customers and partners for successful integrations and support.

DNS

Linux

Python

TCP/IP

🕒 July 28

BCD Travel

10,000+ employees

💼 Consulting

📦 Logistics

🏨 Hospitality

DevOps Engineer responsible for designing, building, and enhancing automation solutions using Microsoft Power Automate. Collaborating with global teams to improve workflows and operational efficiency.

Cloud

RPA

🕒 July 27

QuantumLoopAi

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Senior Azure DevOps Engineer responsible for managing and optimising Azure cloud infrastructure. Join QuantumLoopAI, a healthtech company scaling its AI-platform solutions.

Azure

Cloud

Docker

Kubernetes

Terraform

🕒 July 27

Coforgetech Ltd

-

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

DevOps Engineer responsible for improving customer experience through deployments and integrations. Collaborate on technical support and back-end system integration for enhanced operational efficiency.

Jenkins

Python

Ruby

🕒 July 24

HighLevel

201 - 500

💼 Consulting

📦 Logistics

☁️ SaaS

Senior DevOps Engineer driving cloud cost optimization strategies across GCP, AWS, and Firebase. Collaborating with teams to improve resource efficiency and organizational impact.

AWS

BigQuery

Cloud

EC2

Firebase

Google Cloud Platform

Kubernetes

MongoDB

Node.js

Python

Terraform