
51 - 200 funcionários
🤖 Inteligência Artificial
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal é uma plataforma de mídia generativa para desenvolvedores que oferece acesso a uma grande galeria de modelos generativos de imagens, vídeos, áudio e 3D prontos para produção, juntamente com inferência de GPU sem servidor e clusters de computação sob demanda para treinamento e ajuste fino. A plataforma oferece APIs e SDKs unificados para acessar centenas de modelos abertos ou pesos privados, um mecanismo de inferência de alto desempenho, implantações de GPU sem servidor gerenciado e clusters dedicados com hardware NVIDIA moderno para treinamento em larga escala. fal atende a desenvolvedores e clientes corporativos com recursos como conformidade SOC 2, endpoints privados, análise de uso e suporte empresarial, e está posicionada para construir, implantar e escalar produtos alimentados por IA generativa.
🕒 Julho 28
🌐 Índia, Austrália, +1 outros países – Remoto
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
👻 Score fantasma 21%
🗣️🇺🇸🇬🇧 Inglês obrigatório
Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

51 - 200 funcionários
🤖 Inteligência Artificial
🔌 API
☁️ SaaS
Artificial Intelligence • API • SaaS
fal é uma plataforma de mídia generativa para desenvolvedores que oferece acesso a uma grande galeria de modelos generativos de imagens, vídeos, áudio e 3D prontos para produção, juntamente com inferência de GPU sem servidor e clusters de computação sob demanda para treinamento e ajuste fino. A plataforma oferece APIs e SDKs unificados para acessar centenas de modelos abertos ou pesos privados, um mecanismo de inferência de alto desempenho, implantações de GPU sem servidor gerenciado e clusters dedicados com hardware NVIDIA moderno para treinamento em larga escala. fal atende a desenvolvedores e clientes corporativos com recursos como conformidade SOC 2, endpoints privados, análise de uso e suporte empresarial, e está posicionada para construir, implantar e escalar produtos alimentados por IA generativa.
• Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform
• 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership • Strong systems fundamentals: distributed systems, networking, observability, and incident management • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production • Familiarity with security and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus • A bias toward automation, measurement, and blameless postmortems
• You will have access to our massive GPU cluster for inference and evaluation • Some core technologies we use include Python, torch, diffusers, Kubernetes, and the fal Python SDK
Candidatar-se🕒 Julho 28
DevOps Engineer responsible for designing, building, and enhancing automation solutions using Microsoft Power Automate. Collaborating with global teams to improve workflows and operational efficiency.
🇮🇳 Índia – Remoto
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
Cloud
RPA
🕒 Julho 27
Senior Azure DevOps Engineer responsible for managing and optimising Azure cloud infrastructure. Join QuantumLoopAI, a healthtech company scaling its AI-platform solutions.
🇮🇳 Índia – Remoto
💰 $2.000.000 Convertible note em 2025-02
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 27
DevOps Engineer responsible for improving customer experience through deployments and integrations. Collaborate on technical support and back-end system integration for enhanced operational efficiency.
🇮🇳 Índia – Remoto
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 24
Senior DevOps Engineer driving cloud cost optimization strategies across GCP, AWS, and Firebase. Collaborating with teams to improve resource efficiency and organizational impact.
🇮🇳 Índia – Remoto
💰 Series A em 2021-11
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 21
Senior DevOps Engineer focused on cloud automation and operational reliability for Govtech solutions. Leading technical projects and mentoring engineers in complex application environments.
🗣️🇺🇸🇬🇧 Inglês obrigatório