Search Remote Jobs

Site Reliability Engineer, Provider Operations

šŸ”„ 14 hours ago

šŸ‡ŗšŸ‡ø United States – Remote

ā° Full Time

🟔 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

šŸ‘» Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

šŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of OpenRouter

OpenRouter

1 - 10 employees

Founded 2023

šŸ’¼ Consulting

šŸ“£ Marketing

šŸ¤– Artificial Intelligence

Consulting • Marketing • Artificial Intelligence

OpenRouter is a platform providing a unified interface for interacting with various large language models (LLMs). It focuses on delivering better prices and uptime without the need for subscriptions. The platform supports multiple AI tasks and showcases models like Mistral Small 3 and DeepSeek R1 Distill Qwen 32B, which are designed for fast performance and high accuracy. OpenRouter also features gaming and coding tools that enhance usability across different sectors such as marketing, finance, and education.

šŸ“‹ Description

• Build and own monitoring for every provider and endpoint, covering latency, throughput, error rates, uptime, and output correctness • Set SLOs per provider tier and create actionable alerts • Improve detection of degraded endpoints and coordinate automatic traffic failover with the routing team • Own on-call for provider incidents, including triage, mitigation, provider communication, postmortems, and follow-up closure • Create provider scorecards and SLO reporting • Serve as the technical escalation point for misbehaving provider endpoints • Build continuous canaries and evaluations to detect silent quality regressions • Automate provider-operations tasks such as disabling endpoints, capacity changes, deprecations, and rate-limit tuning • Build endpoint load-testing tools for launch readiness and day-zero traffic • Report to the Provider Operations Manager

šŸŽÆ Requirements

• 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems • Strong observability practice with metrics, tracing, logs, SLOs/error budgets, and actionable alerting • Capable software engineer who prefers writing tools over executing runbooks • Experience with TypeScript and/or Python • Experience with distributed systems failure modes including timeouts, retries, backpressure, and partial outages • Ability to serve as a calm, clear incident commander and communicate with external partners under pressure • Understanding of, or eagerness to learn, LLM inference concepts including streaming, tool calling, prompt caching, throughput/latency tradeoffs, and provider API differences • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company is nice to have • Experience with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is nice to have • Background in routing, load balancing, or traffic management systems is nice to have • Experience with evals or synthetic monitoring for ML systems is nice to have

Apply Now

Similar Jobs

šŸ”„ 15 hours ago

Mirantis

501 - 1000

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ“¦ Logistics

Senior Data Platform Engineer operating PostgreSQL and Kafka on Kubernetes for Mirantis, a Kubernetes-native AI infrastructure company. Ensuring reliable, secure, multi-region data services for enterprise GPU infrastructure.

šŸ”„ 16 hours ago

LMI

1001 - 5000

šŸ“¦ Logistics

šŸ„ Healthcare

šŸŽ–ļø Defense

DevOps Engineer automating cloud infrastructure and CI/CD for LMI’s federal SHEPRD application. Supporting secure, reliable deployments across development, test, and production environments.

šŸ”„ 17 hours ago

Encoura

51 - 200

šŸ’¼ Consulting

šŸ“£ Marketing

šŸ“š Education

Senior DevOps Engineer designing and scaling Azure infrastructure for Encoura, a higher-education technology company. Automating deployments, improving reliability, and supporting real-time student and institutional systems.

šŸ”„ 19 hours ago

Slate Auto

201 - 500

🚘 Automotive

šŸ­ Manufacturing

šŸš— Transport

Senior DevOps Engineer building AWS and Kubernetes infrastructure for Slate’s affordable, customizable vehicles. Operating platforms, CI/CD pipelines and production reliability systems.

šŸ”„ 21 hours ago

Akamai Technologies

5001 - 10000

šŸ”’ Cybersecurity

Senior Site Reliability Engineer maintaining Akamai's Linux kernels, KVM/QEMU virtualization, and cloud infrastructure. Supporting distributed Compute products that make digital experiences faster and more secure.