Release Engineer

Job not on LinkedIn

🕒 July 27

🌏 Anywhere in the World

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 30%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Supabase

Supabase

51 - 200 employees

Founded 2020

☁️ SaaS

🔌 API

🤖 Artificial Intelligence

💰 $80M Series B on 2022-05

SaaS • API • Artificial Intelligence

Supabase is an open source alternative to Firebase, providing a range of backend tools designed to help developers start and scale their applications effectively. It offers features such as a full Postgres database, authentication with Row Level Security, instant APIs, Edge Functions for custom code, real-time data synchronization, and storage for large files. Developers can integrate machine learning models, utilize RESTful APIs, and take advantage of platform-integrated best of breed products. Supabase is designed to be highly portable, extendable, and user-friendly, making it a powerful choice for startups and enterprises looking to innovate quickly and efficiently.

📋 Description

• Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets • Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows • Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies) • Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do • Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load • Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil • Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom • Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal • Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise • Ensure deployments fail fast and safely when health checks degrade • Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds • Partner with product engineering and platform teams to align release practices with reliability and availability targets

🎯 Requirements

• Have 5+ years in SRE, production operations, platform engineering, or release engineering • Have operated production systems at scale and carried on-call for them • Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar) • Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR • Operate confidently on AWS (multiple accounts, IAM, VPC) in production • Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes • Script and automate to eliminate toil rather than absorb it • Communicate clearly with both infrastructure specialists and product engineers • Thrive in async, globally distributed teams • Are comfortable navigating ambiguity and iterating toward better systems over time.

🏖️ Benefits

• Fully Remote • ESOP • Tech Allowance • Health Benefits • Annual Off-Sites • Flexible Work • Professional Development

Apply Now

Similar Jobs

🕒 July 11

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Site Reliability Engineer ensuring reliability of GitLab's user-facing services. Supporting operational excellence through engineering principles and automation.

AWS

Cloud

Google Cloud Platform

Kubernetes

Ruby

Terraform

Go

🕒 July 8

Onchain

51 - 200

💸 Finance

💳 Fintech

₿ Crypto

DevOps Engineer evolving cloud infrastructure behind Lisk's financial platform as we scale. Join a remote-first team to improve security and reliability of money-handling systems.

AWS

Cloud

Terraform

🕒 May 28

Chess.com

501 - 1000

🎮 Gaming

📚 Education

📱 Media

Site Reliability Engineer at Chess.com ensuring infrastructure stability and scalable systems for millions of users. Playing a critical role in supporting rapid feature development and deployment.

Ansible

AWS

Azure

Chef

Cloud

DNS

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

NoSQL

Prometheus

Puppet

Python

TCP/IP

Terraform

Unix

Go

🕒 May 13

Shuru

51 - 200

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

Senior DevOps Engineer helping scale cloud platform from pre-production to production for fintech. Collaborating with teams to enhance infrastructure, deployment, and monitoring processes.

AWS

Azure

Cloud

Google Cloud Platform

Grafana

Kubernetes

Oracle

Postgres

Prometheus

Redis

SQL

Terraform

🕒 April 1

Canonical

501 - 1000

🤖 Artificial Intelligence

Site Reliability / Gitops Engineer supporting and maintaining Canonical’s IT production services. Automating operations with Infrastructure as Code for private and public cloud environments.

Cloud

Distributed Systems

ElasticSearch

Firewalls

Grafana

Linux

Prometheus

Python