Release Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Supabase

Supabase

51 - 200 employees

Founded 2020

☁️ SaaS

🔌 API

🤖 Artificial Intelligence

💰 $80M Series B on 2022-05

SaaS • API • Artificial Intelligence

Supabase is an open source alternative to Firebase, providing a range of backend tools designed to help developers start and scale their applications effectively. It offers features such as a full Postgres database, authentication with Row Level Security, instant APIs, Edge Functions for custom code, real-time data synchronization, and storage for large files. Developers can integrate machine learning models, utilize RESTful APIs, and take advantage of platform-integrated best of breed products. Supabase is designed to be highly portable, extendable, and user-friendly, making it a powerful choice for startups and enterprises looking to innovate quickly and efficiently.

📋 Description

• Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets • Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows • Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies) • Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do • Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load • Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil • Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom • Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal • Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise • Ensure deployments fail fast and safely when health checks degrade • Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds • Partner with product engineering and platform teams to align release practices with reliability and availability targets

🎯 Requirements

• Have 5+ years in SRE, production operations, platform engineering, or release engineering • Have operated production systems at scale and carried on-call for them • Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar) • Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR • Operate confidently on AWS (multiple accounts, IAM, VPC) in production • Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes • Script and automate to eliminate toil rather than absorb it • Communicate clearly with both infrastructure specialists and product engineers • Thrive in async, globally distributed teams • Are comfortable navigating ambiguity and iterating toward better systems over time.

🏖️ Benefits

• Fully Remote • ESOP • Tech Allowance • Health Benefits • Annual Off-Sites • Flexible Work • Professional Development

Apply Now

Similar Jobs

🕒 6 days ago

Social Discovery Group

1001 - 5000

🌍 Social Impact

📱 Media

DevOps Engineer developing internal services and tools in Go for social discovery products. Collaborating on deployment automation in a remote working environment with a global team.

🗣️🇷🇺 Russian Required

Ansible

Grafana

Kubernetes

Linux

Prometheus

Terraform

Go

🕒 July 11

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Site Reliability Engineer ensuring reliability of GitLab's user-facing services. Supporting operational excellence through engineering principles and automation.

AWS

Cloud

Google Cloud Platform

Kubernetes

Ruby

Terraform

Go

🕒 July 9

Growe Talents

11 - 50

💼 Consulting

📣 Marketing

🎯 Recruiter

DevSecOps role to build and own the security function at Growe Talents, assessing security posture and implementing controls across the development lifecycle.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 July 8

Onchain

51 - 200

💸 Finance

💳 Fintech

₿ Crypto

DevOps Engineer evolving cloud infrastructure behind Lisk's financial platform as we scale. Join a remote-first team to improve security and reliability of money-handling systems.

AWS

Cloud

Terraform

🕒 May 28

Chess.com

501 - 1000

🎮 Gaming

📚 Education

📱 Media

Site Reliability Engineer at Chess.com ensuring infrastructure stability and scalable systems for millions of users. Playing a critical role in supporting rapid feature development and deployment.

Ansible

AWS

Azure

Chef

Cloud

DNS

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

NoSQL

Prometheus

Puppet

Python

TCP/IP

Terraform

Unix

Go