Senior Site Reliability Engineer

🕒 July 2

🌐 United States, Canada – Remote

info

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akuity

Akuity

11 - 50 employees

Founded 2021

🏢 Enterprise

☁️ SaaS

Enterprise • SaaS • Cloud

Akuity is a company that provides an end-to-end GitOps platform for Kubernetes, focusing on deployment, promotion, and monitoring with tools like Argo CD, Kargo, and KubeVision. The platform extends Kubernetes APIs to enhance continuous delivery, container orchestration, and event automation. Akuity's solutions aim to simplify infrastructure management, improve security, and increase deployment efficiency, demonstrating significant time savings and improved shipping velocity for developers. Akuity also offers enterprise support for Argo, emphasizing security, compliance, and scalability for Kubernetes deployments.

📋 Description

• Own SLI/SLO/SLA definitions for the Akuity SaaS platform and drive continuous improvement against them • Design, instrument, and maintain observability systems (metrics, logs, traces) across multi-region AWS infrastructure • Identify reliability gaps, lead blameless post-mortems, and close the loop with permanent fixes • Partner with engineering teams to build reliability into new features before they ship to production • Participate in an on-call rotation and act as incident commander for high-severity production events • Build and maintain runbooks, escalation paths, and incident playbooks that keep mean time to resolution low • Drive improvements to alerting fidelity; reduce noise, increase signal, eliminate toil • Lead post-incident reviews with clear timelines, root cause analysis, and follow-through on action items

🎯 Requirements

• 5+ years of SRE, platform engineering, or production operations experience in a SaaS environment • Deep hands-on Kubernetes expertise; you understand the scheduler, networking, storage, and autoscaling at a level where you can debug anything • Strong AWS fundamentals across compute (EC2, EKS), networking (VPC, NLB, Route53), storage (S3, RDS), and IAM • Experience defining and operating against SLOs in production; you've written error budgets, not just read about them • Proficiency with observability tooling (Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent) • Solid scripting and automation skills; Go, Python, Bash, or similar; you automate what you touch • Strong written communication: clear runbooks, sharp incident reports, thoughtful post-mortems • Live within US time zones (Pacific through Eastern), including Canada and other regions

🏖️ Benefits

• Competitive compensation, commensurate with experience • Equity participation in a well-funded, growing company • Fully remote: work from anywhere within US time zones (Pacific through Eastern), including Canada and other regions • Home office stipend and equipment budget • Flexible time off and a culture that respects it • Work directly with the engineers who built Argo CD and Kargo; you'll learn a lot here • US-based employees receive full benefits, including comprehensive health, dental, and vision coverage. Candidates based outside the US will be engaged as contractors.

Apply Now

Similar Jobs

🕒 July 2

Sanity.io

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

SRE managing scalable content operations infrastructure for AI-powered platform. Collaborating with dev teams and ensuring reliability for high request volume systems.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 2

Global Alliant Inc

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Full Stack Software Engineer in Agile teams supporting federal technology initiatives. Responsible for building secure, scalable, cloud-native applications using modern tech stacks.

🕒 July 1

Upstart

1001 - 5000

🚘 Automotive

💼 Consulting

🏥 Healthcare

DevOps Engineer focused on enhancing cloud infrastructure reliability and performance at Upstart. Partnering with cross-functional teams to optimize Kubernetes and AWS cloud services.

🕒 July 1

Buyers Edge Platform

501 - 1000

📦 Logistics

🏭 Manufacturing

🏨 Hospitality

DevOps Engineer improving reliability and operational efficiency of hosted infrastructure and applications for Buyers Edge Platform. Designing and building tooling to streamline workflows in foodservice tech.

🇺🇸 United States – Remote

💰 $425M Private Equity Round on 2024-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 1

Ad Hoc LLC

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevSecOps Engineer working with federal enterprise cloud platform. Designing secure, automated CI/CD pipelines and collaborating with government stakeholders.

🇺🇸 United States – Remote

💵 $145k - $160k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)