Site Reliability Engineer

🕒 August 25

☕ Washington – Remote

infoinfo

💵 $123k - $150k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Kong Inc.

Kong Inc.

201 - 500 employees

Founded 2017

💼 Consulting

📦 Logistics

🔌 API

💰 $100M Series D on 2021-02

Consulting • Logistics • API

Kong Inc. is a company that provides a comprehensive API platform designed to facilitate API management, AI integration, and developer productivity. It offers solutions like Kong Gateway, Kong Konnect, and a variety of other tools targeted at managing and optimizing the API lifecycle. Kong's platform supports multi-cloud environments and is built to deliver high performance and security. It is notably recognized by Gartner as a leader in API management and supports innovations across industries like financial services, healthcare, and technology. The company emphasizes flexibility, security, and speed, making it a favored choice for enterprises looking to enhance their digital services through APIs. Kong also supports a robust community of developers and provides extensive integrations and plugins to streamline API management and operations.

📋 Description

• Operate and scale Kong’s global SaaS platform, Konnect, across regions and clouds • Build, automate, and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD • Design, maintain, and optimize multi-region PostgreSQL, Redis, ClickHouse, and Druid data and caching layers • Operate and improve Kong Gateway and Kong Mesh environments • Develop and maintain CI/CD pipelines and GitOps workflows • Enhance observability and incident response readiness using Datadog, Prometheus, Grafana, and Thanos; define and track SLOs • Collaborate with development and security teams to operate SaaS services in compliance with reliability, security, and regulatory standards • Participate in a global 24/7 on-call rotation and improve operational playbooks and postmortem practices • Lead and contribute to scaling initiatives that improve elasticity, reliability, and cost-efficiency

🎯 Requirements

• BS in Computer Science or equivalent practical experience • Proven experience managing SaaS or PaaS systems at enterprise scale • Deep expertise in Kubernetes, including debugging cluster/networking issues and designing for fault tolerance and scalability • Strong proficiency with Terraform or Terragrunt • Experience with CI/CD pipelines and GitOps workflows, including ArgoCD, Atlantis, and Helm • Proficiency in Go, Python, or Bash • Solid understanding of Linux/Unix systems, DNS, TLS/SSL, HTTP, load balancers, and distributed systems • Experience working with API gateway and service mesh technologies • Familiarity with Kafka and observability platforms such as Datadog, Prometheus, and Grafana • Experience working in a 24/7/365 production support environment • Legally authorized to work in the country where this position will be worked

🏖️ Benefits

• Health, dental, and vision insurance • Paid holidays and PTO • Professional development opportunities and tools

Apply Now

Similar Jobs

🕒 August 25

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

AWS DevSecOps Engineer modernizing CACI’s federal Grant Solutions cloud platform. Automating secure infrastructure, CI/CD, observability, and compliant software delivery.

🇺🇸 United States – Remote

💵 $98.5k - $206.8k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 25

Bitwarden

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior SRE operating Bitwarden Gov’s FedRAMP-compliant cloud infrastructure. Managing reliability, monitoring, incident response, Kubernetes, and security across multi-cloud environments.

🕒 August 25

URUS Group

1001 - 5000

🌾 Agriculture

🤝 B2B

🧬 Biotechnology

DevOps Team Lead operating VAS’s AWS platform for farm management software. Leading IaC, reliability, security, cost optimization, and globally distributed workloads.

🕒 August 25

System Automation Corporation

51 - 200

💼 Consulting

🏥 Healthcare

⚖️ Legal

Site Reliability Engineer operating Azure infrastructure for System Automation’s regulatory-agency SaaS platform. Automating reliability, observability, CI/CD, security, and incident response.

🕒 August 25

Blue River Technology

201 - 500

🌾 Agriculture

🤖 Artificial Intelligence

🔧 Hardware

Senior Site Reliability Engineer scaling Kubernetes platforms and cloud infrastructure. Supporting Blue River Technology’s autonomous robotics products through reliability, security, and observability.