Senior Site Reliability Engineer

🕒 June 23

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of SigNoz

SigNoz

11 - 50 employees

☁️ SaaS

🏢 Enterprise

Software • SaaS • Enterprise

SigNoz is an open-source observability tool designed as an alternative to Datadog and New Relic. It provides a comprehensive platform that integrates application performance monitoring (APM), logs, metrics, traces, exceptions, and alerts in a single tool. SigNoz supports various deployment options including self-hosting and cloud services, and it's built to handle extensive data ingestion, making it suitable for teams of all sizes to monitor, troubleshoot, and enhance application performance efficiently.

📋 Description

• Own the reliability, scalability, and operability of the SigNoz cloud platform • Keep a petabyte-scale observability system fast and dependable • Scale the ingest path — making it robust to bursts while maintaining data freshness • Operate and tune ClickHouse and the data layer for performance and cost • Manage Kubernetes infrastructure: cluster operations, upgrades, multi-tenancy • Help make the observability of SigNoz itself world-class • Work with a high-caliber team across various responsibilities including SLOs/SLIs, incident response, and tooling

🎯 Requirements

• 5–8 years in SRE, infrastructure, or platform/backend roles operating production systems at scale • Deep, practical Kubernetes experience • Strong grasp of distributed systems failure modes, performance debugging, and capacity planning • Comfortable in code (Go preferred) • Loves open source — ideally with prior contributions to OSS projects • Comfortable in a high-ownership, fast-moving, remote-first environment • Strong communication — can write clear runbooks and tech docs and explain trade-offs

🏖️ Benefits

• Remote-first, async-friendly culture

Apply Now

Similar Jobs

🕒 June 19

BETSOL

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Cloud Engineer at BETSOL building and operating cloud portal workloads across Azure and GCP. Focused on DevOps and DevSecOps with AI-first development practices.

Ansible

Azure

Cloud

Google Cloud Platform

Grafana

JavaScript

Jenkins

Kubernetes

Prometheus

Python

Terraform

TypeScript

Vault

🕒 June 15

Jumio Corporation

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

DevOps Engineer IV managing and administering web services in public cloud environments. Collaborating with global teams and mentoring junior engineers at Jumio.

AWS

Cloud

Distributed Systems

Docker

Grafana

Kubernetes

Packer

Prometheus

Python

Terraform

🕒 June 12

Kayzen

51 - 200

📣 Marketing

☁️ SaaS

Lead DevOps Engineer building and scaling the backbone of Kayzen's RTB platform. Join a diverse team impacting Adtech solutions globally.

Ansible

Chef

Grafana

HAProxy

Java

Jenkins

Kubernetes

Linux

NGINX

Prometheus

Puppet

Python

SQL

TCP/IP

Terraform

Unix

🕒 June 9

Thinkahead Consultant Psychologist Pty Ltd

1 - 10

🏥 Healthcare

💼 Consulting

📚 Education

DevSecOps Engineer working with AWS to create secure environments and optimize developer tooling. Collaborating with engineering teams to standardize SDLC and enhance security measures.

AWS

Cloud

Python

SDLC

Terraform

🕒 June 1

OpenAI

201 - 500

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.

AWS