Senior Software Engineer – Reliability

Job not on LinkedIn

🔥 1 hour ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Atlan Stormwater

Atlan Stormwater

51 - 200 employees

Founded 1972

🏭 Manufacturing

📦 Logistics

💼 Consulting

Manufacturing • Logistics • Consulting

Atlan Stormwater is a company dedicated to providing comprehensive stormwater management solutions, focusing on both water quantity and quality. Established in 1972, Atlan manufactures a range of products including stormwater detention and retention systems, hydrocarbon capture technologies, and gross pollutant traps. The company emphasizes sustainable practices by integrating green infrastructure into urban environments to enhance water filtration, mitigate flooding, and promote clean waterways for communities. Atlan also offers design assistance and maintenance services to ensure optimal performance of their systems over time.

📋 Description

• Improve investigation-agent root-cause accuracy so engineers can trust diagnoses without re-checking • Expand auto-remediation from a handful of playbooks to dozens using staged autonomy • Build and run fault-injection benchmarks and evaluation harnesses • Set evaluation pass marks, report results with honest denominators, and validate the harnesses • Design fail-closed checks, kill switches, approval flows, and blast-radius limits for production actions • Build new agents that remove categories of operational toil • Develop clean interfaces, guardrails, and safe defaults for other teams using the reliability platform • Stop production incidents from worsening and permanently remove recurring failure classes • Contribute to Atlan's mature, opinionated codebase and improve its documented architecture and eval-gated pull requests • Collaborate across teams to drive adoption of the reliability platform

🎯 Requirements

• Experience carrying a pager, owning incidents end to end, or working a support or escalation queue • Demonstrated ability to eliminate recurring operational problems rather than merely optimize runbooks • Experience measuring outcomes using adoption metrics, real numbers, and honest denominators • Experience rebuilding a core part of one's work with AI and shipping an AI-native workflow used by others • Experience architecting agents that take autonomous action, including defining guardrails • Understanding of false-positive rates, rollback paths, and blast radius • Experience building platforms for other teams, including clean interfaces, documented failure modes, and safe defaults • Ability to contribute quickly to a high-rigor existing codebase and improve it without rewriting it • Ability to write deterministic code for routing, filtering, and safety before generative steps • Genuine interest in reliability and operational efficiency • Availability for overlap between roughly 11am and 8pm IST

🏖️ Benefits

• Strong base salary • Performance-based variable pay • Impact-driven equity for most roles • Health, dental, vision, and mental health benefits from Day 1 • Flexible health stipends • Flexible time off • Modern leave policies • Accelerated growth and learning opportunities • Global, remote-first, high-trust work environment • Work from anywhere with a diverse team across 15+ countries • Async work environment • Flexibility and ownership over how you work

Apply Now

Similar Jobs

🔥 20 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Site Reliability Engineer automating Linux infrastructure, monitoring, and deployments. Supporting Akamai’s globally distributed cloud and edge platform for reliable, secure digital experiences.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

SaltStack

Splunk

Terraform

🕒 6 days ago

JumpCloud

201 - 500

☁️ SaaS

🔐 Security

🏢 Enterprise

Senior SRE architecting multi-region cloud infrastructure, Kubernetes, disaster recovery, and observability for JumpCloud’s AI-powered IT management platform. Driving reliability, FinOps, automation, and incident management.

AWS

Cloud

Distributed Systems

Google Cloud Platform

HAProxy

Kubernetes

Microservices

NGINX

Python

Terraform

Go

🕒 6 days ago

JumpCloud

201 - 500

☁️ SaaS

🔐 Security

🏢 Enterprise

Site Reliability Engineer strengthening AWS/GCP reliability, observability, Kubernetes, and disaster recovery. Supporting JumpCloud’s AI-powered unified IT management platform through automation and incident response.

AWS

Cloud

Google Cloud Platform

HAProxy

Kubernetes

Microservices

NGINX

Python

Terraform

Vault

Go

🕒 September 23

MFSG

11 - 50

🏭 Manufacturing

🔧 Hardware

🚗 Transport

Site Reliability Engineer automating reliable, compliant digital banking platforms for MFSG Technologies. Managing CI/CD, observability, incident response, and resilient production deployments.

Ansible

AWS

Azure

Cloud

Docker

Kubernetes

Python

Terraform

🕒 September 23

iCert Global

51 - 200

💼 Consulting

📣 Marketing

📚 Education

Lead SRE managing Azure infrastructure, AKS, observability, and major incidents for Icertis’s AI-powered contract intelligence platform. Driving automation, reliability, and cloud-native operations.

AWS

Azure

Cloud

Distributed Systems

Docker

Kubernetes

Python

ServiceNow

Terraform