Senior Site Reliability Engineer

🔥 2 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of JumpCloud

JumpCloud

201 - 500 employees

Founded 2013

☁️ SaaS

🔐 Security

🏢 Enterprise

💰 $66M Series F on 2021-10

SaaS • Security • Enterprise

JumpCloud is a unified identity, device, and access management platform that helps organizations centrally manage and secure their IT infrastructure across multiple operating systems and devices. It provides comprehensive solutions for cross-platform device management, cloud-first directory services, Active Directory modernization, and hybrid work enablement. JumpCloud enhances security through features such as zero trust security, passwordless authentication, and multi-factor authentication. Its platform supports automated onboarding and offboarding, identity lifecycle management, conditional access, and compliance management. JumpCloud integrates with various HR systems and offers SaaS management, allowing companies to streamline IT operations and reduce complexity. Trusted by organizations worldwide, JumpCloud is recognized for its ability to unify and secure digital workplace environments efficiently and effectively.

📋 Description

• Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure on AWS/GCP • Architect, build, and maintain Disaster Recovery processes, multi-region failover automation, and business continuity strategies • Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks • Drive end-to-end observability strategy using Datadog and Golden Signals monitoring • Lead on-call escalation and major incident management while driving adherence to 99.99% availability SLAs • Facilitate blameless post-incident reviews and implement systemic root-cause remediations • Architect, manage, and scale production Kubernetes EKS clusters with GitOps workflows using Argo CD and Kargo • Design and maintain Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments • Build FinOps and cost-optimization dashboards for multi-cloud spend, unit economics, and resource utilization • Write production-grade Python or Go tooling, platform automation, and custom integrations • Champion AI-assisted software development workflows using Cursor, Claude Code, and GitHub Copilot • Author operational runbooks and architecture decision records • Mentor mid-level and junior engineers

🎯 Requirements

• 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems • Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline • Advanced Python or Go capabilities • Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling such as Argo CD • Deep Terraform proficiency across complex multi-account AWS environments, including IAM, VPCs, Transit Gateway, ALB/NLB, and Route53 • Experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and FinOps dashboards • Experience designing and testing multi-region Disaster Recovery architectures and automating failover systems • Experience defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms • Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy systems such as HAProxy or NGINX • Ability to lead technical discussions, write architectural design documents/RFCs, and mentor engineering peers • Strong problem-solving, communication, and collaboration skills • Preferred: basic understanding of chaos engineering; secrets management architectures; DevSecOps practices; identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions • Must be located in and authorized to work in India • Fluent spoken and written English required

🏖️ Benefits

• Remote work within India • Opportunity to work with talent across 15+ countries • Professional growth and expertise-sharing opportunities • Supportive, collaborative work environment • Equal opportunity employment

Apply Now

Similar Jobs

🕒 September 24

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

Site Reliability Engineer modernizing Granicus’s government technology platforms through AIOps, observability, and automation. Improving reliability, incident response, and resilient cloud operations.

Ansible

AWS

Azure

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

ITSM

Kubernetes

Linux

Logstash

Terraform

Unix

🕒 September 23

MFSG

11 - 50

🏭 Manufacturing

🔧 Hardware

🚗 Transport

Site Reliability Engineer automating reliable, compliant digital banking platforms for MFSG Technologies. Managing CI/CD, observability, incident response, and resilient production deployments.

Ansible

AWS

Azure

Cloud

Docker

Kubernetes

Python

Terraform

🕒 September 23

iCert Global

51 - 200

💼 Consulting

📣 Marketing

📚 Education

Lead SRE managing Azure infrastructure, AKS, observability, and major incidents for Icertis’s AI-powered contract intelligence platform. Driving automation, reliability, and cloud-native operations.

AWS

Azure

Cloud

Distributed Systems

Docker

Kubernetes

Python

ServiceNow

Terraform

🕒 September 21

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Senior DevOps Engineer migrating 2,000+ GitHub repositories and CI/CD ecosystems for global IT services company Miratech. Automating enterprise platforms across GitHub, Terraform, AWS/EKS, Kubernetes, and Jenkins.

AWS

Cloud

Jenkins

Kubernetes

Python

Terraform

🕒 September 21

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Senior Observability/DevOps Engineer migrating enterprise Datadog environments for Miratech, a global IT services and consulting company. Automating resilient observability platforms with Datadog, Terraform, AWS, and Kubernetes.

AWS

Kubernetes

ServiceNow

Terraform