Site Reliability Engineer

Job not on LinkedIn

🔥 2 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of JumpCloud

JumpCloud

201 - 500 employees

Founded 2013

☁️ SaaS

🔐 Security

🏢 Enterprise

💰 $66M Series F on 2021-10

SaaS • Security • Enterprise

JumpCloud is a unified identity, device, and access management platform that helps organizations centrally manage and secure their IT infrastructure across multiple operating systems and devices. It provides comprehensive solutions for cross-platform device management, cloud-first directory services, Active Directory modernization, and hybrid work enablement. JumpCloud enhances security through features such as zero trust security, passwordless authentication, and multi-factor authentication. Its platform supports automated onboarding and offboarding, identity lifecycle management, conditional access, and compliance management. JumpCloud integrates with various HR systems and offers SaaS management, allowing companies to streamline IT operations and reduce complexity. Trusted by organizations worldwide, JumpCloud is recognized for its ability to unify and secure digital workplace environments efficiently and effectively.

📋 Description

• Design, deploy, and maintain reliability, availability, and performance for critical JumpCloud systems and APIs across AWS and GCP • Operationalize SLIs, SLOs, and error budgets with core application teams • Build and refine end-to-end observability across microservices and cloud infrastructure using tools such as Datadog • Implement monitoring based on Golden Signals: latency, traffic, errors, and saturation • Participate in on-call rotations, incident response, and blameless post-incident reviews • Manage production Kubernetes EKS clusters using GitOps workflows such as Argo CD and Kargo • Provision and secure multi-cloud infrastructure using modular Terraform • Develop and maintain disaster recovery dashboards, runbooks, multi-region failover automation, and validation tests aligned with RTO/RPO targets • Write production-grade Python or Go scripts and automation tools to eliminate operational toil • Use AI-assisted development tools such as Cursor, Claude Code, and GitHub Copilot for scripting, runbook generation, and incident triage

🎯 Requirements

• 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems • Proficiency in Python or Go for SRE tools, custom automation, and cloud integrations • Production experience with Kubernetes, container orchestration, and GitOps pipelines such as Argo CD • Experience writing, maintaining, and modularizing Terraform configurations • Experience operating AWS workloads, including EKS, IAM, VPC networking, Route53, and ALB/NLB, or GCP workloads • Practical experience with FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards • Experience building disaster recovery dashboards, running failover drills, and configuring monitoring for system health and recovery metrics • Practical experience with Datadog or similar, PagerDuty, alerting hygiene, and SLI/SLO frameworks • Operational experience configuring and troubleshooting production service meshes such as Istio and high-availability proxy solutions such as HAProxy or NGINX • Strong troubleshooting skills and track record of improving operational efficiency through code • Strong team-player orientation and alignment with company core values • Willingness and ability to participate in on-call shifts • Fluent spoken and written English • Preferred: experience with GitHub Actions or GitLab Pipelines • Preferred: basic understanding of chaos engineering or resilience testing • Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator • Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation

🏖️ Benefits

• Remote-first work within India • Opportunity to work in a fast, SaaS-based environment • Professional growth and expertise-sharing opportunities • Collaboration with talented global teams • Employee voice in product and feature development • Supportive executive team and board • Equal opportunity employment • No third-party resumes accepted

Apply Now

Similar Jobs

🔥 17 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer II maintaining observability, Kubernetes, and cloud infrastructure for Akamai’s distributed security platform. Improving reliability, monitoring, and automated remediation.

AWS

Azure

Docker

Kafka

Kubernetes

Linux

Python

SQL

🕒 September 24

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

Site Reliability Engineer modernizing Granicus’s government technology platforms through AIOps, observability, and automation. Improving reliability, incident response, and resilient cloud operations.

Ansible

AWS

Azure

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

ITSM

Kubernetes

Linux

Logstash

Terraform

Unix

🕒 September 23

MFSG

11 - 50

🏭 Manufacturing

🔧 Hardware

🚗 Transport

Site Reliability Engineer automating reliable, compliant digital banking platforms for MFSG Technologies. Managing CI/CD, observability, incident response, and resilient production deployments.

Ansible

AWS

Azure

Cloud

Docker

Kubernetes

Python

Terraform

🕒 September 23

iCert Global

51 - 200

💼 Consulting

📣 Marketing

📚 Education

Lead SRE managing Azure infrastructure, AKS, observability, and major incidents for Icertis’s AI-powered contract intelligence platform. Driving automation, reliability, and cloud-native operations.

AWS

Azure

Cloud

Distributed Systems

Docker

Kubernetes

Python

ServiceNow

Terraform

🕒 September 21

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Senior Observability/DevOps Engineer migrating enterprise Datadog environments for Miratech, a global IT services and consulting company. Automating resilient observability platforms with Datadog, Terraform, AWS, and Kubernetes.

AWS

Kubernetes

ServiceNow

Terraform