Site Reliability Engineer – Engineering Productivity

🔥 18 hours ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Build, safely and incrementally deploy, and operate critical production systems • Focus on scalability, reliability, observability, performance, and security • Monitor, support, and enhance developer experience across services • Build automation to remove toil and operate production systems efficiently • Monitor and respond to alerts, enhance alerting, and establish automated alert handling • Create and maintain incident-response runbooks • Build and deploy new systems with scalability, reliability, and observability as primary requirements • Triage platform and infrastructure issues and assist Arista software engineers with triage • Engage with third-party vendor support • Deploy new systems in a staged manner • Write postmortems and develop solutions to prevent recurring incidents • Plan and communicate production-system maintenance windows • Identify infrastructure issues causing workflow bottlenecks and limitations for product-development teams • Design and implement solutions to resolve infrastructure issues • Survey and adopt infrastructure and platform best practices • Implement fault tolerance and performance improvements to increase system availability • Study open-source system designs and implementation details to improve triage and fix resolution

🎯 Requirements

• At least BSc Computer Science or Engineering + 5 years’ experience, MS Computer Science or Engineering + 5 years’ experience, or equivalent work experience • Knowledge of one or more of Go, Python, or shell scripting to implement medium-complexity automation workflows • Knowledge of Linux or UNIX from an administration and debugging perspective • Hands-on experience operating software systems at scale • Experience in server provisioning, especially from storage and networking perspectives • Strong problem-solving and software troubleshooting skills • Experience with infrastructure-as-code • Experience managing databases such as MariaDB, PostgreSQL, or MongoDB • Experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers • Experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos • Experience managing Elasticsearch clusters • Experience managing Artifactory or Docker registries • Experience managing CI/CD systems such as ArgoCD or Spinnaker • Experience managing version-control systems such as Perforce or Gerrit • Experience with infrastructure-as-code frameworks such as Ansible • Experience managing large Java applications • Experience in storage infrastructure management such as NAS, SAN, or Ceph

🏖️ Benefits

• Great Place to Work awards for engineering, diversity, compensation, and work-life balance • Engineers have complete ownership of their projects • Flat and streamlined management structure • Opportunities to work across various domains • Access to every part of the company • Inclusive environment valuing diversity of thought and perspectives

Apply Now

Similar Jobs

🔥 21 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Site Reliability Engineer maintaining Akamai's Compute services and infrastructure. Improving reliability, automation, observability, and incident response for the distributed cloud and edge platform.

Ansible

Cloud

Docker

Grafana

HAProxy

Jenkins

Linux

NGINX

Prometheus

Python

Redis

SaltStack

Terraform

Go

🕒 Yesterday

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Platform DevOps Engineer managing AWS infrastructure, Kubernetes, and CI/CD pipelines for Miratech’s global IT services. Automating reliable cloud operations with Python, Terraform, Ansible, and DevOps tooling.

Ansible

AWS

Azure

Cloud

Kubernetes

Python

Terraform

.NET

🕒 Yesterday

Fortive

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

📦 Logistics

Site Reliability Engineer maintaining highly available Windows/Linux infrastructure for a customer-facing platform. Automating deployments, observability, CI/CD, and cloud operations with AWS and Azure.

Ansible

ASP.NET

AWS

Azure

Cloud

Docker

Jenkins

Kubernetes

Linux

Python

SQL

Terraform

🕒 Yesterday

Playpower Labs

11 - 50

📚 Education

🤖 Artificial Intelligence

DevOps Engineer running AWS infrastructure, CI/CD, security, and reliability for PlayPower Labs’ EdTech software. Supporting products used by millions of students and teachers.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Jenkins

Kubernetes

Linux

Python

Terraform

🕒 2 days ago

Hyland

1001 - 5000

🤝 B2B

☁️ SaaS

🏢 Enterprise

Senior DevOps Engineer maintaining cloud infrastructure, availability, and performance for Hyland’s enterprise content intelligence platform. Automating deployments and supporting Kubernetes-based Cloud Services.

Ansible

AWS

Cloud

Java

Jenkins

Kubernetes

Linux

Microservices

Packer

Python

Shell Scripting

Terraform