Site Reliability Engineer – Engineering Productivity, DevOps

🔥 3 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Build, deploy safely and incrementally and operate critical production systems with focus on scalability, reliability, observability, performance and security. • Monitor, support and enhance developer experience across services. • Build automation to remove toil and efficiently operate production systems. • Proactively monitor, respond to, and enhance alerts and set up automated alert handling. • Create and maintain the incident response runbooks. • Triage platform/infrastructural issues and help Arista software engineers in their triages. • Engage with 3rd party vendor support. • Write postmortem documents and build solutions to avoid incidents from repeating. • Plan and communicate maintenance windows on production systems. • Work with Arista’s product development teams to identify infrastructural issues that are causing bottlenecks and limitations in their workflows. • Design and implement solutions to resolve them. • Survey and adopt best practices around infrastructure/platform to maintain secure, scalable and fault-tolerant systems. • Study the design and sufficient implementation details of OSS systems for better triage and fix resolution.

🎯 Requirements

• At least BSc Computer Science or Engineering + 3 years’ experience, MS Computer Science or Engineering + 3 years’ experience, or equivalent work experience. • Knowledge of one or more of Go, Python, shell scripting to be able to implement medium complexity automation workflows. • Knowledge of Linux (or UNIX) from administration and debugging perspective. • Hands-on experience in operating software systems (infrastructure, complex applications etc) at scale. • Experience in server provisioning (esp from storage and networking perspective). • Strong problem solving and software troubleshooting skills. • Experience with infrastructure-as-code

🏖️ Benefits

• Health insurance • Retirement plans • Paid time off • Flexible work arrangements • Professional development

Apply Now

Similar Jobs

🔥 14 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

DevOps Engineer operating systems gathering telemetry from the next-gen cloud computing platform. Collaborating across scrum teams for design, development, and maintenance.

Cloud

Distributed Systems

Docker

Hadoop

Java

Kafka

Kubernetes

Linux

Python

Shell Scripting

Spark

TCP/IP

Terraform

🔥 16 hours ago

Valtech

5001 - 10000

🤝 B2B

☁️ SaaS

Site Reliability Engineer connecting software development and operations at Valtech. Delivering reliable speed and infrastructure to enhance customer experience.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

Kafka

Kubernetes

Microservices

Spring Boot

SpringBoot

🕒 Yesterday

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

Join the DevOps team at Fundraise Up, responsible for platform reliability and scalability. Drive technical initiatives and support developers by managing CI/CD and observability tasks.

🗣️🇷🇺 Russian Required

Ansible

Docker

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

🕒 Yesterday

CloudLinux

51 - 200

☁️ SaaS

🔐 Security

🌐 Web 3

DevOps Engineer working with CloudLinux to stabilize and modernize backend infrastructure for Imunify360 team. Collaborate on performance and reliability improvements in a fully remote environment.

Ansible

Apache

Django

Docker

Kafka

Kubernetes

NoSQL

Postgres

Puppet

RabbitMQ

Redis

SaltStack

SQL

🕒 2 days ago

The Codest

51 - 200

💳 Fintech

🛍️ eCommerce

DevOps Engineer developing AI-powered applications and supporting cloud engineering at Codest. Collaborating on microservices and ensuring software best practices in a remote environment.

AWS

Jenkins