Senior Site Reliability Engineer – CloudVision

🔥 7 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Design, build, and deploy production systems with a focus on scalability, reliability, observability, and performance, ensuring systems meet stringent security standards • Develop and maintain comprehensive automation solutions to eliminate toil and streamline operational efficiency across production environments • Proactively monitor production systems, establish intelligent alerting strategies, and implement automated incident response mechanisms to minimise downtime • Create and maintain detailed incident response runbooks; conduct thorough postmortem analyses following incidents to identify root causes and prevent recurrence • Collaborate with software engineering teams to identify and resolve infrastructural bottlenecks, designing innovative solutions that enhance product deployment workflows • Manage and optimise monitoring infrastructure using industry-standard tools, ensuring comprehensive visibility across all systems • Plan, communicate, and execute maintenance windows on production systems with minimal disruption to service availability • Triage platform and infrastructural issues with decisiveness and analytical rigour; engage with third-party vendors and support teams as required • Deploy new systems and updates in a staged, risk-managed manner, ensuring safe and incremental rollouts • Survey and adopt best practices in infrastructure and platform management to maintain secure, scalable, and fault-tolerant systems • Study the design and implementation details of open-source systems to enhance troubleshooting capabilities and accelerate issue resolution • Work transparently with stakeholders to communicate system status, planned maintenance, and infrastructure improvements

🎯 Requirements

• Bachelor's degree in Computer Science, Engineering, or equivalent professional experience (5+ years in a related infrastructure or systems role) • Proficiency in one or more programming languages: Go, Python, or bash shell scripting, with the ability to implement medium-complexity automation workflows • Strong knowledge of Linux or UNIX from both administration and debugging perspectives • Hands-on experience operating software systems, infrastructure, and complex applications at scale in production environments • Demonstrated expertise in infrastructure-as-code principles and practices • Strong problem-solving and software troubleshooting skills with a methodical, analytical approach • Experience with server provisioning, particularly from storage and networking perspectives • Proven ability to work collaboratively within cross-functional teams and communicate technical concepts clearly • Experience with incident response, postmortem analysis, and continuous improvement methodologies • Experience with container orchestration platforms, particularly Kubernetes • Hands-on experience with Docker and virtualisation technologies • Proficiency in managing monitoring stacks, including Prometheus and Grafana • Experience with CI/CD systems such as GitLab tools or Spinnaker • Knowledge of infrastructure-as-code frameworks, particularly Terraform • Experience managing databases such as PostgreSQL or equivalent relational database management systems • Experience with artifact repositories and Docker registries • Familiarity with cloud platforms (Google Cloud Platform, Amazon Web Services, or Microsoft Azure) • Understanding of distributed systems architecture and principles • Experience with performance tuning and system optimisation • Knowledge of security best practices in infrastructure and systems design • On-call support experience and comfort with incident response responsibilities

🏖️ Benefits

• Professional development opportunities • Remote work options

Apply Now

Similar Jobs

🕒 5 days ago

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

IT Infrastructure Support Engineer optimizing critical physical security systems for a global IT provider. Engineering automation tools for a reliable and scalable infrastructure environment.

Ansible

Chef

Cloud

Grafana

IoT

Kubernetes

Linux

Prometheus

Puppet

Python

Terraform

Go

🕒 6 days ago

Sardine

51 - 200

🔒 Cybersecurity

📋 Compliance

💳 Fintech

DevOps Engineer at Sardine improving infrastructure and tooling for a remote-first financial crime platform. Collaborating to ensure reliable, scalable, and cost-efficient systems.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Prometheus

Python

Terraform

Go

🕒 6 days ago

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 June 27

Castillians

51 - 200

DevOps Engineer working with Castille Resources on security measures and tooling for client portal. Engaged in creating security policies, threat assessments, and team training.

🕒 June 25

CloudBees

501 - 1000

🤝 B2B

Strategic DevSecOps Consultant at CloudBees helping enterprises transform software delivery through cloud and DevSecOps strategies. Collaborating with clients to enhance operational efficiency and developer experience.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

Jenkins

Kubernetes

Python

Go