Senior Site Reliability Engineer – CloudVision

🕒 August 4

🇮🇪 Ireland – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Design, build, and deploy production systems focused on scalability, reliability, observability, performance, and security • Develop and maintain automation solutions to eliminate toil and improve operational efficiency • Monitor production systems, establish alerting strategies, and implement automated incident response • Create incident response runbooks and conduct postmortem analyses • Collaborate with software engineering teams to resolve infrastructure bottlenecks and improve deployment workflows • Manage and optimise monitoring infrastructure for comprehensive system visibility • Plan, communicate, and execute production maintenance windows with minimal service disruption • Triage platform and infrastructure issues and engage vendors and support teams as needed • Deploy systems and updates through staged, risk-managed rollouts • Research and adopt infrastructure and platform management best practices • Study open-source system design and implementation details to improve troubleshooting • Communicate system status, maintenance plans, and infrastructure improvements to stakeholders

🎯 Requirements

• Bachelor's degree in Computer Science, Engineering, or equivalent professional experience • 5+ years in a related infrastructure or systems role • Proficiency in Go, Python, or bash shell scripting • Ability to implement medium-complexity automation workflows • Strong Linux or UNIX administration and debugging knowledge • Hands-on experience operating software systems, infrastructure, and complex applications at production scale • Expertise in infrastructure-as-code principles and practices • Strong problem-solving and software troubleshooting skills • Experience with server provisioning, storage, and networking • Ability to collaborate cross-functionally and communicate technical concepts clearly • Experience with incident response, postmortem analysis, and continuous improvement • Experience with Kubernetes, Docker, and virtualisation technologies desirable • Proficiency with Prometheus and Grafana desirable • Experience with GitLab tools or Spinnaker desirable • Knowledge of Terraform desirable • Experience with PostgreSQL or equivalent relational databases desirable • Experience with artifact repositories and Docker registries desirable • Familiarity with Google Cloud Platform, Amazon Web Services, or Microsoft Azure desirable • Understanding of distributed systems architecture desirable • Performance tuning and system optimisation experience desirable • Knowledge of infrastructure and systems security best practices desirable • On-call support and incident response experience desirable

🏖️ Benefits

• Remote work from Ireland • Permanent employment • Work with cross-functional teams and access to various company domains • Complete ownership of projects • Flat and streamlined management structure • Opportunities to work across various domains • Access to every part of the company • Test automation tools and engineering-focused culture • Inclusive environment valuing diversity of thought and perspectives

Apply Now

Similar Jobs

🕒 July 28

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

Site Reliability Engineer automating servers, networks, cameras, access control, and Cisco infrastructure. Building IaC, monitoring, CMDB, and incident-remediation systems for enterprise reliability.

Ansible

Chef

Cloud

Grafana

IoT

Kubernetes

Linux

Prometheus

Puppet

Python

Terraform

Go

🕒 July 27

Sardine

51 - 200

🔒 Cybersecurity

📋 Compliance

💳 Fintech

DevOps Engineer at Sardine improving infrastructure and tooling for a remote-first financial crime platform. Collaborating to ensure reliable, scalable, and cost-efficient systems.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Prometheus

Python

Terraform

Go

🕒 July 27

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 June 27

Castillians

51 - 200

DevOps Engineer working with Castille Resources on security measures and tooling for client portal. Engaged in creating security policies, threat assessments, and team training.

🕒 June 25

CloudBees

501 - 1000

🤝 B2B

Strategic DevSecOps Consultant at CloudBees helping enterprises transform software delivery through cloud and DevSecOps strategies. Collaborating with clients to enhance operational efficiency and developer experience.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

Jenkins

Kubernetes

Python

Go