Site Reliability Engineer – Engineering Productivity

🕒 August 7

🇮🇪 Ireland – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Build, deploy safely and incrementally, and operate critical production systems • Focus on scalability, reliability, observability, performance, and security • Build automation to remove toil and proactively monitor, respond to, and enhance alerts • Create and maintain incident response runbooks • Triage platform and infrastructure issues • Write postmortem documents to prevent recurring incidents • Plan and communicate maintenance windows on production systems • Engage with third-party vendor support as needed • Identify infrastructure bottlenecks with product development teams • Design solutions to enhance developer experience and workflow efficiency • Survey and adopt infrastructure and platform design best practices • Study open-source system implementations to improve triage and resolution

🎯 Requirements

• At least a BSc in Computer Science or Engineering plus 3 years’ experience, an MS in Computer Science or Engineering plus 3 years’ experience, or equivalent work experience • Knowledge of Go, Python, or shell scripting for implementing medium-complexity automation workflows • Knowledge of Linux or UNIX administration and debugging • Hands-on experience operating software systems at scale • Experience with server provisioning, especially storage and networking • Strong problem-solving and software troubleshooting skills • Experience with infrastructure-as-code • Experience managing databases such as MariaDB, PostgreSQL, or MongoDB (desired) • Experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers (desired) • Experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos (desired) • Experience managing Elasticsearch clusters (desired) • Experience managing Artifactory and Docker registries (desired) • Experience managing CI/CD systems such as ArgoCD or Spinnaker (desired) • Experience managing version-control systems such as Perforce or Gerrit (desired) • Experience with infrastructure-as-code frameworks such as Ansible (desired) • Experience managing large Java applications (desired) • Experience managing storage infrastructure such as NAS, SAN, or Ceph (desired)

🏖️ Benefits

• Employees can work remotely • Work-life balance • Inclusive environment • Diversity-focused workplace

Apply Now

Similar Jobs

🕒 August 5

Twilio

5001 - 10000

🔌 API

🤝 B2B

DevOps Engineer shaping Twilio’s OpenTelemetry-first observability platform. Architecting scalable telemetry systems, developer tooling, and APIs for reliable, cost-effective operations.

AWS

Cloud

Distributed Systems

Grafana

Java

Kafka

Kubernetes

Prometheus

Python

Go

🕒 July 27

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 June 27

Castillians

51 - 200

DevOps Engineer working with Castille Resources on security measures and tooling for client portal. Engaged in creating security policies, threat assessments, and team training.

🕒 June 25

CloudBees

501 - 1000

🤝 B2B

Strategic DevSecOps Consultant at CloudBees helping enterprises transform software delivery through cloud and DevSecOps strategies. Collaborating with clients to enhance operational efficiency and developer experience.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

Jenkins

Kubernetes

Python

Go

🕒 April 1

Lacroix Healthcare Consulting, LLC

1 - 10

💼 Consulting

🏥 Healthcare

⚕️ Healthcare Insurance

Senior Site Reliability Engineer focusing on cloud service deployment and maintenance within Grouper’s Platform Operations Team. Collaborating across departments to ensure data integrity and operational efficiency.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Linux

Spark

SQL

Terraform