Site Reliability Engineer – Engineering Productivity

Job not on LinkedIn

🕒 August 7

🇮🇪 Ireland – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Build, deploy safely and incrementally, and operate critical production systems • Focus on scalability, reliability, observability, performance, and security • Build automation to remove toil and proactively monitor, respond to, and enhance alerts • Create and maintain incident response runbooks • Triage platform and infrastructure issues • Write postmortem documents to prevent recurring incidents • Plan and communicate maintenance windows on production systems • Engage with third-party vendor support as needed • Identify infrastructure bottlenecks with product development teams • Design solutions to enhance developer experience and workflow efficiency • Survey and adopt infrastructure and platform design best practices • Study open-source system implementations to improve triage and resolution

🎯 Requirements

• At least a BSc in Computer Science or Engineering plus 3 years’ experience, an MS in Computer Science or Engineering plus 3 years’ experience, or equivalent work experience • Knowledge of Go, Python, or shell scripting for implementing medium-complexity automation workflows • Knowledge of Linux or UNIX administration and debugging • Hands-on experience operating software systems at scale • Experience with server provisioning, especially storage and networking • Strong problem-solving and software troubleshooting skills • Experience with infrastructure-as-code • Experience managing databases such as MariaDB, PostgreSQL, or MongoDB (desired) • Experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers (desired) • Experience managing monitoring stacks such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos (desired) • Experience managing Elasticsearch clusters (desired) • Experience managing Artifactory and Docker registries (desired) • Experience managing CI/CD systems such as ArgoCD or Spinnaker (desired) • Experience managing version-control systems such as Perforce or Gerrit (desired) • Experience with infrastructure-as-code frameworks such as Ansible (desired) • Experience managing large Java applications (desired) • Experience managing storage infrastructure such as NAS, SAN, or Ceph (desired)

🏖️ Benefits

• Employees can work remotely • Work-life balance • Inclusive environment • Diversity-focused workplace

Apply Now

Similar Jobs

🕒 August 5

Twilio

5001 - 10000

🔌 API

🤝 B2B

DevOps Engineer shaping Twilio’s OpenTelemetry-first observability platform. Architecting scalable telemetry systems, developer tooling, and APIs for reliable, cost-effective operations.

AWS

Cloud

Distributed Systems

Grafana

Java

Kafka

Kubernetes

Prometheus

Python

Go

🕒 August 4

Red Hat

10,000+ employees

🏢 Enterprise

Senior Site Reliability Engineer modernizing Red Hat's internal enterprise open-source applications and infrastructure. Automating, securing, and improving highly available cloud and container platforms.

Ansible

Cloud

Docker

Kubernetes

Linux

OpenShift

Prometheus

Python

SDLC

Terraform

🕒 July 28

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

IT Infrastructure Support Engineer optimizing critical physical security systems for a global IT provider. Engineering automation tools for a reliable and scalable infrastructure environment.

Ansible

Chef

Cloud

Grafana

IoT

Kubernetes

Linux

Prometheus

Puppet

Python

Terraform

Go

🕒 July 27

Sardine

51 - 200

🔒 Cybersecurity

📋 Compliance

💳 Fintech

DevOps Engineer at Sardine improving infrastructure and tooling for a remote-first financial crime platform. Collaborating to ensure reliable, scalable, and cost-efficient systems.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Prometheus

Python

Terraform

Go

🕒 July 27

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform