Site Reliability Engineer – Engineering Productivity

Job not on LinkedIn

🕒 May 12

🇵🇱 Poland – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 23%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Arista Networks

Arista Networks

1001 - 5000 employees

Founded 2004

🏢 Enterprise

📡 Telecommunications

💰 $2.6M Post-IPO Debt on 2015-05

Enterprise • Telecommunications • Cloud

Arista Networks is a leader in building scalable high-performance and ultra-low-latency networks for modern data center and cloud computing environments. The company offers a wide range of networking solutions including the Arista Extensible Operating System (EOS) for cloud networking, CloudVision for network automation and observability, and a variety of high-performance switches and routing platforms. Arista's solutions are designed for enterprise WANs, cloud-grade routing, multi-domain segmentation, security, and modern network operating models, making them a pivotal choice for hyperscale data centers and cloud environments.

📋 Description

• Build, deploy safely and incrementally and operate critical production systems with focus on scalability, reliability, observability, performance and security. • Monitor, support and enhance developer experience across services. • Build automation to remove toil and efficiently operate production systems. • Proactively monitor, respond to, and enhance alerts and set up automated alert handling • Create and maintain the incident response runbooks. • Triage platform/infrastructural issues and help Arista software engineers in their triages. • Engage with 3rd party vendor support. • Write postmortem documents and build solutions to avoid incidents from repeating. • Plan and communicate maintenance windows on production systems. • Work with Arista’s product development teams to identify infrastructural issues that are causing bottlenecks and limitations in their workflows. • Design and implement solutions to resolve them. • Survey and adopt best practices around infrastructure/platform to maintain secure, scalable and fault-tolerant systems. • Study the design and sufficient implementation details of OSS systems for better triage and fix resolution.

🎯 Requirements

• At least BSc Computer Science or Engineering + 3 years’ experience, MS Computer Science or Engineering + 3 years’ experience, or equivalent work experience. • Knowledge of one or more of Go, Python, shell scripting to be able to implement medium complexity automation workflows. • Knowledge of Linux (or UNIX) from administration and debugging perspective • Hands-on experience in operating software systems (infrastructure, complex applications etc) at scale • Experience in server provisioning (esp from storage and networking perspective). • Strong problem solving and software troubleshooting skills • Experience with infrastructure-as-code

🏖️ Benefits

• Health insurance • Flexible work arrangements • Professional development opportunities • Paid time off

Apply Now

Similar Jobs

🕒 May 11

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Senior Engineer creating solutions to improve automation and efficiency for Akamai's Compute products and services. Collaborating on deployment, monitoring, and resolving incidents with a focus on reliability and scalability.

Ansible

Distributed Systems

Grafana

Kubernetes

Linux

Prometheus

Python

SaltStack

Terraform

Unix

Go

🕒 April 30

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Senior Site Reliability Engineer designing and operating application deployment for Akamai Cloud. Collaborate with global teams to solve complex engineering challenges and enhance observability infrastructure.

Ansible

Chef

Distributed Systems

Puppet

SaltStack

Terraform

🕒 April 23

RedSky

11 - 50

💼 Consulting

🎖️ Defense

🏥 Healthcare

Venture Builder creating startups from the ground up at Red Sky. Join and build teams pushing boundaries across various industries.

🕒 April 22

CloudLinux

51 - 200

☁️ SaaS

🔐 Security

🌐 Web 3

Lead the evolution of CloudLinux's data platform into a DBaaS model. Design resilient databases and implement automated infrastructure management for high-performance systems.

Airflow

Ansible

Apache

Cloud

ETL

Kubernetes

MongoDB

Postgres

Python

Redis

SQL

Terraform

Zookeeper

Go

🕒 April 20

VirtusLab

201 - 500

💼 Consulting

🏭 Manufacturing

📦 Logistics

DevOps Engineer managing security operations for a rapidly scaling UK insurance leader. Responsible for incident response, security analysis, and integrating cloud infrastructure for insurance solutions.

Cloud