Site Reliability Engineer – Network Infrastructure

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nebius Group

Nebius Group

1001 - 5000 employees

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Nebius Group is building one of the world’s leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.

📋 Description

• Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense) • Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards • Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting) • Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents • Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes • Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast

🎯 Requirements

• Strong production Linux fundamentals and a structured approach to debugging complex systems • Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.) • Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”) • Ability to write and maintain software/automation (Go is common for us; Python is also welcome) • Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows

🏖️ Benefits

• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams

Apply Now

Similar Jobs

🕒 3 days ago

Kestra

51 - 200

☁️ SaaS

🤖 Artificial Intelligence

🏢 Enterprise

Senior DevOps Engineer architecting, building, and scaling infrastructure for Kestra's SaaS platform. Innovating systems for deployment, monitoring, and automation at a remote-first company.

AWS

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

Kafka

Kubernetes

Postgres

Terraform

🕒 4 days ago

Publitas

51 - 200

☁️ SaaS

🛍️ eCommerce

🤝 B2B

Senior DevOps Engineer at Publitas, a remote-first SaaS company. Responsible for AWS/GCP infrastructure and handling high-severity incidents with a focus on operational tasks.

AWS

Google Cloud Platform

🕒 July 22

SAI360

201 - 500

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Senior DevOps Build Engineer at SAI360 optimizing infrastructure automation and CI/CD frameworks. Collaborating with global teams and mentoring engineers in modern DevOps practices.

Ansible

Cloud

Docker

Java

Kubernetes

Linux

Maven

Spring

Terraform

🕒 July 2

EclecticIQ

51 - 200

🔒 Cybersecurity

🏢 Enterprise

☁️ SaaS

DevOps Lead responsible for automating software development and operations at a cybersecurity firm. Collaborating with multiple teams to enhance processes and optimize development experiences.

Ansible

AWS

Cloud

Docker

EC2

Kubernetes

Linux

Packer

Python

SDLC

TCP/IP

Terraform

Go

🕒 June 19

Reddit, Inc.

501 - 1000

💼 Consulting

📣 Marketing

📱 Media

Senior Site Reliability Engineer for Reddit's Ads platform focused on improving reliability and operational excellence. Collaborating with Ads Engineering to build and maintain scalable systems.

Cloud

Distributed Systems

Linux

Python

Go