Senior Site Reliability Engineer, Compute Node Team

Job not on LinkedIn

Likely ghost job

🕒 January 26

🇳🇱 Netherlands – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 64%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nebius Group

Nebius Group

1001 - 5000 employees

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Nebius Group is building one of the world’s leading AI infrastructure companies, focusing on providing the necessary compute, storage, and tools for developers in the AI space. Based in Europe and listed on Nasdaq, Nebius has a global presence with R&D centers across Europe, North America, and Israel. The company's primary offering is an AI-centric cloud platform designed for intensive AI workloads, complemented by various other businesses involved in generative AI development, edtech, and autonomous technology.

📋 Description

• Ensure reliability, availability and performance of compute nodes running VMs • Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability.

🎯 Requirements

• Strong Linux expertise: • deep understanding of Linux user space and kernel space • knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) • clear understanding of system boundaries and constraints at different layers • Virtualization experience: • hands-on experience with QEMU/KVM • understanding of VM lifecycle, performance characteristics and failure modes • Containerization knowledge: • practical experience with containers, namespaces and cgroups • strong understanding of resource isolation and control • Strong debugging skills: • ability to reason about complex system failures • structured, hypothesis-driven approach to incident analysis • SRE mindset: • clear understanding of the SRE role in system design and operations • experience building and operating observability stacks, not just consuming them • ability to turn system behavior into actionable reliability signals.

🏖️ Benefits

• Competitive salary and comprehensive benefits package. • Opportunities for professional growth within Nebius. • Flexible working arrangements. • A dynamic and collaborative work environment that values initiative and innovation.

Apply Now