Senior Site Reliability Engineer, Compute Node Team

Stelle nicht auf LinkedIn

🕒 vor 6 Monaten

🇳🇱 Niederlande – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung ßberprßfen.

Logo of Nebius Group

Nebius Group

1001 - 5000 Mitarbeiter

🤖 Künstliche Intelligenz

🏢 Unternehmen

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Die Nebius Group baut eines der weltweit fßhrenden Unternehmen fßr KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools fßr Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfßgt Nebius ßber eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die fßr rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.

Beschreibung

• Ensure reliability, availability and performance of compute nodes running VMs • Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability.

🎯 Anforderungen

• Strong Linux expertise: • deep understanding of Linux user space and kernel space • knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) • clear understanding of system boundaries and constraints at different layers • Virtualization experience: • hands-on experience with QEMU/KVM • understanding of VM lifecycle, performance characteristics and failure modes • Containerization knowledge: • practical experience with containers, namespaces and cgroups • strong understanding of resource isolation and control • Strong debugging skills: • ability to reason about complex system failures • structured, hypothesis-driven approach to incident analysis • SRE mindset: • clear understanding of the SRE role in system design and operations • experience building and operating observability stacks, not just consuming them • ability to turn system behavior into actionable reliability signals.

🏖️ Vorteile

• Competitive salary and comprehensive benefits package. • Opportunities for professional growth within Nebius. • Flexible working arrangements. • A dynamic and collaborative work environment that values initiative and innovation.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 7 Monaten

KPN

10.000+ Mitarbeiter

📡 Telekommunikation

🛍️ eCommerce

🔒 Cybersecurity

Storage DevOps Engineer responsible for managing enterprise storage solutions in an innovative IT environment at KPN. Focused on automation, performance, and collaboration with Cloud and Connectivity teams.

🇳🇱 Niederlande – Remote

💵 €5.038 - €7.575 / Monat

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇳🇱 Niederländisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 8 Monaten

KPN

10.000+ Mitarbeiter

📡 Telekommunikation

🛍️ eCommerce

🔒 Cybersecurity

DevOps Engineer - Connectivity for KPN's Tech Hub Private Cloud, managing Cisco, F5, and Fortinet components and developing virtual networking solutions.

🇳🇱 Niederlande – Remote

💵 €4.499 - €6.582 / Monat

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇳🇱 Niederländisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 9 Monaten

KPN

10.000+ Mitarbeiter

📡 Telekommunikation

🛍️ eCommerce

🔒 Cybersecurity

DevOps Engineer developing automated solutions for storage and backup within KPN's Private Cloud. Collaborating with a team to optimize IT environments for innovative companies while adhering to compliance and security standards.

🇳🇱 Niederlande – Remote

💵 €5.038 - €7.575 / Monat

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇳🇱 Niederländisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 10 Monaten

KPN

10.000+ Mitarbeiter

📡 Telekommunikation

🛍️ eCommerce

🔒 Cybersecurity

DevOps engineer automating network functions like Firewalling and Loadbalancing for Datacenter platforms. Seeking experienced professional for transformation of legacy environments.

🇳🇱 Niederlande – Remote

💵 €5.038 - €7.575 / Monat

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇳🇱 Niederländisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 10 Monaten

KPN

10.000+ Mitarbeiter

📡 Telekommunikation

🛍️ eCommerce

🔒 Cybersecurity

Automate and operate KPN datacenter network functions, migrating legacy systems to OCP and CI/CD-driven solutions.

🇳🇱 Niederlande – Remote

💵 €5.038 - €7.575 / Monat

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇳🇱 Niederländisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich