
1001 - 5000 Mitarbeiter
🤖 Künstliche Intelligenz
🏢 Unternehmen
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Die Nebius Group baut eines der weltweit führenden Unternehmen für KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools für Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfügt Nebius über eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die für rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.
🕒 vor 20 Tagen
🇳🇱 Niederlande – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
👻 Geisterscore 25%
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

1001 - 5000 Mitarbeiter
🤖 Künstliche Intelligenz
🏢 Unternehmen
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Die Nebius Group baut eines der weltweit führenden Unternehmen für KI-Infrastruktur auf und konzentriert sich darauf, die notwendige Rechenleistung, Speicherkapazität und Tools für Entwickler im KI-Bereich bereitzustellen. Mit Sitz in Europa und an der Nasdaq notiert verfügt Nebius über eine globale Präsenz mit F&E-Zentren in Europa, Nordamerika und Israel. Das zentrale Angebot des Unternehmens ist eine KI-zentrierte Cloud-Plattform, die für rechenintensive KI-Workloads ausgelegt ist, ergänzt durch verschiedene weitere Geschäftsbereiche in den Bereichen Generative KI, Edtech und autonome Technologien.
• Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense) • Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards • Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting) • Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents • Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes • Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast
• Strong production Linux fundamentals and a structured approach to debugging complex systems • Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.) • Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”) • Ability to write and maintain software/automation (Go is common for us; Python is also welcome) • Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows
• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams
Jetzt Bewerben🕒 vor 24 Tagen
Senior DevOps Engineer architecting, building, and scaling infrastructure for Kestra's SaaS platform. Innovating systems for deployment, monitoring, and automation at a remote-first company.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 24 Tagen
Senior DevOps Engineer at Publitas, a remote-first SaaS company. Responsible for AWS/GCP infrastructure and handling high-severity incidents with a focus on operational tasks.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
DevOps Lead responsible for automating software development and operations at a cybersecurity firm. Collaborating with multiple teams to enhance processes and optimize development experiences.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 4 Monaten
Database Reliability Engineer responsible for reliability and performance of ClickHouse core services. Collaborating with teams for process improvements, investigations, and incident response.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 4 Monaten
Cloud/DevOps Engineer responsible for Azure and IT infrastructure optimization at HelloFlex. Work closely with development teams to ensure efficient and secure deployments.
🗣️🇳🇱 Niederländisch erforderlich
🗣️🇺🇸🇬🇧 Englisch erforderlich