Senior Engineer, Network Observability

🕒 vor 2 Monaten

🌐 Vereinigtes Königreich, Irland – Remote

infoinfo

⏰ Vollzeit

🟠 Senior

🛜 Netzwerkingenieur / Netzwerkadministrator

👻 Geisterscore 28%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of CoreWeave

CoreWeave

11 - 50 Mitarbeiter

Gegründet 2017

💼 Beratung

🏭 Fertigung

📦 Logistik

💰 €100.000.000 Debt Financing im 2022-12

Consulting • Manufacturing • Logistics

CoreWeave ist ein spezialisierter Cloud-Anbieter, der eine enorme Menge an GPU-Rechenressourcen auf der schnellsten und flexibelsten Infrastruktur der Branche bereitstellt. Als NVIDIA Elite Cloud Solutions-Anbieter für Compute und Visualisierung entwickelt CoreWeave Cloud-Lösungen für rechenintensive Anwendungsfälle - VFX und Rendering, maschinelles Lernen und KI, Stapelverarbeitung und Pixel-Streaming - die bis zu 35-mal schneller und 80 % kostengünstiger sind als die großen, generischen Public Clouds. Erfahren Sie mehr unter www.coreweave.com.

Beschreibung

• We’re seeking a talented and experienced Senior Engineer for Network Observability to join our Network Observability team. In this role, you will be a key player in designing, developing, and maintaining the monitoring, telemetry, and observability systems that keep CoreWeave’s GPU cloud network operating reliably and at scale. • You’ll focus on building solutions that provide real-time insights into network performance, ensuring that issues are detected proactively and resolved quickly. • Develop, optimize, and maintain network observability platforms. Use your skills in Python and Golang to create and automate collectors, exporters, and dashboards that provide deep visibility into network health and performance. • Collaborate with Network Engineering and Platform teams to ingest and unify logs, metrics, and events from a variety of platforms (Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux, etc.) into a single observability pipeline. • Design and implement scalable telemetry solutions using protocols like gNMI, SNMP, and streaming analytics. Ensure advanced alerting and anomaly detection with tools such as Prometheus, Grafana, and Alertmanager. • Work closely with network developers, site reliability engineers, and security teams to integrate observability solutions across the broader infrastructure. • Participate in design discussions, RFCs, and architectural decisions. • Join a rotating on-call schedule to troubleshoot and resolve observability-related issues. Provide timely support to operations teams, quickly isolating and fixing problems when they arise. • Guide junior team members, share best practices, and foster a culture of continuous learning and improvement within the observability domain.

🎯 Anforderungen

• Deep familiarity with Prometheus, Grafana, Alertmanager, gNMI, and SNMP. Experience writing or extending custom metric collectors/exporters is a plus. • Experience as a Network Engineer, SRE, Software Developer, or Systems Administrator in large-scale environments. A track record of building and operating robust telemetry and monitoring solutions is a plus. • Passion for automating tasks and processes. You find satisfaction in creating workflows that handle repetitive tasks and reduce human error to near zero. • Comfortable containerizing solutions in Kubernetes, designing, building, and deploying container-based workloads efficiently. • Proficient with Python, Go, and Bash, plus familiarity with configuration management and templating tools (e.g., Ansible, Jinja2). . • Strong knowledge of Linux systems and IP networking concepts, with hands-on experience in routing, switching, and network troubleshooting. • Practical knowledge with a variety of platforms, including Arista EOS, NVIDIA Cumulus Linux, Nokia SR OS, and SR Linux. • Collaborative, humble, and always ready to help others while staying open to learning from more senior colleagues.

🏖️ Vorteile

• Family-level Medical Insurance • Family-level Dental Insurance • Generous Pension Contribution • Life Assurance at 4x Salary • Critical Illness Cover • Employee Assistance Programme • Tuition Reimbursement • Work culture focused on innovative disruption

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 3 Monaten

Roc Technologies

201 - 500

💼 Beratung

🎖️ Verteidigung

🏥 Gesundheitswesen

Senior Network Engineer managing high-quality network solutions for enterprise customers across the UK. Involved in discovery, design, implementation, and troubleshooting of network projects.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

🛜 Netzwerkingenieur / Netzwerkadministrator

🗣️🇺🇸🇬🇧 Englisch erforderlich

Cloud

Switching

🕒 vor 4 Monaten

Clir Renewables

51 - 200

⚡ Energie

🤖 Künstliche Intelligenz

☁️ SaaS

Network Engineer managing connections and improving network systems for Clir Renewables. Focus on renewable energy efficiency and cost reduction while collaborating with clients.

🇬🇧 Vereinigtes Königreich – Remote

💵 £40.000 - £45.000 / Jahr

💰 €1.493.600 Venture Round - Clir Renewables im 2024-01

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🛜 Netzwerkingenieur / Netzwerkadministrator

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Endeavor

5001 - 10000

🏨 Gastgewerbe

📣 Marketing

Network Engineer responsible for supporting TKO’s global network infrastructure, including datacenters and cloud. Ensuring security and operational compliance with enterprise networking standards.

🗣️🇺🇸🇬🇧 Englisch erforderlich