Senior Site Reliability Engineer

Vermutlich ein Geisterjob

🕒 vor 2 Monaten

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 61%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Runware

Runware

11 - 50 Mitarbeiter

Gegründet 2023

🤖 Künstliche Intelligenz

🔌 API

📱 Medien

Artificial Intelligence • API • Media

Runware ist eine flexible generative KI-Plattform, die sich auf die Erstellung hochqualitativer Medien für Bilder und Videos über eine erschwingliche und schnelle API spezialisiert hat. Sie ist in der Lage, anspruchsvolle Aufgaben wie Bildgenerierung, Videoinferenz, Upscaling und Hintergrundentfernung auszuführen, was sie zu einem unverzichtbaren Werkzeug für Entwickler macht, die ihre Projekte mit KI-Technologie verbessern möchten. Angetrieben von einem maßgeschneiderten Sonic Inference Engine™ und erneuerbarer Energie, gewährleistet Runware eine schnelle und effiziente Mediengenerierung, ohne dass komplexe Infrastrukturen oder maschinelles Lernen erforderlich sind.

Beschreibung

• Own and improve the reliability, availability and performance of critical production services across the Runware platform • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

🎯 Anforderungen

• Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure • Have experience designing and operating observability systems using metrics, logs and distributed tracing • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation • Bonus • Experience operating high-throughput or low-latency APIs and distributed systems • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads • Experience with RabbitMQ or other distributed messaging and queueing systems • Experience operating MySQL, Redis, ClickHouse or similar production data systems • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments • Experience building automated scaling, capacity management or self-healing systems

🏖️ Vorteile

• Generous paid time off – vacation, sick days, public holidays • Meaningful stock options – share in the upside you create • Remote-first setup – work from home anywhere we can employ you • Flexible hours – own your schedule outside core collaboration blocks • Family leave – paid maternity, paternity, and caregiver time • Company retreats – twice-yearly gatherings in inspiring locations

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 3 Monaten

Omilia - Conversational Intelligence

201 - 500

💼 Beratung

🛡️ Versicherung

✈️ Reisen

Senior Site Reliability Engineer operating and maintaining production clusters while developing observability solutions. Collaborating with teams to enhance platform reliability through automation and monitoring.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

DeepHealth

11 - 50

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

DevOps Engineer managing AWS infrastructure and enhancing platform reliability at DeepHealth. Collaborating with teams to automate processes and improve software delivery.

🇬🇧 Vereinigtes Königreich – Remote

💵 £60.000 - £70.000 / Jahr

💰 €225.000 Grant im 2019-08

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

DeepHealth

11 - 50

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

DevOps Engineer responsible for AWS and Kubernetes platform management at DeepHealth, a healthcare SaaS provider. Ensuring cloud infrastructure is secure and reliable for efficient software delivery.

🇬🇧 Vereinigtes Königreich – Remote

💵 £60.000 - £70.000 / Jahr

💰 €225.000 Grant im 2019-08

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

Fuse Energy

11 - 50

💼 Beratung

📦 Logistik

⚡ Energie

Database Reliability Engineer ensuring reliability, performance, and scalability of database infrastructure at Fuse Energy. Responsible for building and maintaining data pipelines and analytical schemas.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

Intermedia Cloud Communications

1001 - 5000

💼 Beratung

🏥 Gesundheitswesen

⚖️ Rechtswesen

Team Lead DevOps Engineer leading a small team for a cloud communications provider. Overseeing Kubernetes, CI/CD processes, and collaborating with multiple teams.

🇬🇧 Vereinigtes Königreich – Remote

💰 Venture Round im 2017-02

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich