Senior Site Reliability Engineer, Cloud Platform

Stelle nicht auf LinkedIn

🕒 vor 10 Tagen

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Salve.Inno

Salve.Inno

11 - 50 Mitarbeiter

Gegründet 2024

💼 Beratung

📣 Marketing

📦 Logistik

Consulting • Marketing • Logistics

Salve. Inno ist eine Personalvermittlungs- und Beratungsfirma, die außergewöhnliche Talente mit Unternehmen durch personalisierte Einstellungsstrategien und globale Fernbeschaffung verbindet. Das Unternehmen ist auf die Personalbeschaffung für Positionen in Sektoren wie Marketing, Forex und iGaming spezialisiert und bietet Kandidatenbeschaffung, -screening sowie karriereseitengetriebene Einstellerfahrungen, wobei der Schwerpunkt auf DE&I, Kommunikation und innovativem Prozessaufbau liegt. Gegründet im Jahr 2024 und mit Hauptsitz in Gdańsk, Polen, operiert Salve. Inno mit einem kleinen Team und einer globalen Präsenz durch Remote-Jobangebote und Beratungsdienste.

Beschreibung

• Maintain the reliability, availability, and performance of production and pre-production environments • Monitor platform health and improve alerting, automation, and operational processes • Respond to production incidents, participate in root cause analysis, and implement long-term improvements • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards • Partner with software engineers to improve application reliability throughout the development lifecycle • Develop and maintain operational documentation, troubleshooting guides, and runbooks • Automate repetitive operational tasks to improve efficiency and reduce manual intervention • Participate in on-call rotations while continuously improving incident response processes • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams

🎯 Anforderungen

• Bachelor's or Master's degree in Engineering, Computer Science, or a related field • Strong experience operating Kubernetes or other container orchestration platforms • Experience supporting large-scale production services • Hands-on experience with AWS • Experience with Prometheus, Grafana, and ELK • Strong scripting skills (Bash, Python, or Go) • Experience administering Linux-based production environments • Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible • Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing) • Excellent troubleshooting, communication, and collaboration skills • A proactive mindset with a passion for automation and reliability • Nice to have: Experience with SIP or VoIP technologies • Nice to have: Familiarity with MySQL or PostgreSQL • Nice to have: Experience with Redis or other NoSQL databases

🏖️ Vorteile

• Flexible remote working environment • Professional development opportunities, including training and technical learning • Collaborative engineering culture focused on knowledge sharing and continuous improvement • Modern Apple equipment provided

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 16 Tagen

NICE

5001 - 10000

☁️ SaaS

🤖 Künstliche Intelligenz

📡 Telekommunikation

Forward Deployed Engineer building AI agents and full-stack automation for NICE’s customer-experience software. Integrating conversational systems with enterprise platforms and deploying customer self-service solutions.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 17 Tagen

Pliant

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Engineering Manager building Pliant’s Site Reliability function for its B2B payments platform. Establishing SLOs, incident processes, observability, and a new reliability engineering team.

🇬🇧 Vereinigtes Königreich – Remote

💰 €40.000.000 Series B - Pliant im 2025-04

⏰ Vollzeit

🟠 Senior

🔴 Experte

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 24 Tagen

PatSnap

501 - 1000

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Site Reliability Engineering Leader at PatSnap, leading the SRE team ensuring reliability for a global SaaS platform. Overseeing strategy, automation, and team development in cloud technologies.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 24 Tagen

TwinStream

51 - 200

🎖️ Verteidigung

💼 Beratung

📦 Logistik

DevOps Engineer maintaining and deploying cross-domain systems using Docker and AMQP architecture for TwinStream clients. Collaborating with teams and ensuring system performance and availability.

🇬🇧 Vereinigtes Königreich – Remote

💵 £70.000 - £85.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 24 Tagen

RTX

10.000+ Mitarbeiter

🚀 Luft- und Raumfahrt

🎖️ Verteidigung

🏭 Fertigung

Principal Site Reliability Engineer managing AWS infrastructures for Collins Aerospace. Delivering B2B products and ensuring service availability with scalable solutions in aviation technology.

🇬🇧 Vereinigtes Königreich – Remote

💰 €200.000 Grant - RTX im 2024-11

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich