Senior Site Reliability Engineer – Cloud Platform

Stelle nicht auf LinkedIn

🕒 vor 3 Tagen

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Salve.Inno

Salve.Inno

11 - 50 Mitarbeiter

Gegründet 2024

💼 Beratung

📣 Marketing

📦 Logistik

Consulting • Marketing • Logistics

Salve. Inno ist eine Personalvermittlungs- und Beratungsfirma, die außergewöhnliche Talente mit Unternehmen durch personalisierte Einstellungsstrategien und globale Fernbeschaffung verbindet. Das Unternehmen ist auf die Personalbeschaffung für Positionen in Sektoren wie Marketing, Forex und iGaming spezialisiert und bietet Kandidatenbeschaffung, -screening sowie karriereseitengetriebene Einstellerfahrungen, wobei der Schwerpunkt auf DE&I, Kommunikation und innovativem Prozessaufbau liegt. Gegründet im Jahr 2024 und mit Hauptsitz in Gdańsk, Polen, operiert Salve. Inno mit einem kleinen Team und einer globalen Präsenz durch Remote-Jobangebote und Beratungsdienste.

Beschreibung

• Maintain the reliability, availability, and performance of production and pre-production environments • Monitor platform health and improve alerting, automation, and operational processes • Respond to production incidents, participate in root cause analysis, and implement long-term improvements • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards • Partner with software engineers to improve application reliability throughout the development lifecycle • Develop and maintain operational documentation, troubleshooting guides, and runbooks • Automate repetitive operational tasks to improve efficiency and reduce manual intervention • Participate in on-call rotations while continuously improving incident response processes • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams

🎯 Anforderungen

• Bachelor's or Master's degree in Engineering, Computer Science, or a related field • Strong experience operating Kubernetes or other container orchestration platforms • Experience supporting large-scale production services • Hands-on experience with AWS • Experience with Prometheus, Grafana, and ELK • Strong scripting skills in Bash, Python, or Go • Experience administering Linux-based production environments • Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible • Solid understanding of networking fundamentals, including TCP/IP, DNS, load balancing, and routing • Excellent troubleshooting, communication, and collaboration skills • A proactive mindset with a passion for automation and reliability • Nice to have: experience with SIP or VoIP technologies • Nice to have: familiarity with MySQL or PostgreSQL • Nice to have: experience with Redis or other NoSQL databases

🏖️ Vorteile

• Flexible remote working environment • Professional development opportunities, including training and technical learning • Opportunity to work on innovative cloud technologies used by customers worldwide • Collaborative engineering culture focused on knowledge sharing and continuous improvement • Modern Apple equipment provided • Inclusive, respectful workplace

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 7 Tagen

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Künstliche Intelligenz

Site Reliability Engineer maintaining CrowdStrike’s large-scale cybersecurity cloud platform. Automating operations, monitoring distributed systems, and leading incident response for reliable 24x7 service.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 16 Tagen

NICE

5001 - 10000

☁️ SaaS

🤖 Künstliche Intelligenz

📡 Telekommunikation

Forward Deployed Engineer building AI agents and full-stack automation for NICE’s customer-experience software. Integrating conversational systems with enterprise platforms and deploying customer self-service solutions.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 17 Tagen

Pliant

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Engineering Manager building Pliant’s Site Reliability function for its B2B payments platform. Establishing SLOs, incident processes, observability, and a new reliability engineering team.

🇬🇧 Vereinigtes Königreich – Remote

💰 €40.000.000 Series B - Pliant im 2025-04

⏰ Vollzeit

🟠 Senior

🔴 Experte

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 24 Tagen

PatSnap

501 - 1000

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Site Reliability Engineering Leader at PatSnap, leading the SRE team ensuring reliability for a global SaaS platform. Overseeing strategy, automation, and team development in cloud technologies.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 24 Tagen

TwinStream

51 - 200

🎖️ Verteidigung

💼 Beratung

📦 Logistik

DevOps Engineer maintaining and deploying cross-domain systems using Docker and AMQP architecture for TwinStream clients. Collaborating with teams and ensuring system performance and availability.

🇬🇧 Vereinigtes Königreich – Remote

💵 £70.000 - £85.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich