Site Reliability Lead

🕒 vor 2 Monaten

🇬🇧 Vereinigtes Königreich – Remote

💵 £80.000 - £90.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🇬🇧 UK-Skilled-Worker-Visum-Sponsor

infoinfo

👻 Geisterscore 20%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Arbor Education

Arbor Education

51 - 200 Mitarbeiter

📚 Bildung

🤝 B2B

💰 Private Equity Round im 2020-12

Education • B2B • Software

Arbor Education ist ein führender Anbieter von Schulverwaltungssoftware im Vereinigten Königreich und bietet umfassende Lösungen wie Arbor School MIS und Arbor MAT MIS. Ihre Plattform ist darauf ausgelegt, Schulen und Multi-Akademie-Trusts bei der Rationalisierung von Abläufen, der Verbesserung des Datenmanagements und der Steigerung der Produktivität des Personals zu unterstützen. Mit über 7. 000 Schulen, die Arbor nutzen, revolutioniert es die Betriebsweise von Bildungseinrichtungen, indem es wertvolle Zeit und Ressourcen spart, mit effektiven Werkzeugen, die auf verschiedene Schultypen wie Grund-, Sekundar- und Förderschulen zugeschnitten sind.

Beschreibung

• Define and guide system architecture, balancing trade-offs between speed, scalability, maintainability, and security to meet business goals. • Champion accountability from design through to production by ensuring systems are observable and meet agreed Service Level Objectives (SLOs). Drive continuous improvement in platform reliability, performance, and efficiency. • Lead Root Cause Analysis (RCA) when issues occur and contribute to optimizing the incident response process and framework. • Drive automation initiatives across the team to reduce operational toil and improve system efficiency. • Uphold coding standards, promote automated testing, and work with the architecture community to drive technology adoption and share best practices across teams. Ensure production readiness standards for all services. • Lead technical estimation and feasibility assessments, ensuring plans are realistic and aligned with team capacity. Contribute to structured release planning and support post-release reviews. • Mentor and coach engineers through constructive feedback, knowledge sharing, and motivation. Foster alignment and help the team galvanise around technical solutions and goals. • Work closely with Product Managers, Engineering Managers, and other engineers to align technical direction with product strategy. Communicate complex technical concepts clearly to both technical and non-technical stakeholders.

🎯 Anforderungen

• Extensive professional experience in SRE, DevOps, or Platform Engineering on complex, scalable systems. • Extensive expertise with AWS and distributed cloud architectures. • Proven experience operating platforms serving a high volume of requests (~1000 req/sec). • Advanced proficiency with Terraform and configuration management tools. • Strong skills in Python, Go, or a similar language for automation and tooling. • Deep experience with monitoring and observability platforms (e.g., DataDog, Prometheus, or equivalent), plus incident/problem management. • Expert understanding of distributed systems, microservices, and resilience patterns. • Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes, ECS). • Practical experience with building and maintaining CI/CD pipelines for automated deployments. • Demonstrated ability in mentoring and supporting the growth of fellow engineers. • Bonus Skills • Experience with chaos engineering and reliability testing. • Knowledge of security best practices and compliance frameworks. • Background in agile and lean methodologies (Scrum/Kanban). • Contributions to open-source projects or the SRE community.

🏖️ Vorteile

• A dedicated wellbeing team who champion initiatives such as mindfulness, lunch n learns, manager training, mental health first aid training and much more! • 32 days holiday (plus Bank Holidays). This is made up of 25 days annual leave plus 7 extra company wide days given over Easter, Summer & Christmas • Life Assurance paid out at 3x annual salary • Comprehensive wellness benefit provided by AIG Smart Health, which provides a 24/7 virtual GP service, Mental health support, Counselling, and personalised Health Checks • Private Dental Insurance with Bupa • Salary sacrifice Pension provided by Scottish Widows • Enhanced maternity and adoption leave (20 weeks full pay) and paternity (6 weeks full pay) pay • 5 free return to work maternity coaching sessions, helping you adapt to this new exciting time of life! • Access to services such as Calm and Bippit (financial wellbeing coaching) • All of our roles champion flexible working and we are happy to discuss what this means to you • Social committees that plan team, office and company wide events to bring people together and celebrate success • Dedicated professional development training budget (CPD courses, upskilling resources, professional memberships etc) • Volunteer with a charity of your choice for a day each year • Dog friendly offices!

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 3 Monaten

Flosum

201 - 500

🤝 B2B

☁️ SaaS

Salesforce DevOps Evangelist promoting Flosum's DevOps and data management platform. Creating engaging content and speaking at major Salesforce events to enhance brand recognition and community involvement.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 3 Monaten

Intermedia Cloud Communications

1001 - 5000

💼 Beratung

🏥 Gesundheitswesen

⚖️ Rechtswesen

DevOps Engineer deploying and managing application infrastructure for a leading cloud tech provider. Focused on utilizing Kubernetes, GCP, and infrastructure automation tools.

🇬🇧 Vereinigtes Königreich – Remote

💰 Venture Round im 2017-02

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

NICE

5001 - 10000

☁️ SaaS

🤖 Künstliche Intelligenz

📡 Telekommunikation

SRE - NOC role focuses on service reliability, incident response, and operational automation. Precision in dealing with operational toil through engineering practices for global operations at NICE.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Ripjar

51 - 200

💸 Finanzen

📋 Compliance

🤖 Künstliche Intelligenz

DevOps Engineer ensuring reliability and security of infrastructure for software combating financial crime at Ripjar. Focus on continuous improvement and automation within a remote-first team.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Cencora

10.000+ Mitarbeiter

💼 Beratung

📦 Logistik

🏥 Gesundheitswesen

Lead DevOps Engineer overseeing Azure infrastructure and CI/CD pipelines improvements at Cencora. Mentor engineers and align initiatives with business goals in the pharmaceutical consulting sector.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich