
51 - 200 employees
📚 Education
🤝 B2B
đź’° Private Equity Round on 2020-12
Education • B2B • Software
Arbor Education is a rapidly expanding company dedicated to transforming the way schools operate by freeing staff from administrative tasks and enhancing collaboration. Utilizing a Management Information System (MIS), Arbor Education provides tools to improve school processes and educational outcomes for over 5,000 schools. Their mission-driven team, consisting of ex-teachers, education technology engineers, and specialists, is passionate about providing effective solutions to improve the educational sector. Founded in 2011, Arbor Education is driven by a commitment to innovation and the well-being of teachers and students, along with a dedication to diversity and inclusion in its workforce.
🔥 0 minutes ago
🇬🇧 United Kingdom – Remote
⏰ Full Time
🟡 Mid-level
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🇬🇧 UK Skilled Worker Visa Sponsor
đź‘» Ghost score 10%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
📚 Education
🤝 B2B
đź’° Private Equity Round on 2020-12
Education • B2B • Software
Arbor Education is a rapidly expanding company dedicated to transforming the way schools operate by freeing staff from administrative tasks and enhancing collaboration. Utilizing a Management Information System (MIS), Arbor Education provides tools to improve school processes and educational outcomes for over 5,000 schools. Their mission-driven team, consisting of ex-teachers, education technology engineers, and specialists, is passionate about providing effective solutions to improve the educational sector. Founded in 2011, Arbor Education is driven by a commitment to innovation and the well-being of teachers and students, along with a dedication to diversity and inclusion in its workforce.
• Proactively monitor and analyse platform performance • Collaborate with engineering teams to address performance bottlenecks and ensure scalability • Assist engineering teams with implementing and reviewing SLOs • Improve observability through monitoring, alerting and dashboards using tools such as DataDog or Prometheus • Work with other teams to ensure effective monitoring and full coverage • Ensure services are highly available and resilient • Champion best practices for high-availability design • Devise runbooks and run game sessions to test disaster recovery plans, high availability and backups • Assess capacity and plan scaling for current and future business needs • Strategise and implement scalable solutions with the Head of Platform Engineering and Head of SRE • Work with the Platform team, feature teams, second-line support and stakeholders to provide a good level of customer service and embed SRE practices • Respond to and troubleshoot incidents, ensuring rapid resolution and minimising downtime • Participate in blameless postmortems to identify root causes and corrective actions • Develop and maintain playbooks and documentation
• Experience in performance monitoring and analysis • Capacity planning experience • Scripting and automation skills with relevant technologies • Experience with Infrastructure as Code, particularly Terraform • Understanding of relational database technologies and cloud versions such as AWS Aurora • Experience with messaging and distributed asynchronous workloads • Experience with nginx or similar technologies • Familiarity with SRE processes • Awareness of DevOps principles including the 3 ways and 5 ideals • Bonus: Experience with other database technologies and cloud platforms • Bonus: Past experience with enterprise solutions running at scale • Bonus: Familiarity with Kanban and Agile development processes • Bonus: Experience with containerisation, such as Docker • Bonus: Familiarity with software best practices including Refactoring, Clean Code, Domain-Driven Design and Test-Driven Development
• A dedicated wellbeing team championing mindfulness, lunch n learns, manager training, mental health first aid training and more • 32 days holiday plus Bank Holidays (25 days annual leave plus 7 extra company-wide days) • Life Assurance paid out at 3x annual salary • Comprehensive wellness benefit through AIG Smart Health, including 24/7 virtual GP service, mental health support, counselling, and personalised health checks • Private Dental Insurance with Bupa • Salary sacrifice Pension provided by Scottish Widows • Enhanced maternity and adoption leave (20 weeks full pay) and paternity leave (6 weeks full pay) • 5 free return-to-work maternity coaching sessions • Access to Calm and Bippit financial wellbeing coaching • Flexible working • Social committees planning team, office and company-wide events • Opportunity to work alongside a passionate team and see the impact of your work
Apply Nowđź•’ 3 days ago
Senior DevOps Engineer managing Linux, AWS and production systems for Snappy Shopper’s rapid grocery delivery platform. Improving reliability, automation and incident response across UK Q-commerce infrastructure.
đź•’ 4 days ago
Infrastructure Deployment Engineer delivering hardware across 11:11 Systems’ global data centers. Leading deployments, infrastructure validation, documentation, and automation from planning through production.
🇬🇧 United Kingdom – Remote
đź’° Private equity on 2021-10
⏰ Full Time
🟡 Mid-level
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
đź•’ 4 days ago
Infrastructure Deployment Engineer delivering hardware infrastructure across 11:11 Systems’ UK data centers. Leading deployments, automation, documentation, and cross-functional infrastructure projects.
🇬🇧 United Kingdom – Remote
đź’° Private equity on 2021-10
⏰ Full Time
🟡 Mid-level
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
đź•’ 4 days ago
Senior SRE operating AWS/EKS production systems for Salve.Inno Consulting’s recruitment business. Leading incident response, Kubernetes reliability, observability, automation, and disaster recovery.
đź•’ 4 days ago
Senior SRE operating AWS and EKS production systems for Salve.Inno Consulting’s recruitment clients. Owning incidents, reliability, observability and automation in distributed environments.