Site Reliability Engineering Leader

🔥 1 minute ago

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🇬🇧 UK Skilled Worker Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of PatSnap

PatSnap

501 - 1000 employees

Founded 2007

💼 Consulting

🏥 Healthcare

📦 Logistics

💰 $300M Series E on 2021-03

Consulting • Healthcare • Logistics

PatSnap is a leading provider of IP and R&D innovation intelligence platforms. It leverages advanced AI technologies to help companies enhance productivity, drive innovation, and accelerate business growth by providing insights into patent and R&D processes. PatSnap's offerings aim to reduce R&D costs, assist in the discovery of new materials, and support life sciences intelligence. It provides a comprehensive database and analysis tools that are used by innovators to automate sequences, assess risk, and enhance innovation strategies.

📋 Description

• Build, lead and develop the UK SRE team, establishing operational standards, best practices, and reliability goals. • Ensure the high availability, stability, security, and performance of business-critical platforms and services. • Define and drive the operational strategy for our global SaaS platform, ensuring exceptional reliability, availability and performance. • Lead major incident management, acting as the senior escalation point during critical production events. • Establish and monitor reliability metrics, including SLIs, SLOs and operational KPIs. • Drive automation across infrastructure, deployments, monitoring and operational workflows to improve efficiency and reduce manual effort. • Champion the adoption of AI-powered operations, leveraging modern AI technologies to enhance engineering productivity and operational excellence. • Partner with Engineering, Product, Security and Infrastructure teams to improve platform architecture, scalability and operational readiness. • Lead disaster recovery planning, operational resilience initiatives and risk management across the platform. • Continuously evaluate emerging cloud, AI and platform technologies to keep PatSnap at the forefront of engineering excellence.

🎯 Requirements

• Bachelor’s degree in Computer Science or a related field, with at least 8 years of experience in DevOps, SRE, or infrastructure operations. • Proven experience leading technical teams and managing production environments at scale. • Strong expertise in cloud platforms (AWS preferred), Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code, and observability platforms. • Deep understanding of distributed systems, high-availability architectures, and large-scale SaaS environments. • Experience driving automation and operational excellence initiatives. • Hands-on experience using AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar technologies to improve engineering productivity. • Strong problem-solving, leadership, communication, and stakeholder management skills. • Fluent in English; Mandarin is highly desirable to facilitate collaboration with teams across multiple regions.

🏖️ Benefits

• Build technology that powers global innovation • Lead AI-powered engineering • Own a business-critical function • Work with modern technologies • Collaborate globally • Grow your leadership career

Apply Now

Similar Jobs

🔥 5 hours ago

TwinStream

51 - 200

🎖️ Defense

💼 Consulting

📦 Logistics

DevOps Engineer maintaining and deploying cross-domain systems using Docker and AMQP architecture for TwinStream clients. Collaborating with teams and ensuring system performance and availability.

🔥 11 hours ago

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Principal Site Reliability Engineer managing AWS infrastructures for Collins Aerospace. Delivering B2B products and ensuring service availability with scalable solutions in aviation technology.

🕒 Yesterday

Cognativ

11 - 50

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Senior Site Reliability Engineer managing reliability for a distributed, camera-based video monitoring and AI alerting platform. Focusing on operational health, service objectives, and incident response.

🕒 Yesterday

Partnerize

201 - 500

📣 Marketing

☁️ SaaS

🤝 B2B

DevOps Engineer working on deployment, maintenance, and optimization of cloud-based environments. Collaborating with a software development team to shape cutting-edge online marketing technology.

Amazon Redshift

Ansible

AWS

Chef

EC2

Linux

MySQL

Prometheus

Puppet

SaltStack

SQL

🕒 Yesterday

Peratera

11 - 50

💳 Fintech

🤝 B2B

🔌 API

DevOps/SRE Engineer responsible for platform reliability and automation at UK fintech. Building and evolving cloud infrastructure with automation and observability practices.