Site Reliability Engineer – Infrastructure Platforms, Intermediate to Senior Staff

🔥 29 minutes ago

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GitLab

GitLab

1001 - 5000 employees

Founded 2014

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

💰 Secondary Market on 2020-11

Consulting • Marketing • Artificial Intelligence

GitLab is the most comprehensive AI-powered DevSecOps platform, offering tools for automated software delivery, security, and compliance throughout the software development lifecycle. It provides solutions across areas such as AI-assisted development, continuous integration/continuous deployment (CI/CD), source code management, and vulnerability management. GitLab aims to simplify and accelerate software delivery by uniting development, security, and operations on a unified platform. It is particularly recognized for its AI code assistants and has been named a leader in the Gartner Magic Quadrant™ for DevOps Platforms, making it a preferred choice for many enterprises.

📋 Description

• Keep user-facing services and production systems reliable, scalable, and efficient • Build automation and tooling that reduces toil and replaces manual work with repeatable infrastructure-as-code-driven workflows • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling • Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps • Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early • Participate in incident response and post-incident reviews, turning learnings into automation and process improvements • Document runbooks, architecture decisions, and reviews • Collaborate asynchronously across Infrastructure Platforms teams and contribute to continuous reliability improvement at scale

🎯 Requirements

• Position open to candidates based in the United Kingdom only • Experience keeping production systems reliable, combining an operations mindset with real software engineering practice • Experience building net-new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch • Ability to read, debug, and reason about code, including behavior, performance, and failure modes • Experience with infrastructure as code, Kubernetes, and its ecosystem • Hands-on experience with at least one major cloud provider: GCP or AWS • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs • Comfort participating in on-call and incident response • Strong written communication and ability to operate as a manager-of-one in an async, distributed environment • Track record of using automation and increasingly AI to reduce toil and improve team workflows • Alignment with GitLab's values and commitment to working in accordance with them

🏖️ Benefits

• Benefits to support your health, finances, and well-being • Flexible Paid Time Off • Team Member Resource Groups • Equity Compensation & Employee Stock Purchase Plan • Growth and Development Fund • Parental Leave

Apply Now

Similar Jobs

🔥 14 hours ago

Arbor Education

51 - 200

📚 Education

🤝 B2B

Site Reliability Engineer improving availability, scalability and observability for Arbor’s school management platform. Supporting over 7,000 schools and trusts through reliable, resilient services.

🕒 3 days ago

IBM

10,000+ employees

💼 Consulting

🏭 Manufacturing

📦 Logistics

Senior DevOps Engineer managing Linux, AWS and production systems for Snappy Shopper’s rapid grocery delivery platform. Improving reliability, automation and incident response across UK Q-commerce infrastructure.

🕒 5 days ago

11:11 Systems

1001 - 5000

🏢 Enterprise

🔒 Cybersecurity

🤝 B2B

Infrastructure Deployment Engineer delivering hardware across 11:11 Systems’ global data centers. Leading deployments, infrastructure validation, documentation, and automation from planning through production.

🕒 5 days ago

11:11 Systems

1001 - 5000

🤝 B2B

🔒 Cybersecurity

🏢 Enterprise

Infrastructure Deployment Engineer delivering hardware infrastructure across 11:11 Systems’ UK data centers. Leading deployments, automation, documentation, and cross-functional infrastructure projects.

🕒 5 days ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior SRE operating AWS and EKS production systems for Salve.Inno Consulting’s recruitment clients. Owning incidents, reliability, observability and automation in distributed environments.