Senior Site Reliability Engineer, SRE

Job not on LinkedIn

🔥 12 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Investigo

Investigo

201 - 500 employees

We’re a leading recruiter with a team of over 250 consultants working in London, Guildford, Milton Keynes, St Albans, Birmingham, New York, Philadelphia, and San Diego.

📋 Description

• Operate, harden and extend production OpenShift / OKD / Kubernetes clusters across on-premises and hybrid environments. • Supporting migrations, helping modernise the underlying compute and infrastructure layer. • Own CI/CD processes across the full lifecycle of platform and application components. • Own and mature GitOps deployment practices, particularly using tools such as Argo CD. • Support cloud-native application delivery using tools such as Helm and Kustomize. • Maintain and improve core platform services including Keycloak, ingress, observability, certificate management, service mesh and container registry capabilities. • Build and operate observability across logs, metrics, traces, alerting, SLOs and error budgets. • Improve platform hardening in line with secure and regulated environment requirements. • Automate repeatable operational tasks using tools such as Ansible, Terraform, Helm, Kustomize, Go, Python or similar. • Lead incident response activity, support blameless post-mortems and drive systemic fixes. • Partner with networking and security teams on platform integration, segmentation, load balancing and accreditation evidence. • Create and maintain clear technical documentation, runbooks, design notes and operational guidance. • Mentor engineers and act as a senior technical authority across cloud and Kubernetes operations. • Participate in an on-call rota, with appropriate compensation.

🎯 Requirements

• Strong experience running production Kubernetes environments, not just consuming or deploying into them. • Strong Linux fundamentals, including systems, networking, storage and performance troubleshooting. • Experience with Kubernetes distributions such as OKD, OpenShift, vanilla Kubernetes, Rancher, EKS, AKS or GKE. • Infrastructure as code experience, including Ansible plus Terraform or equivalent. • Experience with Helm, Kustomize or similar cloud-native deployment tooling. • GitOps and CI/CD experience managing full application and component lifecycles, using tools such as Argo CD, Flux, GitHub Actions or similar. • Observability across logs, metrics and traces, using tools such as Prometheus, Grafana, Elastic Stack, LGTM and OpenTelemetry. • Experience with identity and access technologies such as OIDC, SAML, SCIM or Keycloak. • Experience with virtualisation or infrastructure platforms such as KVM, libvirt or VMware. • Scripting or tooling experience using Go, Python, shell scripting or similar. • Strong troubleshooting, problem-solving and analytical skills. • Experience working in secure, regulated or enterprise-scale environments. • Strong written and verbal communication skills, with the ability to produce clear documentation, runbooks, post-mortems and technical guidance. • Eligibility to hold UK SC clearance. • Desirable (Not Essential) • Specific OpenShift or OKD experience, including operators, MachineConfig or SCCs. • Service mesh experience such as Istio or Linkerd. • Policy engine experience such as OPA, Gatekeeper or Kyverno. • Software supply chain security experience, including SBOMs, image signing, admission control or tools such as Sigstore. • Storage experience such as Ceph, Longhorn, OpenShift Data Foundation or equivalent. • Networking experience including BGP, VXLAN, Palo Alto or Juniper technologies. • AI, ML or GPU-enabled platform operations. • CKA, CKAD, CKS, Red Hat certifications or equivalent. • Active or recent UK SC clearance. • Recognised open-source contributions to the Kubernetes ecosystem.

🏖️ Benefits

• Private Medical • Health Cash Plan • 4x Life Assurance • Inclusive Culture: Enjoy an inclusive culture and environment. • Holiday: Generous holiday allowance. • Learning: Access to continuous learning and development opportunities. • Bonus Potential: Bonus potential based on performance and business-related factors. • Discounts: Discounts on a wide range of products and services. • Pension: Pension scheme contributions. • EV Car Scheme • Regular Pay Reviews • More Benefits: Explore additional benefits on our career site.

Apply Now

Similar Jobs

🕒 Yesterday

GitLab

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Senior Backend Engineer developing cloud-native and self-managed deployment environments for GitLab. Building consistency in deployment across development and production environments.

Cloud

Kubernetes

🕒 2 days ago

Runware

11 - 50

🤖 Artificial Intelligence

🔌 API

📱 Media

Site Reliability Engineer ensuring reliability and performance of Runware's AI platforms. Collaborating across software, infrastructure, and operations to enhance observability and reduce incidents.

Distributed Systems

Kubernetes

MySQL

PHP

Python

RabbitMQ

Redis

Go

🕒 2 days ago

Ensono

1001 - 5000

Senior DevOps Consultant at Ensono delivering complex projects with deep engineering skills and a focus on quality. Engage in project lifecycle and collaborate with client teams in a remote setting.

AWS

Azure

Cloud

Google Cloud Platform

Java

JavaScript

Kubernetes

Node.js

Terraform

🕒 5 days ago

Salve.Inno

11 - 50

🎯 Recruiter

🤝 B2B

Senior Site Reliability Engineer responsible for maintaining and improving cloud platform reliability at Salve.Inno Consulting. Collaborating with engineering teams to implement best practices and drive operational excellence.

Ansible

AWS

DNS

Grafana

Kubernetes

Linux

Prometheus

Python

TCP/IP

Terraform

Go

🕒 5 days ago

Salve.Inno

11 - 50

🎯 Recruiter

🤝 B2B

Senior Site Reliability Engineer for Salve.Inno Consulting enhancing cloud platform reliability and driving operational excellence through observability and automation.

Ansible

AWS

DNS

Grafana

Kubernetes

Linux

Prometheus

Python

TCP/IP

Terraform

Go