Senior Site Reliability Engineer, Observability

Job not on LinkedIn

🔥 0 minutes ago

🌐 United Kingdom, Ireland – Remote

infoinfo

💵 £80k - £95k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of IAM Cloud

IAM Cloud

51 - 200 employees

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

IAM Cloud is a company that specializes in cloud-based solutions for IT management, offering products like Cloud Drive Mapper, which provides secure access to cloud storage without syncing, and forthcoming solutions like Authentic and Identity Exchange (IDx) for enhanced security and identity management. Founded on a philosophy of solving software-related problems, IAM Cloud emphasizes security and compliance, holding ISO27001 certification. They cater to a diverse range of clients globally, from small organizations to entities with over 250,000 users, and foster a customer-centric and partnership-driven approach. Notably, IAM Cloud is fully employee-owned, prioritizing customers over investor-driven growth, and has been recognized as a global partner of Microsoft.

📋 Description

• Define and evolve the observability strategy, target architecture and platform standards • Establish instrumentation, telemetry destinations, retention, cost and governance policies • Design Grafana dashboards and alerting around services and customer journeys • Establish SLOs, SLIs and error budgets for customer-facing services • Rationalise alerting so every page is actionable, owned and documented with a runbook • Lead the move to a modern incident response and on-call platform integrated with Microsoft Teams • Design rota, escalation and post-incident review processes • Champion observability-first practices across the engineering team • Create telemetry governance policy covering Log Analytics tables, retention, sampling and cardinality • Select, deliver and cut over to replacement incident response and on-call tooling before the current tool's end-of-life • Define SLOs and SLIs with dashboards and burn-rate alerting • Design and implement OpenTelemetry instrumentation across .NET services and collector pipelines • Own Azure Monitor, Log Analytics and Application Insights configuration as code, using Bicep first • Own Grafana data sources, dashboards-as-code, alert rules and team conventions • Run incident response and on-call tooling, including routing, escalation, integrations and Teams workflows • Contribute observability standards to the Kubernetes and Prometheus direction • Lead incident response, run blameless post-incident reviews and turn findings into engineering work • Track and reduce MTTR, alert volume, incident recurrence and telemetry cost per service • Evaluate observability tooling and AI-assisted operations capabilities and provide recommendations • Maintain standards, runbooks, decision records and rationale for thresholds • Participate in a light-touch shared on-call rotation

🎯 Requirements

• 6+ years in engineering, with 3+ years in a reliability, observability or production-operations role owning outcomes • Defined SLOs, SLIs and error budgets for real services • Experience reducing alert noise and rationalising alerting estates • Hands-on OpenTelemetry experience, including instrumenting services, running collectors, and making sampling and cardinality decisions • Experience managing telemetry cost and retention governance • Deep Azure Monitor, Log Analytics and Application Insights experience, including KQL • Strong Grafana skills, including dashboards, alerting and managing both as code • Infrastructure-as-code experience; Bicep is the standard, with strong Terraform experience transferable • Clear written and spoken communication at C1 level or above, or native-level business English • Experience selecting, implementing or migrating incident-management and on-call platforms (nice to have) • Prometheus-based monitoring and alerting, including exporters, recording rules and cardinality management (nice to have) • Kubernetes in production, ideally AKS (nice to have) • Software development background, ideally .NET (nice to have) • Familiarity with AI-assisted operations and coding tools (nice to have) • Azure certifications AZ-104, AZ-400 or AZ-305, or CKA (nice to have) • Experience in a small or scale-up environment where you were the observability function (nice to have)

🏖️ Benefits

• Guaranteed £2,000 pay rise every year, separate from merit or promotion increases • 26 days holiday plus public holidays, with the option to swap public holidays • Birthday and work anniversary day off every year • Up to 4 weeks per year working from anywhere in the world • Genuinely flexible, fully remote working • Generous Becoming a Parent leave • Medical scheme including remote GP, dental, optical, and diagnostics (UK employees) • Access to Support Room confidential counselling, therapy, and coaching • Help@Hand employee assistance programme • 5x salary life insurance through Unum • Up to 10% pension match (UK employees) • Annual learning and development budget per function • In-person team events a few times a year • Async-friendly working • Shared, light on-call rotation

Apply Now

Similar Jobs

🕒 September 16

Adapty.io

51 - 200

☁️ SaaS

🤝 B2B

Senior DevOps Engineer owning Adapty's Proxmox, Ceph, Kubernetes, and cloud infrastructure. Supporting the SaaS platform powering mobile app subscriptions at global scale.

🕒 September 10

Jones Lang LaSalle Americas, Inc.

10,000+ employees

🏠 Real Estate

🤝 B2B

💼 Consulting

Senior Reliability Engineer optimizing reliability, maintenance, and lifecycle asset management for JLL’s real estate services. Supporting a Life Sciences client across EMEA locations.

🕒 September 8

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Site Reliability Engineer maintaining reliable, scalable production infrastructure for GitLab’s DevSecOps platform. Automating Kubernetes, cloud, observability, and incident-response operations across distributed Infrastructure Platforms teams.

🕒 September 8

Arbor Education

51 - 200

📚 Education

🤝 B2B

Site Reliability Engineer improving availability, scalability and observability for Arbor’s school management platform. Supporting over 7,000 schools and trusts through reliable, resilient services.

🕒 September 4

IBM

10,000+ employees

💼 Consulting

🏭 Manufacturing

📦 Logistics

Senior DevOps Engineer managing Linux, AWS and production systems for Snappy Shopper’s rapid grocery delivery platform. Improving reliability, automation and incident response across UK Q-commerce infrastructure.