
51 - 200 employees
💼 Consulting
📣 Marketing
📦 Logistics
Consulting • Marketing • Logistics
IAM Cloud is a company that specializes in cloud-based solutions for IT management, offering products like Cloud Drive Mapper, which provides secure access to cloud storage without syncing, and forthcoming solutions like Authentic and Identity Exchange (IDx) for enhanced security and identity management. Founded on a philosophy of solving software-related problems, IAM Cloud emphasizes security and compliance, holding ISO27001 certification. They cater to a diverse range of clients globally, from small organizations to entities with over 250,000 users, and foster a customer-centric and partnership-driven approach. Notably, IAM Cloud is fully employee-owned, prioritizing customers over investor-driven growth, and has been recognized as a global partner of Microsoft.
🔥 0 minutes ago
🌐 United Kingdom, Ireland – Remote
💵 £80k - £95k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
💼 Consulting
📣 Marketing
📦 Logistics
Consulting • Marketing • Logistics
IAM Cloud is a company that specializes in cloud-based solutions for IT management, offering products like Cloud Drive Mapper, which provides secure access to cloud storage without syncing, and forthcoming solutions like Authentic and Identity Exchange (IDx) for enhanced security and identity management. Founded on a philosophy of solving software-related problems, IAM Cloud emphasizes security and compliance, holding ISO27001 certification. They cater to a diverse range of clients globally, from small organizations to entities with over 250,000 users, and foster a customer-centric and partnership-driven approach. Notably, IAM Cloud is fully employee-owned, prioritizing customers over investor-driven growth, and has been recognized as a global partner of Microsoft.
• Define and evolve the observability strategy, target architecture and platform standards • Establish instrumentation, telemetry destinations, retention, cost and governance policies • Design Grafana dashboards and alerting around services and customer journeys • Establish SLOs, SLIs and error budgets for customer-facing services • Rationalise alerting so every page is actionable, owned and documented with a runbook • Lead the move to a modern incident response and on-call platform integrated with Microsoft Teams • Design rota, escalation and post-incident review processes • Champion observability-first practices across the engineering team • Create telemetry governance policy covering Log Analytics tables, retention, sampling and cardinality • Select, deliver and cut over to replacement incident response and on-call tooling before the current tool's end-of-life • Define SLOs and SLIs with dashboards and burn-rate alerting • Design and implement OpenTelemetry instrumentation across .NET services and collector pipelines • Own Azure Monitor, Log Analytics and Application Insights configuration as code, using Bicep first • Own Grafana data sources, dashboards-as-code, alert rules and team conventions • Run incident response and on-call tooling, including routing, escalation, integrations and Teams workflows • Contribute observability standards to the Kubernetes and Prometheus direction • Lead incident response, run blameless post-incident reviews and turn findings into engineering work • Track and reduce MTTR, alert volume, incident recurrence and telemetry cost per service • Evaluate observability tooling and AI-assisted operations capabilities and provide recommendations • Maintain standards, runbooks, decision records and rationale for thresholds • Participate in a light-touch shared on-call rotation
• 6+ years in engineering, with 3+ years in a reliability, observability or production-operations role owning outcomes • Defined SLOs, SLIs and error budgets for real services • Experience reducing alert noise and rationalising alerting estates • Hands-on OpenTelemetry experience, including instrumenting services, running collectors, and making sampling and cardinality decisions • Experience managing telemetry cost and retention governance • Deep Azure Monitor, Log Analytics and Application Insights experience, including KQL • Strong Grafana skills, including dashboards, alerting and managing both as code • Infrastructure-as-code experience; Bicep is the standard, with strong Terraform experience transferable • Clear written and spoken communication at C1 level or above, or native-level business English • Experience selecting, implementing or migrating incident-management and on-call platforms (nice to have) • Prometheus-based monitoring and alerting, including exporters, recording rules and cardinality management (nice to have) • Kubernetes in production, ideally AKS (nice to have) • Software development background, ideally .NET (nice to have) • Familiarity with AI-assisted operations and coding tools (nice to have) • Azure certifications AZ-104, AZ-400 or AZ-305, or CKA (nice to have) • Experience in a small or scale-up environment where you were the observability function (nice to have)
• Guaranteed £2,000 pay rise every year, separate from merit or promotion increases • 26 days holiday plus public holidays, with the option to swap public holidays • Birthday and work anniversary day off every year • Up to 4 weeks per year working from anywhere in the world • Genuinely flexible, fully remote working • Generous Becoming a Parent leave • Medical scheme including remote GP, dental, optical, and diagnostics (UK employees) • Access to Support Room confidential counselling, therapy, and coaching • Help@Hand employee assistance programme • 5x salary life insurance through Unum • Up to 10% pension match (UK employees) • Annual learning and development budget per function • In-person team events a few times a year • Async-friendly working • Shared, light on-call rotation
Apply Now🕒 September 16
Senior DevOps Engineer owning Adapty's Proxmox, Ceph, Kubernetes, and cloud infrastructure. Supporting the SaaS platform powering mobile app subscriptions at global scale.
🇬🇧 United Kingdom – Remote
💰 $2M Seed Round - Adapty on 2022-03
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 September 10
Senior Reliability Engineer optimizing reliability, maintenance, and lifecycle asset management for JLL’s real estate services. Supporting a Life Sciences client across EMEA locations.
🕒 September 8
Site Reliability Engineer maintaining reliable, scalable production infrastructure for GitLab’s DevSecOps platform. Automating Kubernetes, cloud, observability, and incident-response operations across distributed Infrastructure Platforms teams.
🇬🇧 United Kingdom – Remote
💰 Secondary Market on 2020-11
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 September 8
Site Reliability Engineer improving availability, scalability and observability for Arbor’s school management platform. Supporting over 7,000 schools and trusts through reliable, resilient services.
🇬🇧 United Kingdom – Remote
💰 Private Equity Round on 2020-12
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🇬🇧 UK Skilled Worker Visa Sponsor
🕒 September 4
Senior DevOps Engineer managing Linux, AWS and production systems for Snappy Shopper’s rapid grocery delivery platform. Improving reliability, automation and incident response across UK Q-commerce infrastructure.