Senior Site Reliability Engineer

🕒 August 21

🌐 United Kingdom, Germany, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🇬🇧 UK Skilled Worker Visa Sponsor

infoinfo

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Megaport

Megaport

201 - 500 employees

Founded 2013

📡 Telecommunications

Networking • Cloud Computing • Telecommunications

Megaport is a leading provider of global private connectivity solutions that enable simplified network interconnection. The company offers a platform for deploying secure, scalable, and agile networks that interconnect data centers, clouds, and virtual points of presence. Megaport's services allow users to create secure and dynamic network connections on-demand, without hardware or long-term contracts, offering flexibility and speed to businesses. By partnering with global service providers, data center operators, and systems integrators, Megaport ensures robust and widespread network access across 930+ locations in 25 countries. Its smart software tools and APIs allow for easy network management, making it a trusted choice for cloud networking and hybrid cloud solutions.

📋 Description

• Improve production reliability and system resilience within an SRE-scoped team • Champion high standards of work and industry best practices • Communicate with teams and stakeholders at all stages • Bring fresh ideas and encourage others • Solve complex technical problems • Work across numerous technologies in a fast-changing industry • Participate in on-call rotation, incident response, and blameless post-incident reviews • Write code, handle alerts, improve solutions, and support others • Engage stakeholders in requirements analysis and demonstrations • Ensure systems are secure, maintainable, and available • Continually evolve skills through peer reviews and research • Contribute to customer success and company goals

🎯 Requirements

• 5+ years administering Linux systems and related infrastructure in production environments • A collaborative SRE mindset, with familiarity around SLIs/SLOs/SLAs, error budgets, blast radius, and blameless postmortems • A focus on automation, reducing toil, and preventing problem recurrence • A track record of writing runbooks that work for the broader team, not just yourself • Strong Kubernetes and broader ecosystem fundamentals • Cloud infrastructure experience; AWS strongly preferred and bare-metal is a bonus • Strong tool development - Bash, plus either Python or Go preferred, or similar • Infrastructure-as-code tooling experience - Terraform preferred • CI/CD and version control, GitHub preferred • Database experience - one of Postgres, Cassandra, or ClickHouse preferred • Experience operating a production observability stack (metrics, logs, traces), with an eye for signal over noise • Comfortable working on live production infrastructure, with strong troubleshooting instincts and ownership of incident response • A history of continual professional development • A self-directed style suited to an async, globally distributed team, and comfortable picking up adjacent work when the situation calls for it

🏖️ Benefits

• Flexible working environment – a remote-first culture with coworking options available. • Generous leave plans – including 4 weeks of paid annual leave, parental leave, birthday leave, and a purchased annual leave program. • Health and wellness support – through a wellness allowance and employee wellbeing initiatives. • Comprehensive learning support – generous study and training allowance plus 5 days of paid study leave • Creative, modern workspaces – designed to inspire when you're not working remotely • Motivated, inclusive team – work alongside industry experts and fresh talent • Recognition programs – celebrate achievements with our *Legend* and *Kudos* awards

Apply Now

Similar Jobs

🕒 August 20

Trust Payments

501 - 1000

💳 Fintech

🤝 B2B

Senior DevOps Engineer maintaining secure, scalable AWS infrastructure and internal platforms. Supporting Trust Payments’ global fintech and payments solutions.

🇬🇧 United Kingdom – Remote

💰 Series unknown on 2013-12

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 19

Aker Systems

51 - 200

🤖 Artificial Intelligence

🔒 Cybersecurity

Lead DevOps Engineer building secure AWS cloud infrastructure and analytical data platforms for enterprise clients. Automating deployments, supporting production systems, and mentoring delivery teams.

🇬🇧 United Kingdom – Remote

💰 Venture Round on 2020-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 14

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Artificial Intelligence

Site Reliability Engineer maintaining CrowdStrike’s large-scale cybersecurity cloud platform. Automating operations, monitoring distributed systems, and leading incident response for reliable 24x7 service.

🕒 August 5

NICE

5001 - 10000

☁️ SaaS

🤖 Artificial Intelligence

📡 Telecommunications

Forward Deployed Engineer building AI agents and full-stack automation for NICE’s customer-experience software. Integrating conversational systems with enterprise platforms and deploying customer self-service solutions.

🕒 August 4

Pliant

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Engineering Manager building Pliant’s Site Reliability function for its B2B payments platform. Establishing SLOs, incident processes, observability, and a new reliability engineering team.

🇬🇧 United Kingdom – Remote

💰 $40M Series B - Pliant on 2025-04

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)