Senior Reliability Engineer

🔥 12 hours ago

🇷🇴 Romania – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Betfair Romania Development

Betfair Romania Development

1001 - 5000 employees

💼 Consulting

📣 Marketing

🏥 Healthcare

Consulting • Marketing • Healthcare

Betfair Romania Development is the largest technology hub of Flutter Entertainment Plc, an FTSE 100 company, with over 1,900 people powering the world’s leading betting and gaming brands. With a presence in Cluj-Napoca, they offer a diverse portfolio of proprietary brands including Betfair, PokerStars, and Paddy Power, serving over 18 million customers globally. The company specializes in software development, online gaming, and customer support, and is dedicated to providing immersive and safe experiences in the realms of sports and casino betting.

📋 Description

• Partner with application and platform teams to understand service architecture, dependencies, critical customer journeys, and reliability risks • Define meaningful SLIs and SLOs connecting technical service behaviour with customer and business outcomes • Support adoption of error budgets and connect reliability performance to engineering priorities • Assess services against reliability and production-readiness expectations across observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery • Use reliability maturity assessments and scorecards to prioritise impactful improvements • Support incident response and investigate complex production issues using logs, metrics, traces, dependencies, and deployment information • Review incidents and recurring operational issues to identify systemic risks and longer-term improvements • Improve runbooks, alerting, escalation paths, and operational practices • Design automation and self-service capabilities to reduce engineering toil • Build tooling and automation for reliability workflows, operational readiness, investigation, and remediation • Partner with Observability Engineering on telemetry needed to measure reliability and troubleshoot issues • Translate performance tests, capacity assessments, Gamedays, chaos experiments, and failover testing into reliability improvements • Support peak-event readiness by reviewing service health, SLOs, dependencies, capacity risks, and reliability findings • Improve reliability of critical customer journeys through monitoring, dependency understanding, failure handling, and recovery practices • Contribute to Reliability Engineering standards, patterns, documentation, and reusable golden paths • Integrate reliability capabilities into CI/CD and developer workflows, including SLO-as-Code, telemetry validation, production-readiness checks, and automated operational controls • Use AI-assisted investigation and automation to accelerate troubleshooting and reduce manual effort • Share knowledge and support engineers in developing stronger SRE and production-engineering practices

🎯 Requirements

• Strong hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering • Good understanding of core SRE principles, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation • Experience defining or working with SLIs and SLOs for production services • Strong production troubleshooting skills across applications, infrastructure, networks, databases, and service dependencies • Good understanding of distributed systems concepts and reliability patterns such as retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation • Experience with observability platforms such as Datadog or equivalent, including logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring • Experience participating in incident response, post-incident reviews, and meaningful follow-up actions • Hands-on experience with Kubernetes and cloud infrastructure, preferably AWS • Experience with infrastructure-as-code tools such as Terraform and modern CI/CD environments • Strong automation and software engineering skills, with proficiency in at least one modern programming language such as Go, Java, Python, or JavaScript • Experience replacing repetitive operational activities with automation or self-service • Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering • Ability to understand application architecture and identify reliability risks across service and infrastructure dependencies • Strong analytical and problem-solving skills • Good communication skills and ability to explain reliability concepts and recommendations to engineers and engineering leadership • Ability to work across multiple teams, balance competing priorities, and drive work through to measurable outcomes • Mindset focused on automation, continuous improvement, knowledge sharing, and solving systemic problems

🏖️ Benefits

• Hybrid & remote working options • €1,000 per year for self-development • Company share scheme • 25 days of annual leave per year • 20 days per year to work abroad • 5 personal days/year • Flexible benefits: travel, sports, hobbies • Extended health, dental and travel insurances • Customized well-being programmes • Career growth sessions • Thousands of online courses through Udemy • A variety of engaging office events

Apply Now

Similar Jobs

🕒 6 days ago

accesa.eu

1001 - 5000

💼 Consulting

🏭 Manufacturing

🏥 Healthcare

Senior DevOps Engineer automating Azure cloud platforms, APIs, and identity workflows for Accesa’s financial-services customers. Improving secure, reliable banking operations through DevOps practices.

Azure

Cloud

Docker

Java

Kubernetes

Linux

OpenShift

Python

SQL

🕒 September 10

eMAG

5001 - 10000

🛍️ eCommerce

🏪 Marketplace

DevOps Engineer securing eMAG’s multi-cloud infrastructure and Kubernetes platforms. Automating cloud governance, CI/CD, identity management, and production operations.

AWS

Cloud

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Python

Terraform

Go

🕒 September 8

Spyrosoft

1001 - 5000

🚘 Automotive

🏥 Healthcare

💼 Consulting

Senior DevOps Engineer improving Kubernetes, cloud infrastructure, and CI/CD for Spyrosoft’s German engineering client. Fully remote within the EU with occasional Nuremberg travel.

🗣️🇩🇪 German Required

Azure

Cloud

Kubernetes

Terraform

🕒 September 4

Nelnet

5001 - 10000

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer managing Azure infrastructure, releases, CI/CD pipelines, and test environments. Supporting Nelnet’s student lending, payments, and education services.

Angular

Azure

Cloud

DNS

Docker

Firewalls

Java

JavaScript

Kubernetes

Linux

Node.js

NoSQL

Python

Redis

SDLC

Shell Scripting

SQL

TCP/IP

.NET

🕒 September 3

Tether.to

11 - 50

₿ Crypto

💳 Fintech

💸 Finance

DevOps Engineer building secure CI/CD, IaC, and release pipelines for Tether’s blockchain-powered digital finance platform. Automating cross-platform deployments across web, desktop, and mobile.

Android

Ansible

AWS

Distributed Systems

Docker

Firewalls

Grafana

iOS

JavaScript

Linux

Microservices

Prometheus

PyTorch

Shell Scripting

Tensorflow

Terraform

TypeScript

C++