Senior Site Reliability Engineer, SRE

🕒 July 9

🏢🏡 Berlin – Hybrid

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 22%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 1GLOBAL

1GLOBAL

WebsiteLinkedIn

501 - 1000 employees

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

1GLOBAL is a leading telecommunications company dedicated to creating better connections between people, devices, and businesses through innovative mobile solutions. Based in the UK, the company specializes in eSIM technology, internet telephony, and global communication services. By investing significantly in its infrastructure, 1GLOBAL aims to revolutionize connectivity on a global scale, enabling seamless interactions across international enterprises, IoT device makers, and telecommunications operators. Through its unique offerings, 1GLOBAL connects billions of individuals and devices, fostering a world where communication is improved and more efficient.

📋 Description

• Act as a senior technical contributor within the SRE team, mentoring peers and setting the technical bar for reliability engineering. • Define, measure, and maintain SLIs and SLOs for core infrastructure and customer-facing services. • Plan and execute redundancy and resilience testing across service, infrastructure, and networking layers — validating failover, HA configurations, and disaster recovery readiness. • Design and implement automated recovery mechanisms, self-healing workflows, and intelligent alerting systems. • Drive incident response, root-cause analysis, and blameless post-mortems, and ensure implementation and tracking of corrective and preventive actions derived from them to achieve continuous improvement. • Develop and enhance observability (metrics, logs, traces) using Prometheus, Grafana, Loki, and OpenTelemetry. • Partner with Infrastructure and DevOps teams to ensure deployment safety, rollback policies, and configuration consistency. • Proactively identify weaknesses through fault-injection, load, and chaos testing. • Continuously reduce operational toil through automation and reliability tooling. • Contribute to on-call practices, improving alert quality, runbooks, escalation procedures, and incident management processes. • Perform capacity planning, performance benchmarking, and resilience audits across systems. • Ensure compliance with security, reliability, and availability standards. • Create and maintain internal documentation, playbooks, and operational guidelines for peers and users. • Contribute to cloud cost-optimization initiatives, including reserved capacity planning, autoscaling design, storage tiering, workload right-sizing, and continuous anomaly detection.

🎯 Requirements

• A minimum of 5 years of experience in Site Reliability, Systems, or Infrastructure Engineering (including 2+ years in a dedicated SRE role). • Strong expertise in Linux systems engineering, distributed systems, and networking. • Proven experience building and running high-availability, mission-critical production systems. • Hands-on experience with redundancy and failover testing, disaster recovery, and high-availability architecture validation. • Deep understanding of monitoring, observability, and incident management principles. • Experience with Prometheus, Grafana, Loki, Thanos, and OpenTelemetry or similar tools. • Proficiency in Python, Go, and Bash for automation and reliability tooling. • Strong knowledge of Kubernetes, container orchestration, and service mesh architectures. • Experience with AWS (EKS, EC2, VPC) and on-premises infrastructure integration. • Proficiency in Infrastructure as Code tools such as Terraform. • Understanding of networking fundamentals (routing, load balancing, BGP, DNS, VXLAN, etc.). • Excellent analytical and problem-solving skills, capable of operating under pressure. • Strong communication and collaboration skills across distributed and cross-functional teams.

🏖️ Benefits

• Growth Opportunities: Advance your career in one of the fastest growing telecommunications companies, expanding over 100% year-on-year under the leadership of successful tech entrepreneurs. • Major Transaction Exposure: Be in the driver’s seat for transactions that will have an impact on the future telco industry. • Work with a Talented Team: From the Board and the Founders to the Senior Management Team, you will collaborate daily with the most capable and renowned external advisors, and constantly being exposed to talented and driven individuals. • Dynamic Work Environment: Thrive in a collaborative, fast-paced workplace where innovation is encouraged, and every contribution counts. • Professional Development: Work alongside industry experts to enhance your skills and knowledge in a cutting-edge field. • International Experience: Gain opportunities to work in different 1GLOBAL offices around the world as you grow within the company. • Open Communication Culture: Join a team where your ideas are heard, and open dialogue is encouraged, fostering a supportive and transparent work environment. • Get Things Done Attitude: Be part of a results-driven team that values efficiency, creativity, and the drive to make a tangible impact in the industry.

Apply Now

Similar Jobs

🕒 June 24

PROMOS consult

201 - 500

💼 Consulting

📦 Logistics

🏥 Healthcare

WebsiteLinkedIn

DevOps Engineer responsible for creating optimal conditions for modern software development. Automating build, test, and deployment processes while collaborating with development and infrastructure teams.

🗣️🇩🇪 German Required

Docker

Grafana

Kubernetes

OpenShift

Prometheus

🕒 June 19

Kittl

11 - 50

📣 Marketing

💼 Consulting

📦 Logistics

WebsiteLinkedIn

Senior DevOps Engineer on Kittl’s Platform team managing Kubernetes infrastructure and improving CI/CD pipelines. Collaborating across teams to enhance operational excellence in an AI-driven environment.

🏢🏡 Berlin – Hybrid

💰 Series A on 2023-01

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Distributed Systems

Grafana

Kubernetes

Terraform

TypeScript

Vault

🕒 March 28

Alexander Thamm [at]

201 - 500

🤖 Artificial Intelligence

💼 Consulting

WebsiteLinkedIn

(Senior) DevOps Engineer specializing in ML solutions implementation and management in Germany. Focused on CI/CD pipelines, automation, and cloud services.

🗣️🇩🇪 German Required

AWS

Azure

Cloud

Kubernetes

Python

Terraform

🕒 March 26

Stackmeister

11 - 50

💼 Consulting

📣 Marketing

☁️ SaaS

WebsiteLinkedIn

DevOps Engineer responsible for implementing software integration and cloud architectures. Involves continuous delivery and deployment processes with technologies like AWS, Azure, and Kubernetes.

🗣️🇩🇪 German Required

AWS

Azure

Cloud

Docker

Jenkins

Kubernetes

Terraform

🕒 January 16

M2. technology & project consulting GmbH

11 - 50

💼 Consulting

📦 Logistics

🏢 Enterprise

WebsiteLinkedIn

DevOps Engineer focusing on client requirements and cloud architecture. Engaging in migrations and advising on complex data infrastructures.

🗣️🇩🇪 German Required

Ansible

AWS

Azure

Cloud

DNS

Google Cloud Platform

Jenkins

Linux

NoSQL

Tableau

Terraform