Senior Site Reliability Engineer – m/f/d

🔥 14 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Flip

Flip

51 - 200 employees

Founded 2018

👥 HR Tech

☁️ SaaS

🤖 Artificial Intelligence

HR Tech • SaaS • Artificial Intelligence

Flip is the AI employee experience platform for frontline and deskless workers. The platform unites internal communications, HR self-service (payslips, time-off, onboarding), AI-native intranet, AI workflows, and frontline identity management into a single, mobile-first app that works on any device — including shared phones without corporate email. Flip includes an AI layer (Flip Intelligence / Ask AI), an app builder and integrations with existing tools (SharePoint, Teams, HR systems), and targets industries with large deskless workforces. Customers cited include Ben & Jerry’s, Bosch, KFC and PENNY, demonstrating use across retail, food service, manufacturing and logistics.

📋 Description

• Co-own the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid growth across the globe. • Drive the resilience strategy: Define how we approach global scaling, zero-downtime deployments, rollback mechanisms and disaster recovery, and make sure the platform stays available around the clock. • Evolve our observability stack: Improve our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can trust. • Improve our IaC Platform: Eliminate toil at the source, and make our infrastructure truly self-service for engineering teams. • Lead in incidents: Take a leading role in platform-related major incidents, drive blameless post-mortems for the squad, and translate findings into systemic improvements. • Mentor within the squad: Coach teammates, run RFCs and design reviews inside the team, and help engineers grow into stronger SREs. • Shape our roadmap: Partner with your squad to define the platform's direction.

🎯 Requirements

• 5+ years of hands-on experience as a Site Reliability Engineer (SRE), Platform Engineer, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Backend Engineer with a strong infrastructure focus. • Proven track record building and operating **high-throughput, highly available systems** in production. • Deep, production-level experience with **Kubernetes** on any Hyperscaler. • Strong experience with modern observability stacks (e.g. Prometheus, Mimir, VictoriaMetrics, Dash0, Loki, ELK) and a clear point of view on SLIs, SLOs and error budgets. • Solid software development skills in **Go** (strongly preferred, since our IaC runs on Pulumi in Go) or Python. • Hands-on experience with Infrastructure as Code (Pulumi, OpenTofu, Terraform) and **GitOps** (e.g. ArgoCD) **+ CI/CD pipeline design**. • Demonstrated ability to lead complex infrastructure initiatives from design to production - including writing RFCs and driving architecture decisions within your team. • Experience mentoring engineers and raising the technical bar within a team. • Comfortable owning major incidents end-to-end and turning learnings into systemic change. • Strong communication skills and business-fluent English. • Willingness to participate in on-call rotations to ensure the reliability of our platform.

🏖️ Benefits

• Work mode: We’re remote-first, giving you flexibility to work from home. At the same time, we deeply value the power of in-person collaboration. Depending on the role, you’ll join occasional team events, workshops, or meetings in our Berlin or Stuttgart offices - always with plenty of notice. The exact balance will be discussed during your interview. • Work-Life-Balance: We don't want you to grow roots to your desk chair. That's why we cover the costs of your E-Gym-Wellpass membership and offer job bike leasing. • Celebrating success: Expect highly motivated and committed people in a relaxed working atmosphere. • Be part of something bigger: You actively shape Flip in your role. Along the way, you are an enabler of the rapid growth process of a young tech company and grow towards your goals, fun is guaranteed. • Happy to be a Flipster: Stay tuned for regular team events and culture days that bring us together as Flipsters. • Working abroad: At Flip you can also work abroad in the European Union. Let's talk about remote work in the interview.

Apply Now

Similar Jobs

🔥 3 hours ago

Qdrant

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

Senior Software Engineer developing and operating cloud-native infrastructure at Qdrant. Designing, building, and optimizing solutions for AI-driven applications across global teams.

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Grafana

Kubernetes

Microservices

Prometheus

Go

🔥 14 hours ago

Codesphere

51 - 200

☁️ SaaS

🏢 Enterprise

🏛️ Government

Site Reliability Engineer at Codesphere ensuring cloud infrastructure reliability and efficiency. Involves collaboration with development teams and maintaining production systems for continuous delivery.

Ansible

Cloud

Kubernetes

Terraform

Go

🔥 16 hours ago

SimScale

51 - 200

☁️ SaaS

🤝 B2B

🤖 Artificial Intelligence

Senior SRE / Platform Engineer at SimScale responsible for managing and improving cloud infrastructure. Focus on AWS, EKS, observability, disaster recovery, and security controls.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Java

Kubernetes

Linux

Prometheus

Python

Rust

Terraform

Go

🔥 18 hours ago

mittwald

51 - 200

🤝 B2B

☁️ SaaS

Senior DevOps Engineer developing and managing a Kubernetes-based hosting platform at Mittwald. Ensuring system stability, performance, scalability, and making informed technological decisions.

🗣️🇩🇪 German Required

Kubernetes

Linux

Python

Go

🔥 18 hours ago

mobilistics GmbH

11 - 50

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

DevOps Engineer specializing in Kubernetes cluster design and management. Working remotely in a dynamic team within cloud-native technology environments.

🗣️🇩🇪 German Required

Ansible

Cloud

Kubernetes

Linux