Site Reliability Engineer

🕒 March 24

🇩🇪 Germany – Remote

💵 €50k - €70k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 54%

infoinfo

🗣️🇩🇪 German Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Deepslate

Deepslate

11 - 50 employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Consulting • Healthcare • Insurance

Deepslate is a Europe-focused speech-to-speech Voice AI company that provides ultra-fast, GDPR-compliant conversational voice agents for enterprise customers. The company offers production-ready infrastructure with self-hosting or EU cloud options, low end-to-end latency (≈250–400 ms), multilingual support across 27+ languages, and advanced contextual reasoning for real-world customer interactions. Deepslate targets verticals like banking, insurance, retail, and healthcare to automate high-volume customer conversations while emphasizing data protection and cost-efficient scaling.

📋 Description

• Your mission is to build an infrastructure so resilient that potential outages are caught and mitigated before they even happen. • Design, build, and manage our cloud infrastructure using modern tools (Pulumi). • Orchestrate and optimize our Kubernetes clusters for complex, compute-heavy AI workloads. • Implement a flawless monitoring setup using Datadog and OpenTelemetry. • Establish and manage our on-call and alerting processes (using PagerDuty) and champion a culture of blameless post-mortems. • Build and maintain highly automated integration testing and deployment pipelines. • Define and monitor our service-level metrics, turning reliability into a measurable core component of our product development cycle. • Ruthlessly automate away toil to allow the engineering team to focus on innovation instead of maintenance. • Ensure our infrastructure is not only highly available but also secured against external threats.

🎯 Requirements

• Kubernetes: Deep, hands-on experience in setting up, managing, and scaling self-hosted Kubernetes clusters in production. • Infrastructure as Code: Strong experience with modern IaC, ideally with Pulumi (or deep Terraform knowledge alongside a willingness to adopt Pulumi). • Observability: You are a pro with Datadog and OpenTelemetry. • Alerting & Incident Management: Proven experience with PagerDuty (or similar tools). • Integration Testing & CI/CD: Hands-on experience setting up robust testing and deployment pipelines. • Fluent German (spoken and written). • Startup mindset: Comfortable navigating the ambiguity and rapid change of an early-stage codebase. • Extreme ownership: You take responsibility beyond processing tickets and drive reliability improvements end-to-end.

🏖️ Benefits

• Health insurance • Professional development

Apply Now

Similar Jobs

🕒 March 24

Reflow

11 - 50

☁️ SaaS

🏢 Enterprise

🤝 B2B

DevOps & Security Engineer evolving Reflow's cloud setup and strengthening security posture for a workflow intelligence platform.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Terraform

🕒 March 18

KLAR-GRUPPE

1 - 10

💼 Consulting

🎯 Recruiter

🤝 B2B

DevOps Engineer designing and automating stable IT systems for a healthcare company. Responsible for CI/CD pipelines and technical support in customer interactions.

🗣️🇩🇪 German Required

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Linux

Terraform

🕒 March 17

virtual7 GmbH

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior NixOS / DevOps Engineer supporting clients with modern, automated software infrastructure processes. Join our team at virtual7 to drive digitalization in the public sector.

🗣️🇩🇪 German Required

Docker

Haskell

Python

Rust

Go

🕒 February 6

PandaDoc

501 - 1000

☁️ SaaS

🤝 B2B

⚡ Productivity

Senior Site Reliability Engineer ensuring reliable service with minimal downtime at PandaDoc. Driving efforts in observability, incident management, and maintaining reliable operations.

AWS

Distributed Systems

Django

Grafana

Java

Kafka

Kubernetes

Postgres

Python

RabbitMQ

Spring

Spring Boot

SpringBoot

🕒 January 28

Famedly GmbH

11 - 50

🏥 Healthcare

💼 Consulting

📦 Logistics

Site Reliability Engineer for a healthcare startup improving medical communication. Designing SRE practices to enhance system reliability and performance.

🗣️🇩🇪 German Required

Kubernetes