Senior Platform, Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇩🇪 Germany – Remote

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Contabo

Contabo

201 - 500 employees

Founded 2003

☁️ SaaS

SaaS

Contabo is a provider of cloud hosting services offering a range of virtual private servers (VPS), dedicated servers, and bare metal servers with a focus on affordability and performance. Known for its fast provisioning and award-winning customer support, Contabo operates globally with data centers in 9 regions and 12 locations. The company emphasizes German quality since 2003 and provides solutions for individuals and businesses seeking efficient and cost-effective cloud computing resources. Notable features include the Contabo API for programmatic access to resources, custom image deployments, and strong emphasis on security.

📋 Description

• Take architectural ownership of shared infrastructure services • Own API gateway and ingress, persistent storage, secrets and identity infrastructure, edge security/WAF, and observability services • Investigate partially under-documented systems and remediate known weaknesses • Evaluate aging components using build-versus-buy reasoning • Decide which legacy components should be fixed, replaced, or retired • Design and establish an on-call process • Create runbooks for platform and infrastructure incidents • Mature observability practices with distributed tracing, SLOs/SLIs, and meaningful dashboards • Operate and build on Prometheus, Grafana, Alloy, and OpenTelemetry • Contribute to system design reviews and apply infrastructure fundamentals to services • Advise multiple teams on shared infrastructure services • Consistently document and distribute technical knowledge • Reduce platform-wide single points of failure and resolve known risks • Travel occasionally for datacenter visits and team offsites

🎯 Requirements

• 7+ years of experience in platform, infrastructure, or SRE roles • Ideally, end-to-end responsibility for a private cloud or IaaS platform • Hands-on production experience with distributed storage systems, especially Ceph • Experience with Kubernetes persistent storage, Longhorn or comparable • Experience operating API gateways and ingress, such as Kong or Nginx Ingress • Experience debugging CORS and rate limiting • Experience with CDN/edge security, DDoS mitigation, WAF configuration, and firewall-rule design • Strong grounding in load balancing, caching, sharding/replication, consistency models, and message queues • Experience designing or maturing observability with tracing, metrics, and SLOs/SLIs is desirable • Experience with Vault, Keycloak, NATS, on-call processes, and incident runbooks is a plus • Composure working with grown, incompletely documented systems • Ability to act as a technical go-to person across teams and share/document knowledge • Professional fluency in English • German language skills, CKA/CKS or Ceph certifications, and Proxmox/OpenStack experience are a plus

🏖️ Benefits

• Remote or hybrid with flexible hours • Same technical equipment at home as in the office • Workation across the EU and in the summer office in Mallorca • Access to EGYM Wellpass and thousands of fitness and wellness facilities • Discounts on many products and services through the corporate benefits program • 30 days of vacation • Additional days off on Christmas Eve and New Year’s Eve • Extra day off for social activities through Volunteer Day • Individual professional and personal development opportunities • Modern, conveniently located offices across European locations • International, diverse work environment • Company events and team activities • Ownership and creative freedom • Open feedback culture

Apply Now

Similar Jobs

🕒 3 days ago

Recare

51 - 200

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Senior platform engineer operating AWS, Kubernetes, and SageMaker AI infrastructure for Recare, a German healthcare SaaS company. Ensuring reliable, scalable production workloads in regulated healthcare.

AWS

Cloud

DynamoDB

Kubernetes

Postgres

Python

🕒 4 days ago

Mondo

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Platform Engineer securing Mondoo’s cloud infrastructure and policy-as-code platform. Automating controls across Kubernetes, AWS, Azure, GCP, and on-premises environments.

Ansible

AWS

Azure

Chef

Cloud

DNS

Docker

Google Cloud Platform

Kubernetes

Puppet

Python

TCP/IP

Terraform

🕒 5 days ago

MAIA

1 - 10

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Senior DevOps engineer owning Linux, cloud, security, and observability infrastructure for MAIA’s enterprise AI platform. Shaping reliable production operations and LLM infrastructure.

Docker

Grafana

Linux

Postgres

Prometheus

Terraform

🕒 September 3

DATAGROUP

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

AI Platform Engineer building DATAGROUP’s sovereign LLM platform, RAG services, and cloud-native infrastructure. Operating Kubernetes, CI/CD, Linux, databases, and AI services.

🗣️🇩🇪 German Required

Azure

Cloud

Docker

Kubernetes

Linux

Postgres

Python

🕒 September 3

evoila

201 - 500

💼 Consulting

Senior Kubernetes Platform Engineer consulting on enterprise Kubernetes and cloud-native platforms for evoila Germany GmbH. Designing, implementing, and modernizing scalable DevOps environments.

🗣️🇩🇪 German Required

Cloud

Flux

Kubernetes

Linux

OpenShift

Terraform

VMware