Senior DevOps / Platform Engineer, AI Infrastructure

🔥 19 hours ago

🇩🇪 Germany – Remote

💵 €75k - €85k / year

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MAIA

MAIA

1 - 10 employees

Founded 2021

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Artificial Intelligence • SaaS • Enterprise

MAIA is a DSGVO-compliant AI-powered platform designed to enhance productivity and knowledge management within businesses. It serves as an intelligent assistant that enables users to efficiently manage, search, and interpret enterprise data while ensuring data privacy and security. Developed by Prodlane GmbH, MAIA offers seamless integration across various departments and supports multiple document formats for comprehensive, quick-response solutions.

📋 Description

• Take technical ownership of MAIA’s production infrastructure • Operate and evolve self-administered Linux systems across virtual machines and dedicated servers • Manage containers, networks, reverse proxies, and API gateways • Improve Infrastructure as Code, GitHub Actions workflows, deployment processes, automated checks, versioning, and rollback capabilities • Make deployments safe, repeatable, and easy for engineers to use • Operate and improve PostgreSQL in production, including performance analysis, connection pooling, capacity planning, backups, and tested restore procedures • Develop the self-hosted observability stack connecting metrics, logs, traces, and actionable alerts • Strengthen infrastructure security through IAM, least privilege, secrets management, TLS, vulnerability scanning, and patch management • Implement technical controls for ISO 27001 with continuously generated, understandable, and auditable evidence • Shape AI infrastructure by integrating and evaluating model and inference providers • Evaluate providers based on reliability, latency, throughput, cost, and operational effort • Explore self-hosted LLM inference, potentially including GPU infrastructure and vLLM • Improve incident response and operational resilience through root-cause investigation, runbooks, documentation, and lasting improvements • Improve internal developer experience by reducing manual work and creating clear interfaces and workflows • Make infrastructure and inference costs transparent and inform build, buy, and hosting decisions • Assess the existing platform, identify risks, explain options, and take improvements through to reliable production operation • Collaborate with software engineers, the CTO, company leadership, and distributed team members • Work as a hands-on senior individual contributor without people management

🎯 Requirements

• Several years of experience operating production SaaS systems on Linux servers administered directly by you or your team • Experience limited to fully managed hyperscaler services is not sufficient • Production experience with Docker and Docker Compose • Understanding of reverse proxies or API gateways such as Traefik or Kong • Practical experience operating PostgreSQL, including personally configured and tested backups and restores, performance analysis, and connection pooling • Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or comparable systems • Experience managing infrastructure through Terraform, Pulumi, or comparable Infrastructure as Code tooling • Experience operating an observability stack using Grafana, Loki, Prometheus, Sentry, or comparable technologies • Understanding of IAM, least privilege, secrets management, vulnerability scanning, and patching • Ability to design infrastructure reliable for customers and straightforward for engineers to use • Fluent English communication • German helpful but not required • Understanding of how RAG systems work and where they commonly fail in production • Ability to discuss LLM serving trade-offs, including latency, throughput, reliability, cost, and operational complexity • Familiarity with models, inference providers, and the wider GenAI market • Strong platform engineering foundation and technical curiosity to develop deeper AI infrastructure expertise • Ability to independently operate business-critical production systems • Ability to identify and prioritise infrastructure work without detailed instructions • Ability to lead technical improvements across team boundaries • Reliable written and remote communication • Accountability after changes are deployed • Helpful but not required: experience with Hetzner, NixOS, self-hosted Supabase, ISO 27001 controls and audit evidence, SRE practices, GPUs, or LLM serving technologies such as vLLM

🏖️ Benefits

• €75,000–€85,000 gross annual salary, depending on experience and scope • Opportunity to participate in our VSOP • Permanent, full-time position • Flexible working hours • Fully remote work from anywhere within Germany • Regular opportunities to meet and work with the team in Leipzig, with travel and accommodation covered • Access to a WellPass fitness membership

Apply Now

Similar Jobs

🕒 September 3

DATAGROUP

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

AI Platform Engineer building DATAGROUP’s sovereign LLM platform, RAG services, and cloud-native infrastructure. Operating Kubernetes, CI/CD, Linux, databases, and AI services.

🗣️🇩🇪 German Required

Azure

Cloud

Docker

Kubernetes

Linux

Postgres

Python

🕒 September 3

evoila

201 - 500

💼 Consulting

Senior Kubernetes Platform Engineer consulting on enterprise Kubernetes and cloud-native platforms for evoila Germany GmbH. Designing, implementing, and modernizing scalable DevOps environments.

🗣️🇩🇪 German Required

Cloud

Flux

Kubernetes

Linux

OpenShift

Terraform

VMware

🕒 September 2

Interval Group

51 - 200

📣 Marketing

⚖️ Legal

📦 Logistics

Lead Platform Engineer rebuilding a leading e-commerce and technology company’s sovereign OpenStack cloud on bare-metal hardware. Automating infrastructure with IaC, AIOps, and self-healing systems.

🗣️🇩🇪 German Required

Ansible

Cloud

JavaScript

Linux

OpenStack

Python

React

Terraform

VMware

Vue.js

🕒 August 29

Deluxe

5001 - 10000

💼 Consulting

📣 Marketing

📦 Logistics

AI Platform Engineer operating and automating production AI services for Deluxe. Improving model deployment, observability, reliability, and delivery across global media workflows.

Airflow

AWS

Cloud

Distributed Systems

Docker

Kubernetes

Python

Ray

🕒 August 18

MTG AG

51 - 200

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

AI Engineer building secure AI/LLM platforms for MTG AG, a German encryption and cybersecurity specialist. Developing AI gateways, RAG solutions, model routing, and observability for enterprise use.

🗣️🇩🇪 German Required

Cloud