Expert DevOps – Observability Operations

Job not on LinkedIn

🔥 0 minutes ago

🗣️🇩🇪 German Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of C4 Group

C4 Group

11 - 50 employees

Founded 2008

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

C4 Group is an IT and management consulting firm that provides recruitment, project and interim management, M&A and post-merger-integration, DevOps & cloud operations, business intelligence and data analytics, SAP/S4 migration, and security & infrastructure services. Founded in 2008, it works primarily with energy, pharmaceutical/life-sciences, and financial-services clients to deliver tailored business, IT and M&A solutions, leveraging strategic partnerships and a consultant network to implement digital transformation projects including AI/ML, RPA, predictive maintenance and energy-optimization use cases.

📋 Description

• Execute Kubernetes platform changes, upgrades, maintenance, and restoration measures according to approved procedures, with implementation records and execution documentation • Configure, enhance, and maintain observability capabilities including metrics, logs, distributed tracing, health checks, dashboards, and alerting rules • Analyze operational anomalies and incidents; produce root cause analyses, restoration recommendations, and improvement requirements • Create and maintain operational playbooks, standard operating procedures, incident guides, escalation criteria, and knowledge artifacts • Configure and deliver platform observability and security components with implementation documentation and configuration records • Create and maintain operational documentation, decision records, handover packages, and an Operational Manual for supported environments • Develop and deliver operational governance guidelines, training materials, and knowledge-transfer packages • Prepare the application for productive operation on new infrastructure • Establish monitoring, alerting, support, automation, and incident-management processes • Support stable, scalable, maintainable operations and enable operational ownership through manuals and playbooks

🎯 Requirements

• 8+ years of DevOps experience working with containerized Java application services and servers • Profound knowledge of the Grafana LGTM stack: Grafana, Prometheus, Loki, Mimir, and Tempo • Profound knowledge of Kubernetes infrastructure • Profound knowledge of ArgoCD or equivalent continuous deployment tools • Profound knowledge and experience with incident and problem management • Shell scripting experience with Windows PowerShell and Linux • Experience with Splunk • English language skills at C1 level • German language skills at C1 level • Profound understanding of deployment concepts for complex microservice-oriented Java application platforms • Experience developing applications • Experience with L2 support of complex business applications • Experience creating and maintaining operations playbooks • Experience in IT provider management, IT monitoring, and IT operations management • Experience with Kafka, MySQL, PostgreSQL, SPARK, Harness, Helm, JFrog Artifactory, Azure DevOps, and OpenTelemetry

Apply Now