Observability Platform Engineer

Job not on LinkedIn

đŸ”„ 0 minutes ago

🇰🇿 Kazakhstan – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

đŸ—ïž Platform Engineer

đŸ‘» Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

‱ Design, build, and operate metrics, logging, distributed tracing, and alerting platform components ‱ Build high-volume, high-cardinality telemetry pipelines for large infrastructure fleets ‱ Define and implement SLO/SLI frameworks and alerting strategies ‱ Partner with service delivery and operations teams to build incident-focused observability ‱ Integrate observability tooling with incident management workflows, root-cause analysis, and post-incident reviews ‱ Improve detection speed and reduce MTTD and MTTR ‱ Contribute to AI-assisted operations tooling, including automated triage, anomaly detection, and engineer-assist tools ‱ Own the reliability, scalability, and security of the observability stack ‱ Document architecture, runbooks, and operational practices

🎯 Requirements

‱ Proven experience designing and building observability platforms for large-scale, production infrastructure environments ‱ Strong hands-on experience with metrics, logging, and distributed tracing tooling, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, and Jaeger/Tempo, or equivalents ‱ Experience with high-volume telemetry pipelines and tradeoffs involving cardinality, retention, cost, and query latency ‱ Strong software engineering skills in at least one relevant language, such as Go, Python, or Rust ‱ Experience with Kubernetes and cloud-native infrastructure ‱ Solid understanding of SLO/SLI/error-budget practices and low-noise alerting design ‱ Comfortable working in a fast-moving environment ‱ Strong communication skills and ability to work directly with operations/service delivery teams ‱ Preferred: observability for GPU/HPC or specialized high-performance compute environments ‱ Preferred: eBPF-based observability tooling ‱ Preferred: AIOps/ML-based anomaly detection or automated triage systems ‱ Preferred: managed services or MSP context ‱ Preferred: contributions to open-source observability projects

đŸ–ïž Benefits

‱ Professional development and training ‱ Attend conferences and working groups ‱ Company outings, happy hours, hackathons, and tech talks ‱ Competitive compensation package with a strong benefits plan ‱ Remote work arrangement

Apply Now