Observability Platform Engineer

đŸ”„ 0 minutes ago

🌐 Kazakhstan, Latvia – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

đŸ—ïž Platform Engineer

đŸ‘» Ghost score 19%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

‱ Design, build, and operate observability platform components for metrics, logging, distributed tracing, and alerting ‱ Build telemetry pipelines for high-cardinality, high-volume data from large infrastructure fleets ‱ Optimize telemetry cost, retention, and query performance ‱ Define and implement SLO/SLI frameworks and alerting strategies ‱ Partner with service delivery and operations teams to support incident response needs ‱ Integrate observability tooling with incident management workflows, root-cause analysis, and post-incident reviews ‱ Improve detection speed and reduce MTTD and MTTR ‱ Contribute to AI-assisted operations tooling, including automated triage, anomaly detection, and engineer-assist tools ‱ Own the reliability, scalability, and security of the observability stack ‱ Document architecture, runbooks, and operational practices

🎯 Requirements

‱ Proven experience designing and building observability platforms for large-scale, production infrastructure environments ‱ Strong hands-on experience with metrics, logging, and distributed tracing tooling, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or equivalents ‱ Experience with high-volume telemetry pipelines and tradeoffs involving cardinality, retention, cost, and query latency ‱ Strong software engineering skills in at least one commonly used language, such as Go, Python, or Rust ‱ Experience with Kubernetes and cloud-native infrastructure ‱ Solid understanding of SLO/SLI/error-budget practices and low-noise alerting design ‱ Comfortable working in a fast-moving environment ‱ Strong communication skills and ability to work directly with operations/service delivery teams ‱ Experience building observability for GPU/HPC or other specialized, high-performance compute environments ‱ Experience with eBPF-based observability tooling ‱ Familiarity with AIOps/ML-based anomaly detection or automated triage systems ‱ Experience in a managed services or MSP context ‱ Contributions to open-source observability projects

đŸ–ïž Benefits

‱ Work with an established Silicon Valley leader in the cloud infrastructure industry ‱ Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies ‱ Be a part of cutting-edge, open-source innovation ‱ Professional development and training ‱ Attend conferences and working groups ‱ Company outings, happy hours, hackathons, and tech talks ‱ Competitive compensation package with a strong benefits plan ‱ Remote work arrangement

Apply Now