Observability Platform Engineer

đŸ”„ 0 minutes ago

🇰🇿 Kazakhstan – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

đŸ—ïž Platform Engineer

đŸ‘» Ghost score 19%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

‱ Design, build, and operate metrics, logging, distributed tracing, and alerting platform components ‱ Build high-volume, high-cardinality telemetry pipelines for large infrastructure fleets ‱ Optimize telemetry pipelines for cost, retention, and query performance ‱ Define and implement SLO/SLI frameworks and alerting strategies that reduce noise ‱ Partner with service delivery and operations teams to build incident-focused observability capabilities ‱ Integrate observability tooling with incident management workflows, root-cause analysis, and post-incident reviews ‱ Improve detection speed and reduce MTTD and MTTR across the platform ‱ Contribute to AI-assisted operations tooling, including automated triage, anomaly detection, and engineer-assist tools ‱ Own the reliability, scalability, and security of the observability stack ‱ Document architecture, runbooks, and operational practices

🎯 Requirements

‱ Proven experience designing and building observability platforms for large-scale, production infrastructure environments ‱ Strong hands-on experience with metrics, logging, and distributed tracing tooling, such as Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or equivalents ‱ Experience with high-volume telemetry pipelines and tradeoffs involving cardinality, retention, cost, and query latency ‱ Strong software engineering skills in at least one language commonly used in this space, such as Go, Python, or Rust ‱ Experience with Kubernetes and cloud-native infrastructure ‱ Solid understanding of SLO/SLI/error-budget practices and low-noise alerting design ‱ Comfortable working in a fast-moving environment where the platform is built alongside the infrastructure it monitors ‱ Strong communication skills and ability to work directly with operations/service delivery teams ‱ Experience building observability for GPU/HPC or other specialized high-performance compute environments ‱ Experience with eBPF-based observability tooling ‱ Familiarity with AIOps/ML-based anomaly detection or automated triage systems ‱ Experience in a managed services or MSP context ‱ Contributions to open-source observability projects

đŸ–ïž Benefits

‱ Work with an established Silicon Valley leader in the cloud infrastructure industry ‱ Work with exceptionally passionate, talented and engaging colleagues ‱ Help Fortune 500 and Global 2000 customers implement next-generation cloud technologies ‱ Be part of cutting-edge, open-source innovation ‱ Professional development and training ‱ Attend conferences and working groups ‱ Company outings, happy hours, hackathons, and tech talks ‱ Competitive compensation package with a strong benefits plan

Apply Now