
501 - 1000 employees
đŒ Consulting
đ„ Healthcare
đŠ Logistics
Consulting âą Healthcare âą Logistics
Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.
đ„ 1 hour ago
đ Kazakhstan, United States â Remote
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ» Ghost score 19%
Improve your chances of getting an interview by checking your resume score before you apply.

501 - 1000 employees
đŒ Consulting
đ„ Healthcare
đŠ Logistics
Consulting âą Healthcare âą Logistics
Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.
âą Define what reliability means for a GPU-accelerated AI platform and make it measurable âą Own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) âą Derive meaningful SLIs from signals emitted by the platform âą Expose SLIs to Platform Administrators through a clean API âą Define SLIs and SLOs across Kubernetes, bare-metal hosts, and NVIDIA infrastructure, including BMC, InfiniBand, NVLink, and UFM âą Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems âą Establish alerting and error-budget practices that maximize signal and minimize noise âą Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected âą Diagnose reliability and performance issues across the observability stack and drive their resolution âą Work across hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack âą Communicate reliability definitions effectively across teams
âą 5+ years in SRE, platform reliability, or a closely related software/infrastructure role âą Strong software engineering skills, such as Go or Python, with experience building and operating APIs or services in production âą Demonstrated experience defining SLIs/SLOs and error budgets for real production systems âą Hands-on experience with observability tooling, including metrics, logging, and tracing, such as Prometheus/VictoriaMetrics, OpenTelemetry, and Grafana âą Solid understanding of Kubernetes and the signals it and its workloads emit âą Strong written and verbal communication with technical audiences âą Preferred: Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM) âą Preferred: Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API âą Preferred: Familiarity with VictoriaMetrics/VictoriaLogs at scale âą Preferred: Proven experience in sovereign or high-security air-gapped environments
âą Professional development and training âą Attend conferences and working groups âą Company outings, happy hours, hackathons, and tech talks âą Competitive compensation package with a strong benefits plan âą Remote work arrangement
Apply Nowđ August 20
DevOps Engineers building resilient infrastructure for a large-scale international IT product. Managing databases, CI/CD, Kubernetes, monitoring, and production operations.
đ°đż Kazakhstan â Remote
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)