Senior Site Reliability Engineer, Golang, Kubernetes

đŸ”„ 12 hours ago

🌐 Kazakhstan, United States – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 20%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

‱ Define what reliability means for a GPU-accelerated AI platform and make it measurable ‱ Own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) ‱ Derive meaningful SLIs from signals emitted by the platform ‱ Expose SLIs to Platform Administrators through a clean API ‱ Define SLIs and SLOs across Kubernetes, bare-metal hosts, and NVIDIA infrastructure, including BMC, InfiniBand, NVLink, and UFM ‱ Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems ‱ Establish alerting and error-budget practices that maximize signal and minimize noise ‱ Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected ‱ Diagnose reliability and performance issues across the observability stack and drive their resolution ‱ Work across hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack ‱ Communicate reliability definitions effectively across teams

🎯 Requirements

‱ 5+ years in SRE, platform reliability, or a closely related software/infrastructure role ‱ Strong software engineering skills, such as Go or Python, with experience building and operating APIs or services in production ‱ Demonstrated experience defining SLIs/SLOs and error budgets for real production systems ‱ Hands-on experience with observability tooling, including metrics, logging, and tracing, such as Prometheus/VictoriaMetrics, OpenTelemetry, and Grafana ‱ Solid understanding of Kubernetes and the signals it and its workloads emit ‱ Strong written and verbal communication with technical audiences ‱ Preferred: Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM) ‱ Preferred: Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API ‱ Preferred: Familiarity with VictoriaMetrics/VictoriaLogs at scale ‱ Preferred: Proven experience in sovereign or high-security air-gapped environments

đŸ–ïž Benefits

‱ Professional development and training ‱ Attend conferences and working groups ‱ Company outings, happy hours, hackathons, and tech talks ‱ Competitive compensation package with a strong benefits plan ‱ Remote work arrangement

Apply Now

Similar Jobs

🕒 August 20

Emerging Travel Group

1001 - 5000

🏹 Hospitality

📩 Logistics

đŸ›ïž eCommerce

DevOps Engineers building resilient infrastructure for a large-scale international IT product. Managing databases, CI/CD, Kubernetes, monitoring, and production operations.

🇰🇿 Kazakhstan – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Consul

Docker

Grafana

Kafka

Kubernetes

Linux

Postgres

Prometheus

Python

Redis

Terraform

Go