Senior Site Reliability Engineer, SRE

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

• Contribute to the design, development, and operation of cloud-based AI solutions built on the CNCF ecosystem, including Kubernetes • Deploy AI infrastructure built on NVIDIA-certified hardware according to engineering architecture and implementation designs • Ensure the reliability, security, and performance of container infrastructure • Mentor team members and Mirantis customers • Work with geographically distributed international teams on technical challenges and process improvements • Develop, implement, maintain, and troubleshoot cloud and AI infrastructure based on open source software • Collaborate with stakeholders to gather and refine technical requirements • Optimize system performance, reliability, and scalability • Troubleshoot, debug, and resolve complex technical issues • Participate in code reviews • Stay current with cloud operations and development trends and best practices • Design and implement AI-driven automation across the DevOps lifecycle • Facilitate knowledge transfer to customers during delivery phases • Define technical strategies and ensure seamless integration of cloud and software services

🎯 Requirements

• 5+ years of professional experience in DevOps, focused on cloud and infrastructure technologies • Experience with Kubernetes and/or OpenStack • Experience with high-performance data center processing, networking, and storage • Exposure to Golang and working knowledge of Python, JavaScript, or other programming languages • Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines • Exceptional problem-solving and debugging skills across networking, storage, Linux, and Kubernetes • Knowledge of performance optimization and security • Ability to lead technical tasks and collaborate with diverse teams • Ability to make independent judgment calls when working directly with customers • Excellent written and spoken English • Excellent customer-facing communication skills • Commitment to innovation, continuous learning, and high-quality results • Ability to travel up to 25%, including internationally • Bachelor's degree in Computer Science or a related field, or equivalent experience • At least 5 years of DevOps or Software Development experience or a similar role • Nice-to-have: network and/or storage architecture experience • Nice-to-have: high-performance computing or GPU infrastructure experience, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise • Nice-to-have: open source community presence, upstream contributions, or conference presentations • Nice-to-have: experience with Rancher, OpenShift, or VMware

🏖️ Benefits

• Professional development and training • Attend conferences and working groups • Company outings, happy hours, hackathons, and tech talks • Competitive compensation package with a strong benefits plan • Remote work option

Apply Now

Similar Jobs

🔥 12 hours ago

ArangoDB

51 - 200

🏥 Healthcare

📦 Logistics

💼 Consulting

Site Reliability Engineer maintaining Kubernetes and cloud infrastructure for Arango’s contextual AI data platform. Automating operations, observability, CI/CD, and reliability for enterprise AI systems.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

Terraform

Go

🕒 3 days ago

Devoteam

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data AWS DevSecOps professional evolving secure cloud data platforms for Devoteam’s large-organization clients. Advising teams on architecture, security, governance and observability in a remote Spain-based role.

Airflow

AWS

Cloud

DynamoDB

ETL

Hadoop

Python

Spark

SQL

Terraform

🕒 3 days ago

Logicalis Spain

1001 - 5000

💼 Consulting

DevOps Engineer operando y automatizando plataformas Kubernetes para Logicalis Spain, proveedor de servicios IT empresariales. Mejorando CI/CD, infraestructura cloud y servicios gestionados.

🗣️🇪🇸 Spanish Required

Ansible

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Grafana

Jenkins

Kubernetes

OpenShift

Prometheus

Python

Terraform

🕒 3 days ago

Tempo Software

201 - 500

☁️ SaaS

🏢 Enterprise

⚡ Productivity

Senior Site Reliability Engineer building AWS infrastructure, CI/CD pipelines, and Kubernetes platforms for Tempo’s enterprise productivity software. Automating reliability, observability, security, and cloud deployments.

Ansible

AWS

Cloud

Docker

Java

Kotlin

Kubernetes

Linux

Terraform

🕒 August 7

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior SRE strengthening reliability for Affirm’s honest, flexible buy-now-pay-later platform. Building incident lifecycle, observability and resilient backend practices for global engineering teams.

AWS

Distributed Systems

Kotlin

Kubernetes

MySQL

Python