Senior Site Reliability Engineer, SRE

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

• Design, develop, deploy, operate, maintain, and troubleshoot cloud-based AI infrastructure solutions based on open source software • Deploy AI infrastructure built on NVIDIA-certified hardware according to engineering architecture and implementation designs • Ensure the reliability, security, performance, and scalability of container infrastructure • Work with geographically distributed international teams on technical challenges and process improvements • Collaborate with stakeholders to gather and refine technical requirements • Troubleshoot, debug, and resolve complex technical issues across networking, storage, Linux, and Kubernetes • Participate in code reviews • Design and implement AI-driven automation across the DevOps lifecycle • Facilitate knowledge transfer to customers during delivery phases • Mentor team members and Mirantis customers • Define technical strategies and make independent technical decisions when working with customers • Stay current with cloud operations and development trends and best practices

🎯 Requirements

• 5+ years of professional experience in DevOps, focused on cloud and infrastructure technologies, including Kubernetes and/or OpenStack • Experience with high-performance data center processing, networking, and storage • Exposure to Golang and working knowledge of Python, JavaScript, and other programming languages • Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines • Exceptional problem-solving and debugging skills across networking and storage, Linux, and Kubernetes • Knowledge of performance optimization and security • Ability to lead technical tasks and collaborate with diverse teams • Ability to make independent judgment calls while working directly with customers with limited day-to-day oversight • Excellent written and spoken English • Excellent customer-facing communication skills • Commitment to innovation, continuous learning, and high-quality results • Ability to travel up to 25%, including internationally if needed • Bachelor's degree in Computer Science or a related field, or equivalent experience • At least 5 years of DevOps or Software Development experience or experience in a similar role • Extensive network and/or storage architecture experience is nice to have • Experience with high-performance computing or GPU infrastructure is nice to have • Open source community presence is nice to have • Experience with Rancher, OpenShift, and VMware is nice to have

🏖️ Benefits

• Professional development and training • Attend conferences and working groups • Company outings, happy hours, hackathons, and tech talks • Competitive compensation package with a strong benefits plan • Work with passionate, talented and engaging colleagues • Be part of cutting-edge, open-source innovation • Openness, collaboration, risk-taking, and continuous growth valued

Apply Now

Similar Jobs

🕒 July 28

OpsMill

11 - 50

☁️ SaaS

🤖 Artificial Intelligence

🤝 B2B

Product Reliability Engineer enhancing on-prem deployment reliability for OpsMill's Infrahub while partnering with customers on troubleshooting and improvements. Building diagnostics tools and resolving complex issues in Kubernetes environments.

Distributed Systems

Kubernetes

Python

Rust

Go

🕒 July 27

Talentgrator

11 - 50

🎯 Recruiter

👥 HR Tech

🎲 Gambling

Experienced DevOps Engineer responsible for server administration, automation, and system management at an iGaming company. Collaborating with development teams and implementing CI/CD pipelines for improved workflows.

Ansible

DNS

Docker

ElasticSearch

Flux

Grafana

Kafka

Kubernetes

MongoDB

NGINX

Postgres

Prometheus

Python

RabbitMQ

Redis

TCP/IP

Terraform

Go