Site Reliability Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of HostPapa

HostPapa

51 - 200 employees

Founded 2006

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

HostPapa is a web hosting company that provides a variety of hosting solutions including shared web hosting, WordPress hosting, VPS hosting, and reseller hosting. They also offer additional services such as domain registration, website building tools, business email services, and security features like SSL certificates. HostPapa prides itself on its 24/7 award-winning customer support available globally, ensuring that customers receive assistance in their preferred language and timezone. With a focus on small businesses, HostPapa aims to empower customers to achieve their online goals with reliable and high-performance hosting services.

📋 Description

• Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services • Influence system architecture with a focus on reliability, scalability, and operability • Reduce operational toil through automation and process improvement • Design and operate the observability stack across metrics, logs, and traces • Develop alerting strategies and dashboards for platform and business health • Design and maintain high-availability architectures, redundancy, failover, and disaster recovery strategies • Conduct capacity planning, load testing, and performance optimization • Lead production incident coordination, communication, and service restoration • Own blameless postmortems and drive improvements to reduce incidents, MTTR, and customer impact • Improve reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing • Partner with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability • Maintain runbooks and operational documentation • Promote SRE best practices across engineering teams • Support other tasks or projects as assigned

🎯 Requirements

• 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer • Strong ownership of production systems • Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms • Hands-on experience with Datadog, Grafana, and Elasticsearch/Kibana • Solid understanding of Linux, networking, and distributed systems fundamentals • Experience with Docker and Kubernetes • Strong scripting and automation skills using Python and/or Bash • Experience participating in on-call rotations and production incident response • Strong written and spoken English • Cloud experience, preferably with Azure; AWS and/or GCP experience valued • Experience with hybrid or on-premises integrations beneficial • Experience defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing are advantageous or considered assets

🏖️ Benefits

• A competitive salary that values you and your unique skill sets • Career advancement & professional development opportunities • Flexible work arrangements to support work/life balance • 24/7 award-winning customer support • Diversity and inclusion • Accommodation may be provided in all parts of the hiring process

Apply Now

Similar Jobs

🔥 11 hours ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE deploying and operating Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis, a cloud-native infrastructure company. Improving reliability, security, scalability, and automation.

Cloud

Distributed Systems

JavaScript

Kubernetes

Linux

Microservices

Open Source

OpenShift

OpenStack

Python

VMware

Go

🔥 23 hours ago

ArangoDB

51 - 200

🏥 Healthcare

📦 Logistics

💼 Consulting

Site Reliability Engineer maintaining Kubernetes and cloud infrastructure for Arango’s contextual AI data platform. Automating operations, observability, CI/CD, and reliability for enterprise AI systems.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

Terraform

Go

🕒 3 days ago

Devoteam

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data AWS DevSecOps professional evolving secure cloud data platforms for Devoteam’s large-organization clients. Advising teams on architecture, security, governance and observability in a remote Spain-based role.

Airflow

AWS

Cloud

DynamoDB

ETL

Hadoop

Python

Spark

SQL

Terraform

🕒 3 days ago

Logicalis Spain

1001 - 5000

💼 Consulting

DevOps Engineer operando y automatizando plataformas Kubernetes para Logicalis Spain, proveedor de servicios IT empresariales. Mejorando CI/CD, infraestructura cloud y servicios gestionados.

🗣️🇪🇸 Spanish Required

Ansible

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Grafana

Jenkins

Kubernetes

OpenShift

Prometheus

Python

Terraform

🕒 3 days ago

Tempo Software

201 - 500

☁️ SaaS

🏢 Enterprise

⚡ Productivity

Senior Site Reliability Engineer building AWS infrastructure, CI/CD pipelines, and Kubernetes platforms for Tempo’s enterprise productivity software. Automating reliability, observability, security, and cloud deployments.

Ansible

AWS

Cloud

Docker

Java

Kotlin

Kubernetes

Linux

Terraform