Site Reliability Engineer

🕒 August 18

🇪🇸 Spain – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of HostPapa

HostPapa

51 - 200 employees

Founded 2006

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

HostPapa is a web hosting company that provides a variety of hosting solutions including shared web hosting, WordPress hosting, VPS hosting, and reseller hosting. They also offer additional services such as domain registration, website building tools, business email services, and security features like SSL certificates. HostPapa prides itself on its 24/7 award-winning customer support available globally, ensuring that customers receive assistance in their preferred language and timezone. With a focus on small businesses, HostPapa aims to empower customers to achieve their online goals with reliable and high-performance hosting services.

📋 Description

• Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services • Influence system architecture with a focus on reliability, scalability, and operability • Reduce operational toil through automation and process improvement • Design and operate the observability stack across metrics, logs, and traces • Develop alerting strategies and dashboards for platform and business health • Design and maintain high-availability architectures, redundancy, failover, and disaster recovery strategies • Conduct capacity planning, load testing, and performance optimization • Lead production incident coordination, communication, and service restoration • Own blameless postmortems and drive improvements to reduce incidents, MTTR, and customer impact • Improve reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing • Partner with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability • Maintain runbooks and operational documentation • Promote SRE best practices across engineering teams • Support other tasks or projects as assigned

🎯 Requirements

• 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer • Strong ownership of production systems • Experience operating highly available, enterprise-grade, multi-tenant SaaS platforms • Hands-on experience with Datadog, Grafana, and Elasticsearch/Kibana • Solid understanding of Linux, networking, and distributed systems fundamentals • Experience with Docker and Kubernetes • Strong scripting and automation skills using Python and/or Bash • Experience participating in on-call rotations and production incident response • Strong written and spoken English • Cloud experience, preferably with Azure; AWS and/or GCP experience valued • Experience with hybrid or on-premises integrations beneficial • Experience defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing are advantageous or considered assets

🏖️ Benefits

• A competitive salary that values you and your unique skill sets • Career advancement & professional development opportunities • Flexible work arrangements to support work/life balance • 24/7 award-winning customer support • Diversity and inclusion • Accommodation may be provided in all parts of the hiring process

Apply Now

Similar Jobs

🕒 August 17

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE deploying and operating Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis, a cloud-native infrastructure company. Improving reliability, security, scalability, and automation.

Cloud

Distributed Systems

JavaScript

Kubernetes

Linux

Microservices

Open Source

OpenShift

OpenStack

Python

VMware

Go

🕒 August 14

Logicalis Spain

1001 - 5000

💼 Consulting

DevOps Engineer operando y automatizando plataformas Kubernetes para Logicalis Spain, proveedor de servicios IT empresariales. Mejorando CI/CD, infraestructura cloud y servicios gestionados.

🗣️🇪🇸 Spanish Required

Ansible

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Grafana

Jenkins

Kubernetes

OpenShift

Prometheus

Python

Terraform

🕒 August 14

Tempo Software

201 - 500

☁️ SaaS

🏢 Enterprise

⚡ Productivity

Senior Site Reliability Engineer building AWS infrastructure, CI/CD pipelines, and Kubernetes platforms for Tempo’s enterprise productivity software. Automating reliability, observability, security, and cloud deployments.

Ansible

AWS

Cloud

Docker

Java

Kotlin

Kubernetes

Linux

Terraform

🕒 July 31

Exoscale

51 - 200

☁️ SaaS

🤖 Artificial Intelligence

Site Reliability Engineer at Exoscale focusing on maintaining and designing operating systems and hypervisor internals. Join a dynamic multicultural team to enhance product services across Europe.

Linux

Python

Go

🕒 July 30

Merlin Digital Partner

11 - 50

☁️ SaaS

⚡ Productivity

🤝 B2B

Site Reliability Engineer ensuring maximum system availability and performance for a SaaS solutions provider in Spain. Collaborating to automate and monitor systems while ensuring best practices in reliability and security.

🗣️🇪🇸 Spanish Required

AWS

Azure

Cloud

Google Cloud Platform

Python

Terraform