Senior Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇪🇸 Spain – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Lodgify

Lodgify

51 - 200 employees

💼 Consulting

📣 Marketing

📦 Logistics

💰 $30M Series B on 2022-11

Consulting • Marketing • Logistics

Lodgify is a comprehensive vacation rental property management software company that helps property owners and managers streamline their operations. It offers features such as a reservation system, unified inbox, guest management, task management, automation, reporting and analytics, and owner statements. Lodgify supports integrations with major booking platforms like Airbnb, Booking. com, Vrbo, and Expedia, providing a channel manager to sync calendars, rates, and reservations. It also enables users to create custom websites with a booking widget, supporting platforms like WordPress, Squarespace, Wix, and more. With additional tools like dynamic pricing, a mobile app, and AI-powered automation, Lodgify aims to increase productivity and enhance guest experiences for short-term rental businesses. The company provides free onboarding, personalized support, and a variety of resources to ensure a smooth transition to their platform, appealing to individual hosts, property managers, and various types of accommodations.

📋 Description

• Define meaningful SLIs, SLOs, and reliability targets for the platform. • Collaborate with software engineering teams on observability, SLIs, SLOs, and reliability best practices. • Improve production readiness through service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness. • Improve reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure. • Build actionable observability using metrics, logs, traces, and golden signals with Datadog, Prometheus, and Grafana. • Implement operational and security best practices through guidelines, policies, and automation. • Reduce alert noise and improve signal quality. • Automate repetitive operational work using Python or other languages. • Implement self-service Internal Developer Platform features via APIs and Kubernetes operators. • Improve deployment safety, rollbackability, and release observability. • Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms. • Participate in on-call, troubleshoot, coordinate incident response, and facilitate blameless post-incident reviews. • Execute disaster recovery drills and analyze cloud/platform usage for cost and resource-efficiency gains. • Strengthen observability, reduce operational toil, improve incident response, define SRE standards, and help teams own their systems in production.

🎯 Requirements

• 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure. • Understanding and application of SRE practices, including SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery. • Ability to design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals. • Experience troubleshooting complex distributed systems. • Ability to write maintainable software to automate operational tasks and reduce manual intervention. • Experience with stateful production systems such as relational databases, caches, queues, or streaming platforms. • Ability to balance reliability, performance, cost, and delivery speed pragmatically. • Comfort working in a transitional environment where SRE practices are being introduced. • Effective collaboration with Engineering, Platform, Security, and Product stakeholders. • Clear communication, strong documentation, and interest in coaching teams toward stronger production ownership. • Initiative and accountability, including raising risks early and driving improvements through completion. • All applications and CVs must be submitted in English.

🏖️ Benefits

• Remote Flexibility: The freedom to work from home any day that works for you. • 25 working days of paid vacation • Jornada Intensiva in August • Premium health, dental, and mental health support via Alan; pre-existing conditions are covered. • €150/month meal allowance on your Alan card • 50% off Ametller Origen prepared dishes at the office • Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo) • Table, ergonomic chair, and monitor for home setup • Free Spanish classes • Cash rewards for employee referrals • Daily office breakfast • Monthly team events • Inclusive, international work environment

Apply Now

Similar Jobs

🕒 5 days ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

DevSecOps Engineer securing Akamai’s cloud-edge platform through CI/CD automation, container orchestration, and infrastructure as code. Building scalable storage solutions for enterprises.

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

Go

🕒 August 18

ArangoDB

51 - 200

🏥 Healthcare

📦 Logistics

💼 Consulting

Site Reliability Engineer maintaining Kubernetes and cloud infrastructure for Arango’s contextual AI data platform. Automating operations, observability, deployments, and reliability across distributed database systems.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

Go

🕒 August 18

HostPapa

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Site Reliability Engineer improving CloudBlue’s multi-tenant SaaS reliability, scalability, and observability. Operating Kubernetes platforms, leading incident response, and strengthening cloud infrastructure for service providers worldwide.

AWS

Azure

Cloud

Distributed Systems

Docker

ElasticSearch

Google Cloud Platform

Grafana

Kubernetes

Linux

Python

🕒 August 17

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE deploying and operating Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis, a cloud-native infrastructure company. Improving reliability, security, scalability, and automation.

Cloud

Distributed Systems

JavaScript

Kubernetes

Linux

Microservices

Open Source

OpenShift

OpenStack

Python

VMware

Go

🕒 August 14

Devoteam

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data AWS DevSecOps professional evolving secure cloud data platforms for Devoteam’s large-organization clients. Advising teams on architecture, security, governance and observability in a remote Spain-based role.

Airflow

AWS

Cloud

DynamoDB

ETL

Hadoop

Python

Spark

SQL

Terraform