Senior Site Reliability Engineer

🕒 August 27

🇪🇸 Spain – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Lodgify

Lodgify

51 - 200 employees

💼 Consulting

📣 Marketing

📦 Logistics

💰 $30M Series B on 2022-11

Consulting • Marketing • Logistics

Lodgify is a comprehensive vacation rental property management software company that helps property owners and managers streamline their operations. It offers features such as a reservation system, unified inbox, guest management, task management, automation, reporting and analytics, and owner statements. Lodgify supports integrations with major booking platforms like Airbnb, Booking. com, Vrbo, and Expedia, providing a channel manager to sync calendars, rates, and reservations. It also enables users to create custom websites with a booking widget, supporting platforms like WordPress, Squarespace, Wix, and more. With additional tools like dynamic pricing, a mobile app, and AI-powered automation, Lodgify aims to increase productivity and enhance guest experiences for short-term rental businesses. The company provides free onboarding, personalized support, and a variety of resources to ensure a smooth transition to their platform, appealing to individual hosts, property managers, and various types of accommodations.

📋 Description

• Define meaningful SLIs, SLOs, and reliability targets for the platform. • Collaborate with software engineering teams on observability, SLIs, SLOs, and reliability best practices. • Improve production readiness through service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness. • Improve reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure. • Build actionable observability using metrics, logs, traces, and golden signals with Datadog, Prometheus, and Grafana. • Implement operational and security best practices through guidelines, policies, and automation. • Reduce alert noise and improve signal quality. • Automate repetitive operational work using Python or other languages. • Implement self-service Internal Developer Platform features via APIs and Kubernetes operators. • Improve deployment safety, rollbackability, and release observability. • Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms. • Participate in on-call, troubleshoot, coordinate incident response, and facilitate blameless post-incident reviews. • Execute disaster recovery drills and analyze cloud/platform usage for cost and resource-efficiency gains. • Strengthen observability, reduce operational toil, improve incident response, define SRE standards, and help teams own their systems in production.

🎯 Requirements

• 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure. • Understanding and application of SRE practices, including SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery. • Ability to design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals. • Experience troubleshooting complex distributed systems. • Ability to write maintainable software to automate operational tasks and reduce manual intervention. • Experience with stateful production systems such as relational databases, caches, queues, or streaming platforms. • Ability to balance reliability, performance, cost, and delivery speed pragmatically. • Comfort working in a transitional environment where SRE practices are being introduced. • Effective collaboration with Engineering, Platform, Security, and Product stakeholders. • Clear communication, strong documentation, and interest in coaching teams toward stronger production ownership. • Initiative and accountability, including raising risks early and driving improvements through completion. • All applications and CVs must be submitted in English.

🏖️ Benefits

• Remote Flexibility: The freedom to work from home any day that works for you. • 25 working days of paid vacation • Jornada Intensiva in August • Premium health, dental, and mental health support via Alan; pre-existing conditions are covered. • €150/month meal allowance on your Alan card • 50% off Ametller Origen prepared dishes at the office • Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo) • Table, ergonomic chair, and monitor for home setup • Free Spanish classes • Cash rewards for employee referrals • Daily office breakfast • Monthly team events • Inclusive, international work environment

Apply Now

Similar Jobs

🕒 August 21

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

DevSecOps Engineer securing Akamai’s cloud-edge platform through CI/CD automation, container orchestration, and infrastructure as code. Building scalable storage solutions for enterprises.

🕒 August 17

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE deploying and operating Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis, a cloud-native infrastructure company. Improving reliability, security, scalability, and automation.

🕒 August 14

Logicalis Spain

1001 - 5000

💼 Consulting

DevOps Engineer operando y automatizando plataformas Kubernetes para Logicalis Spain, proveedor de servicios IT empresariales. Mejorando CI/CD, infraestructura cloud y servicios gestionados.

🗣️🇪🇸 Spanish Required

🕒 August 14

Tempo Software

201 - 500

☁️ SaaS

🏢 Enterprise

⚡ Productivity

Senior Site Reliability Engineer building AWS infrastructure, CI/CD pipelines, and Kubernetes platforms for Tempo’s enterprise productivity software. Automating reliability, observability, security, and cloud deployments.

🕒 July 31

Exoscale

51 - 200

☁️ SaaS

🤖 Artificial Intelligence

Site Reliability Engineer at Exoscale focusing on maintaining and designing operating systems and hypervisor internals. Join a dynamic multicultural team to enhance product services across Europe.