Senior Site Reliability Engineer – AWS, Kubernetes, Databricks

Job not on LinkedIn

🔥 1 minute ago

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Compass

Compass

10,000+ employees

🏠 Real Estate

📱 Media

Real Estate • Media

Compass is a real-estate-focused content and services site that provides detailed market analysis, buying/selling/renting guides, mortgage and financing information, and home improvement and renovation advice. The site offers resources for homebuyers, sellers, renters, agents, and real estate investors — including articles on appraisals, affordable housing, investment strategies, staging and property maintenance. Compass aims to help users make informed decisions across the housing lifecycle through timely market updates and practical how-to content.

📋 Description

• Ensure the reliability, availability, and performance of systems and applications in production; • Design, implement, and evolve automations for deployment, monitoring, scalability, and platform operations; • Administer and evolve Kubernetes environments, ensuring stability and high availability; • Respond to critical incidents, conducting root cause analyses (RCA) and implementing preventive actions; • Define, track, and refine reliability indicators such as SLIs, SLOs, and SLAs; • Manage and optimize CI/CD pipelines to enable continuous and secure delivery; • Implement and maintain infrastructure as code (IaC), ensuring standardization and automation of environments; • Develop and strengthen observability, monitoring, metrics, and alerting practices; • Collaborate with development teams to build resilient, scalable, and high-performance applications; • Document operational procedures, contingency plans, and operational best practices; • Drive continuous improvement of the platform to reduce recurring failures and increase service reliability.

🎯 Requirements

• Strong experience as a Site Reliability Engineer (SRE), DevOps Engineer, or Platform Engineer; • Experience managing Kubernetes clusters in production environments; • Experience with cloud infrastructure, preferably AWS; • Experience with infrastructure automation using Terraform and Infrastructure as Code (IaC); • Experience with CI/CD pipelines and continuous integration/delivery processes; • Experience with observability, monitoring, and incident management tools; • Experience troubleshooting distributed applications and critical production environments; • Experience with Databricks and Apache Spark; • Experience with AWS CodePipeline; • Knowledge of SQL databases, data warehousing, and data modeling; • Experience with Amazon EC2, Amazon S3, and AWS Lambda; • Knowledge of cloud architecture and Microsoft Azure environments; • Knowledge of high availability, scalability, resilience, and incident recovery best practices; • Analytical ability to identify root causes and implement continuous improvements; • Advantage: Databricks certifications.

Apply Now

Similar Jobs

🔥 12 hours ago

Cloud Architect & DevOps Senior role at POWER DATA, focusing on critical cloud architecture and data sovereignty protocols.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Cloud

Google Cloud Platform

Kafka

Kubernetes

RabbitMQ

🕒 Yesterday

Smarthis

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Lead Cloud Platform & DevOps Engineer helping drive digital transformation across Latin America. Experience with Azure, DevOps practices, and strong leadership in large-scale cloud projects.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

Cloud

Terraform

🕒 3 days ago

Inmetrics

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

SRE/Cloud Infrastructure Specialist managing cloud environments and optimizing costs at Inmetrics. Proposing improvements in CI/CD pipelines and cloud resource management for clients' internal teams.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Cloud

Google Cloud Platform

Grafana

Java

Kubernetes

Prometheus

Python

Terraform

🕒 July 22

Seox | Inteligência digital para Publishers

11 - 50

💼 Consulting

📣 Marketing

📱 Media

Profissional de Infraestrutura e DevOps working on optimizing client news portals with PHP, Nginx, and Cloudflare

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

DNS

Linux

NGINX

Oracle

PHP

🕒 July 20

Verity Group

51 - 200

💼 Consulting

🤖 Artificial Intelligence

🔒 Cybersecurity

SRE/DevOps Engineer managing cloud-native platforms and CI/CD pipelines for Verity's digital transformation projects. Focused on reliability, performance, and automation in high-availability environments.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

ElasticSearch

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Terraform