Senior SRE Engineer – Observability & Reliability

Job not on LinkedIn

🔥 0 minutes ago

🇧🇷 Brazil – Remote

⏳ Contract/Temporary

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 13%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GrooveTech

GrooveTech

51 - 200 employees

Founded 2017

🤝 B2B

🏢 Enterprise

🎯 Recruiter

B2B • Enterprise • Recruitment

GrooveTech is a Brazilian IT services firm that provides managed staff augmentation and squads, software quality assurance, 24x7 NOC monitoring, and strategic IT consulting including digital due diligence for M&A. They supply managed teams (Executives of IT, PM/PO, technical leads and business partners), daily reporting, and focus on ROI by reducing turnover and improving project performance through testing automation, monitoring (on-prem and cloud), and tailored delivery models.

📋 Description

• Lead the administration and evolution of Datadog as the central observability platform, including APM, traces, log pipelines, metrics, RUM, Synthetics, SLOs, monitors, and Cloud SIEM • Ensure comprehensive observability coverage for production environments on AWS, including ECS, Lambda, API Gateway, DynamoDB, EventBridge/SQS, CloudFront, and WAF • Define and track SLIs and SLOs for critical journeys, monitoring error budget consumption • Design signal-focused alerts by eliminating false positives, duplicate notifications, and inefficient monitors • Integrate and audit security logs, including CloudTrail, GuardDuty, Inspector, Macie, and VPC Flow Logs • Lead critical incident response, including triage, event correlation, transparent communication, and leadership of RCAs and postmortems • Maintain technical governance by creating runbooks and playbooks and supporting Change Management processes in Jira Service Management • Support PCI DSS compliance by producing monitoring, logging, and FIM evidence, as well as supporting audits and penetration testing • Manage Observability FinOps by optimizing Datadog costs related to retention, log indexing, and metric cardinality • Serve as the technical reference and owner for platform observability and reliability

🎯 Requirements

• Strong, hands-on track record (8+ years) in SRE, Observability, or the support of critical environments • Deep, advanced expertise in the Datadog platform, including APM, Traces, Logs, Synthetics, SLOs, and Dashboards • Strong knowledge of AWS architecture focused on operational troubleshooting, including CloudWatch, CloudTrail, ECS, Lambda, queues, and events • Experience managing incidents, problems, and changes (ITIL or equivalent) • Proven experience creating structured technical documentation, including runbooks, playbooks, and postmortems • Ability to execute, troubleshoot, and deliver quickly, without remaining limited to the theoretical or architectural level • Ability to independently manage the observability discipline and prioritize backlogs without the need for micromanagement • Excellent adaptability to dynamic routines, openness to feedback, and a focus on continuous improvement • Excellent interpersonal skills for aligned and transparent collaboration with engineering, QA, and product teams • Advanced to fluent English • Experience in payments, fintech, or PCI DSS-regulated environments • Knowledge of Cloud SIEM and AWS security tools, including GuardDuty, Inspector, and Macie • Observability as Code practices using Terraform and the Datadog API • Ability to read and analyze Go and/or TypeScript/Node.js code to support diagnostics

🏖️ Benefits

• Life insurance • Access to the Wellhub ecosystem (formerly Gympass) • Support from technical and management back-office teams to ensure accelerated onboarding and continued support throughout your journey

Apply Now

Similar Jobs

🔥 1 minute ago

Finaya

11 - 50

💳 Fintech

☁️ SaaS

🔌 API

Engenheira SRE/DevOps garantindo confiabilidade de serviços financeiros críticos da Finaya. Gerenciando AWS, Kubernetes, observabilidade, incidentes, Terraform e pipelines CI/CD.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Cloud

ElasticSearch

Kubernetes

Python

Terraform

Go

🕒 6 days ago

DOMVS iT

51 - 200

💼 Consulting

🏥 Healthcare

🏢 Enterprise

DevOps Sênior construindo pipelines CI/CD e infraestrutura AWS/Terraform para assistente financeiro de IA da DOMVS iT. Trabalho 100% remoto, com integração via WhatsApp.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Cloud

Terraform

🕒 September 17

Metal Toad

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

DevOps Engineer designing and maintaining scalable AWS infrastructure for Metal Toad, a strategic AI partner. Supporting high-availability applications, managed services, automation, security, and disaster recovery.

AWS

Cloud

Linux

Python

TCP/IP

🕒 September 2

Sabiá Administração

51 - 200

🎲 Gambling

📣 Marketing

SRE garantindo disponibilidade, resiliência e resposta a incidentes na plataforma iGaming da Cometa Gaming. Automatizando infraestrutura, CI/CD, Kubernetes e observabilidade.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

Flux

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Terraform

🕒 August 19

Inflect

11 - 50

☁️ SaaS

Senior DevOps consultant architecting reliable AWS EKS platforms for Inflect’s digital infrastructure marketplace. Delivering Terraform, SRE, OpenTelemetry, and CI/CD solutions.

AWS

Cloud

NFS

Terraform