SRE, Site Reliability Engineer

Job not on LinkedIn

🔥 12 hours ago

🇧🇷 Brazil – Remote

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 15%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sabiá Administração

Sabiá Administração

51 - 200 employees

Founded 2024

🎲 Gambling

📣 Marketing

Gambling • Marketing

Sabiá Administração is a strategic management hub for brands in the betting sector. We work with intelligence, ethics and performance to transform ideas into sustainable businesses in iGaming. Our mission is to professionalize our operations and position them among the leading ones in the Brazilian market. We are proudly responsible for the brands: Lotogreen, BR4BET and Gol de Bet. Here, you can follow the behind-the-scenes of our ecosystem, our achievements and the evolution of those who believe that the union of talents transforms the game.

📋 Description

• Actively participate in production incidents, including root cause investigation, mitigation, and blameless postmortems • Define and track SLIs, SLOs, and Error Budgets • Reduce toil through automation and perform advanced troubleshooting on Linux and distributed systems • Create and evolve CI/CD pipelines and GitOps strategies using tools such as Argo CD and Flux • Provision and maintain infrastructure with Terraform, Terragrunt, and Ansible across AWS, GCP, or Azure • Manage Docker- and Kubernetes-based environments • Implement and enhance observability with Datadog, Grafana, and Prometheus • Support development teams with a focus on DevEx, resilience, performance, and cloud-native security • Turn incidents into continuous improvements for the reliability and availability of Cometa Gaming’s platform

🎯 Requirements

• Previous hands-on experience as an SRE, DevOps Engineer, or in an equivalent role, with real-world experience responding to production incidents (on-call) • Strong knowledge of Kubernetes, containers, and cloud infrastructure (AWS, GCP, or Azure) • Proficiency in CI/CD, IaC, environment automation, and Site Reliability Engineering concepts (SLOs, Error Budgets, and toil reduction) • Familiarity with monitoring/telemetry and security tools in cloud-native environments • Proactive approach to investigating complex problems and a passion for engineering best practices • Bachelor’s degree • English proficiency

Apply Now

Similar Jobs

🕒 August 19

Inflect

11 - 50

☁️ SaaS

Senior DevOps consultant architecting reliable AWS EKS platforms for Inflect’s digital infrastructure marketplace. Delivering Terraform, SRE, OpenTelemetry, and CI/CD solutions.

AWS

Cloud

NFS

Terraform

🕒 August 17

Metal Toad

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

DevOps Engineer designing and maintaining AWS infrastructure for Metal Toad, a cloud consulting firm serving enterprise customers. Building scalable, secure systems and supporting mission-critical applications through automation, monitoring, and incident response.

AWS

Cloud

Linux

Python

TCP/IP

🕒 July 28

In All Media

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

DevOps Manager leading Azure cloud and Kubernetes environments for a power management platform. Managing CI/CD, incident lifecycle, and a technical team within a remote LATAM context.

Azure

Cloud

Grafana

Kubernetes

🕒 July 21

Uberall

201 - 500

🍽️ Food & Beverage

🏨 Hospitality

🚘 Automotive

DevSecOps Engineer building, maintaining, and scaling secure cloud infrastructure and CI/CD pipelines with a security-first mindset. Collaborating with a small team to ensure platform security and compliance.

AWS

Cloud

EC2

Kubernetes

Microservices

Shell Scripting

Terraform

🕒 July 16

CIAL Dun & Bradstreet

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

SRE role managing critical incidents and observability for SaaS solutions in Latin America. Responsibilities include monitoring, incident response, and improving platform observability.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

Grafana

Prometheus

Python