SRE Engineer

🔥 8 minutes ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Verity Group

Verity Group

51 - 200 employees

Founded 2010

💼 Consulting

🤖 Artificial Intelligence

🔒 Cybersecurity

Consulting • Artificial Intelligence • Cybersecurity

Verity Group is a digital transformation and innovation consulting firm that focuses on delivering real results through technology. With expertise in areas such as AI migration, hyperautomation, cybersecurity, and digital engineering, Verity Group partners with ambitious companies to develop strategic business solutions. The firm prides itself on prioritizing depth over volume, ensuring tailored services that drive efficiency and growth throughout the technology journey, from strategy to execution.

📋 Description

• Define and track SLIs, SLOs, SLAs, MTTR, and MTTD • Implement observability, monitoring, alerting, and APM • Monitor latency, traffic, errors, saturation, availability, and performance • Prevent, identify, and resolve incidents • Lead root cause analyses and define actions to prevent recurrence • Identify risks, bottlenecks, and single points of failure • Support the design of resilient, scalable, and highly available solutions • Automate operational activities and reduce manual tasks • Operate and evolve Kubernetes and Docker environments • Support capacity planning, business continuity, and disaster recovery strategies • Participate in deployments and support application stabilization • Work with teams to improve reliability from the solution design stage • Create and maintain dashboards, alerts, procedures, and operational documentation • Promote a culture of reliability, observability, and continuous improvement

🎯 Requirements

• Experience as a Site Reliability Engineer, SRE, or in an equivalent role • Hands-on experience with cloud environments using GCP, AWS, and/or Azure • Knowledge of Kubernetes and Docker • Experience with observability, monitoring, alerting, and APM • Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD • Experience managing, investigating, and resolving incidents • Knowledge of application and infrastructure troubleshooting • Experience administering Linux environments • Knowledge of networking, security, performance, and high availability • Experience with automation and Infrastructure as Code • Experience with CI/CD pipelines • Strong communication skills and the ability to work with cross-functional teams • Analytical, proactive, collaborative, and prevention-oriented mindset • Experience with GKE, EKS, or AKS • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools • Experience with the ELK Stack, Elasticsearch, and Kibana • Knowledge of Terraform and Ansible • Experience with mission-critical environments and distributed systems • Experience in financial institutions or regulated environments • Experience with capacity management and cloud cost optimization • Knowledge of disaster recovery and business continuity • Experience defining and managing error budgets • Cloud, Kubernetes, or SRE certifications • Availability for employment under the CLT regime, as indicated in the application form

🏖️ Benefits

• Meal allowance • Food allowance • Home office allowance • Health insurance • Dental insurance • Life insurance • Birthday day off • TotalPass / Wellhub • Boon Saúde app • Discount partnerships • Partnerships with businesses and educational institutions • Welcome kit • Onboarding • Verity Learning • Verity Break • #VerityComVocê • Great Place to Work certification

Apply Now

Similar Jobs

🔥 43 minutes ago

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Site Reliability Engineer operating monitoring platforms and improving cloud reliability for CI&T’s AI transformation solutions. Automating observability, incident response, and cybersecurity remediation.

AWS

Azure

Cloud

Cyber Security

DNS

Google Cloud Platform

Grafana

JavaScript

Linux

Python

Splunk

Terraform

TypeScript

Go

🔥 6 hours ago

Franq

51 - 200

🛡️ Insurance

💼 Consulting

💳 Fintech

DevOps/SRE administrando Kubernetes, AWS/GCP e automação de infraestrutura na Franq. Fortalecendo observabilidade, CI/CD, disponibilidade e confiabilidade de sistemas financeiros.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Apache

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Java

Kubernetes

Linux

NoSQL

Prometheus

Python

SQL

Terraform

🕒 4 days ago

SoftDesign

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

DevOps Engineer building Azure platforms, Kubernetes infrastructure, and CI/CD automation at SoftDesign. Advancing observability, DevSecOps, and developer experience.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

Cloud

Grafana

Kubernetes

Prometheus

Terraform

Vault

🕒 4 days ago

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

Especialista DevOps construindo e arquitetando plataformas AKS multiambiente na GFT Technologies. Automatizando infraestrutura Azure, alta disponibilidade e migração de aplicações para contêiner.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

DNS

Docker

Flux

Kubernetes

Node.js

Terraform

Vault

🕒 4 days ago

Aiphoria

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior DevOps Engineer operating Kubernetes, cloud, and GPU inference infrastructure for an AI product company. Building secure, observable, and scalable remote production platforms.

Ansible

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Kafka

Kubernetes

Linux

Microservices

Postgres

Prometheus

Python

Terraform