SRE Engineer

🔥 16 minutes ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Verity Group

Verity Group

51 - 200 employees

Founded 2010

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Consulting • Healthcare • Insurance

<Verity Group> is a Brazil-based digital transformation and innovation consultancy that develops and accelerates products and services for enterprises using modern engineering, cloud, and AI-driven approaches. The company offers application modernization, digital experience design, technology outsourcing, and consulting services, and operates an AI orchestration platform called Verity Quantum to integrate and deploy AI agents. Verity Group serves B2B clients in banking, insurance, healthcare and other industries, has 200+ consultants and 1,000+ projects over 15+ years, and focuses on end-to-end strategy-to-execution delivery to generate measurable business value.

📋 Description

• Define and track SLIs, SLOs, SLAs, MTTR, and MTTD. • Implement observability, monitoring, alerting, and APM. • Monitor latency, traffic, errors, saturation, availability, and performance. • Prevent, identify, and resolve incidents. • Lead root cause analyses and define actions to prevent recurrence. • Identify risks, bottlenecks, and single points of failure. • Support the design of resilient, scalable, and highly available solutions. • Automate operational activities and reduce manual tasks. • Operate and evolve Kubernetes and Docker environments. • Support capacity planning, business continuity, and disaster recovery strategies. • Participate in deployments and support application stabilization. • Collaborate with teams to improve reliability from the solution design stage onward. • Create and maintain dashboards, alerts, procedures, and operational documentation. • Promote a culture of reliability, observability, and continuous improvement.

🎯 Requirements

• Experience as a Site Reliability Engineer, SRE, or in an equivalent role. • Hands-on experience with cloud environments using GCP, AWS, and/or Azure. • Knowledge of Kubernetes and Docker. • Experience with observability, monitoring, alerting, and APM. • Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD. • Experience managing, investigating, and resolving incidents. • Knowledge of application and infrastructure troubleshooting. • Experience administering Linux environments. • Knowledge of networking, security, performance, and high availability. • Experience with automation and Infrastructure as Code. • Experience with CI/CD pipelines. • Strong communication skills and the ability to work with multidisciplinary teams. • Analytical, proactive, collaborative, and prevention-oriented mindset. • Preferred: Experience with GKE, EKS, or AKS. • Preferred: Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools. • Preferred: Experience with the ELK Stack, Elasticsearch, and Kibana. • Preferred: Knowledge of Terraform and Ansible. • Preferred: Experience with mission-critical environments and distributed systems. • Preferred: Experience working in financial institutions or regulated environments. • Preferred: Experience with capacity management and cloud cost optimization. • Preferred: Knowledge of disaster recovery and business continuity. • Preferred: Experience defining and managing error budgets. • Preferred: Cloud, Kubernetes, or SRE certifications.

🏖️ Benefits

• Meal voucher • Food allowance • Home office allowance • Medical insurance • Dental insurance • Life insurance • Birthday day off • TotalPass / Wellhub app • Boon Saúde • Discount partnerships • Agreements with businesses and educational institutions • Welcome kit • Verity onboarding • Verity Learning Interval • Great Place to Work certification

Apply Now

Similar Jobs

🔥 7 hours ago

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Software Architect building secure .NET applications and AI-enabled solutions. Designing scalable architectures and mentoring developers at CI&T, an enterprise technology transformation company.

ASP.NET

AWS

Cloud

JavaScript

MySQL

Postgres

React

SDLC

SQL

.NET

🕒 Yesterday

Spread Tecnologia

1001 - 5000

💼 Consulting

📣 Marketing

🏥 Healthcare

Analista DevOps na Spread Tecnologia, empresa de inovação e soluções digitais. Configuração de pipelines DevSecOps, infraestrutura, segurança e plataformas corporativas.

🗣️🇧🇷🇵🇹 Portuguese Required

Apache

DNS

Docker

ElasticSearch

Firewalls

Grafana

JUnit

Kubernetes

Linux

Maven

MySQL

NGINX

Node.js

Postgres

Python

Selenium

TCP/IP

VMware

🕒 2 days ago

Stefanini Brasil

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Engenheira SRE aprimorando observabilidade de aplicações críticas na Stefanini, ecossistema latino-americano de tecnologia. Evolução de monitoramento, resposta a incidentes e automação operacional com Datadog, ELK e Python.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Cloud

ElasticSearch

Java

Kubernetes

Python

Terraform

🕒 2 days ago

Viasoft Korp | Industry ERP

51 - 200

🏭 Manufacturing

🤝 B2B

Analista DevOps evoluindo infraestrutura cloud, Kubernetes e CI/CD. Automatizando operações do ERP autônomo da Korp, empresa brasileira de tecnologia para indústria e distribuição.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Cloud

Grafana

Jenkins

Kubernetes

Linux

🕒 2 days ago

Commit

501 - 1000

🔒 Cybersecurity

DevOps Engineer modernizing AWS infrastructure and production support. Transforming CI/CD by moving the company’s AWS codebase into GitHub Actions.

AWS

Docker

EC2

Python

Terraform