SRE Specialist

Job not on LinkedIn

🔥 9 minutes ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Quality Digital

Quality Digital

1001 - 5000 employees

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

Quality Digital is a leading technology company in Brazil with over 34 years of experience, specializing in accelerating digital transformation for businesses. They provide innovative digital solutions that enhance operational performance, governance, and customer communication through the use of specialized teams and digital platforms. Their services include IT optimization, automation, digital strategy, and AI analytics to create high-value experiences across multiple channels.

📋 Description

• Work in Site Reliability Engineering, with a strong focus on observability • Collaborate across different technology teams to map, understand, and monitor the complete chain of services and systems • Lead technical assessments and discovery workshops with Infrastructure, Networking, Database, Cloud, Security, Development, Architecture, Integration, Middleware, API, and Systems teams • Map the architecture and dependency chain of key systems • Build end-to-end Service Mapping, including infrastructure, applications, integrations, APIs, databases, queues, external services, and other dependencies • Identify integration paths and assess the impact of components on service availability • Link technical components to their respective business services • Identify monitoring and observability gaps • Define and implement an observability strategy covering metrics, logs, traces, events, digital experience, and availability • Create service-oriented monitoring to track complete transaction and integration flows • Define reliability indicators such as SLIs, SLOs, and service availability • Build dashboards, alerts, correlations, and diagnostic mechanisms • Reduce MTTD and MTTR and accelerate root-cause identification • Participate in corporate projects from the early stages, ensuring appropriate observability requirements • Support the evolution of the observability culture and promote best practices among technical teams

🎯 Requirements

• Bachelor’s degree or equivalent higher education qualification • Knowledge of application architecture and distributed systems • Knowledge of server infrastructure and operating systems • Knowledge of networks and communication protocols • Knowledge of databases • Knowledge of REST APIs and system integrations • Knowledge of microservices • Knowledge of containers and Kubernetes • Knowledge of cloud and hybrid environments • Knowledge of queues and messaging • Knowledge of load balancers, proxies, and gateways • Knowledge of DNS, HTTP/HTTPS, TCP/IP, and TLS • Knowledge of APM and Distributed Tracing • Knowledge of centralized log management and analysis • Knowledge of infrastructure metrics and monitoring • Knowledge of Service Mapping and Dependency Mapping • Knowledge of SLI, SLO, SLA, and Error Budget concepts • Knowledge of Incident Management and root-cause analysis • Intermediate-level English is desirable • Experience with Datadog, Elastic Stack/Elasticsearch/Logstash/Kibana/Elastic Agent, Sensedia/API Management, Zabbix, OpenTelemetry, Prometheus, and Grafana is desirable

🏖️ Benefits

• An environment conducive to learning and professional growth • Performance reviews and feedback focused on continuous development • Meal and/or food allowance • Medical and dental insurance • Pharmacy partnerships offering discounts on medications • Childcare assistance in accordance with current policy • Partnership with SESC, offering a variety of cultural and leisure activities • Partnerships for language and technology studies, as well as access to a course platform • Payroll-deducted loans at attractive rates • Financial education program • Corporate University and learning paths • Referral Program with opportunities for prizes and bonuses • Group life insurance

Apply Now

Similar Jobs

🕒 Yesterday

RD Station

1001 - 5000

📣 Marketing

💼 Consulting

📦 Logistics

Analista de SRE na RD Station, empresa de soluções para crescimento e simplificação de negócios. Evoluindo confiabilidade, observabilidade e automação de infraestrutura em nuvem.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Kubernetes

Linux

Redis

🕒 Yesterday

GT

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE improving observability, incident response, and reliability for Feeld’s inclusive dating app. Building dependable systems with Node.js, TypeScript, and AWS.

AWS

JavaScript

Node.js

TypeScript

🕒 Yesterday

Inmetrics

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

Especialista SRE garantindo disponibilidade, segurança e desempenho de sistemas na Inmetrics. Automatizando infraestrutura cloud, observabilidade, Kubernetes e práticas DevOps para clientes em transformação digital.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Azure

Cloud

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Prometheus

Python

Splunk

Terraform

Go

🕒 Yesterday

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior DevOps Engineer maintaining Azure, Kubernetes, CI/CD, and GitOps infrastructure. Supporting CI&T’s enterprise AI transformation solutions and international client operations.

Azure

Cloud

Docker

Grafana

Jenkins

Kubernetes

Postgres

Prometheus

Python

Terraform

Vault

🕒 2 days ago

opinov8

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior GCP DevOps Engineer supporting cloud infrastructure for a media-technology platform powering TV, radio, and digital advertising. Automating deployments, monitoring systems, and resolving production issues.

AWS

Azure

Cassandra

Cloud

Docker

Google Cloud Platform

Hadoop

Java

Jenkins

Linux

NGINX

PHP

Subversion