Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CI&T

CI&T

5001 - 10000 employees

Founded 1995

💼 Consulting

🏥 Healthcare

📣 Marketing

💰 $5.5M Venture Round on 2014-04

Consulting • Healthcare • Marketing

CI&T is a global tech transformation specialist focusing on helping organizations navigate their technology journey. With services spanning from application modernization and cloud solutions to AI-driven data analytics and customer experience, CI&T empowers businesses to accelerate their growth and maximize operational efficiency. The company emphasizes digital product design, strategy consulting, and immersive experiences, ensuring a robust support system for enterprises in various industries.

📋 Description

• Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring • Proactively analyze logs, error tracking, traces, and metrics to find failures, regressions, and anomalies; triage, reproduce, and drive issues to resolution • Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring customer-impacting issues reach the right people • Instrument services with meaningful metrics, structured logs, and distributed traces • Partner with engineers to improve code observability • Write post-incident reviews, track remediation items, and feed lessons learned into monitors, runbooks, and system design • Define, measure, and report service level indicators for key customer-facing services • Improve reliability, scalability, and cost efficiency of cloud infrastructure, CI/CD pipelines, and release processes • Automate repetitive operational work • Maintain and improve runbooks, escalation paths, and operational documentation • Partner with enterprise InfoSec to remediate cybersecurity risks, including SSL/TLS cleanup, security headers, and DNS configuration hygiene • Track security findings from scans and audits through verified closure • Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features

🎯 Requirements

• 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role • Hands-on experience administering and building in platforms such as New Relic, Grafana, Splunk, or Dynatrace, including dashboards, monitors, log management, and APM • Strong troubleshooting and root-cause analysis skills • Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash • Comfort reading application code to understand failures • Experience operating services in AWS, GCP, or Azure • Solid grasp of networking, containers, and Linux fundamentals • Familiarity with infrastructure as code such as Terraform, CloudFormation, or Pulumi • Familiarity with CI/CD tooling such as GitHub Actions • Experience with on-call responsibilities, incident response, and post-incident review processes • Clear written and verbal communication, including explaining reliability concerns and trade-offs to non-technical stakeholders

🏖️ Benefits

• Health and dental insurance • Meal and food allowance • Childcare assistance • Extended paternity leave • Partnership with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass • Profit Sharing and Results Participation (PLR) • Life insurance • Continuous learning platform (CI&T University) • Discount club • Free online platform dedicated to physical, mental, and overall well-being • Pregnancy and responsible parenting course • Partnerships with online learning platforms • Language learning platform • Inclusion support and accommodations during the selection process • Dedicated Health and Well-being team, inclusion specialists, and affinity groups

Apply Now

Similar Jobs

🔥 5 hours ago

Franq

51 - 200

🛡️ Insurance

💼 Consulting

💳 Fintech

DevOps/SRE administrando Kubernetes, AWS/GCP e automação de infraestrutura na Franq. Fortalecendo observabilidade, CI/CD, disponibilidade e confiabilidade de sistemas financeiros.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Apache

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Java

Kubernetes

Linux

NoSQL

Prometheus

Python

SQL

Terraform

🕒 3 days ago

SoftDesign

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

DevOps Engineer building Azure platforms, Kubernetes infrastructure, and CI/CD automation at SoftDesign. Advancing observability, DevSecOps, and developer experience.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

Cloud

Grafana

Kubernetes

Prometheus

Terraform

Vault

🕒 4 days ago

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

Especialista DevOps construindo e arquitetando plataformas AKS multiambiente na GFT Technologies. Automatizando infraestrutura Azure, alta disponibilidade e migração de aplicaçþes para contêiner.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

DNS

Docker

Flux

Kubernetes

Node.js

Terraform

Vault

🕒 4 days ago

Aiphoria

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior DevOps Engineer operating Kubernetes, cloud, and GPU inference infrastructure for an AI product company. Building secure, observable, and scalable remote production platforms.

Ansible

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Kafka

Kubernetes

Linux

Microservices

Postgres

Prometheus

Python

Terraform

🕒 4 days ago

KnowBe4

1001 - 5000

🔒 Cybersecurity

☁️ SaaS

📚 Education

Senior Site Reliability Engineer building reliable AWS infrastructure for KnowBe4’s workforce security platform. Improving Terraform, CI/CD, observability, and distributed systems in Brazil.

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

JavaScript

Python

Ruby

Terraform