Cloud Engineer – Site Reliability Engineer, Senior

🔥 5 minutes ago

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GrooveTech

GrooveTech

51 - 200 employees

Founded 2017

🤝 B2B

🏢 Enterprise

🎯 Recruiter

B2B • Enterprise • Recruitment

GrooveTech is a Brazilian IT services firm that provides managed staff augmentation and squads, software quality assurance, 24x7 NOC monitoring, and strategic IT consulting including digital due diligence for M&A. They supply managed teams (Executives of IT, PM/PO, technical leads and business partners), daily reporting, and focus on ROI by reducing turnover and improving project performance through testing automation, monitoring (on-prem and cloud), and tailored delivery models.

📋 Description

• Lead the technical evolution of infrastructure on GCP, primarily GKE, Compute Engine, Cloud SQL, Cloud Storage, VPC, Load Balancing, IAM, Cloud DNS, Secret Manager, Pub/Sub, and BigQuery • Design architectures with a focus on high availability, scalability, resilience, security, performance, and cost efficiency • Apply SRE best practices by defining and tracking SLIs, SLOs, and Error Budgets • Monitor reliability indicators such as availability, latency, throughput, error rate, saturation, capacity, MTTR, and MTBF • Administer and perform advanced troubleshooting on Kubernetes/GKE in production • Define autoscaling strategies, capacity planning, resource management, and workload isolation • Evolve the observability strategy using New Relic, OpenTelemetry, Prometheus, Grafana, and native cloud tools • Create dashboards, alerts, and technical and business metrics to ensure actionable observability • Serve as Incident Commander during critical incidents and lead blameless postmortems • Manage and evolve infrastructure as code using Terraform • Develop and maintain reusable modules, configuration patterns, state management, and Terraform pipelines • Manage CI/CD pipelines with GitHub Actions and deployment processes using ArgoCD/GitOps • Implement rolling, blue/green, and canary deployment strategies • Develop automation using Go, Python, and Bash • Build Platform Engineering solutions: self-service interfaces, templates, automations, and golden paths • Work with cloud networking, security, governance, IAM/RBAC, secrets, certificates, and encryption • Contribute to FinOps initiatives, resource optimization, rightsizing, and waste reduction • Design and validate Disaster Recovery strategies including backup, restore, failover, RPO, and RTO • Perform performance and capacity planning to anticipate bottlenecks and growth requirements

🎯 Requirements

• 5+ years of experience managing critical production infrastructure in a public cloud environment • Strong experience with Google Cloud Platform (GCP) • Hands-on experience with Kubernetes in production environments, preferably GKE • Advanced knowledge of Terraform, including modules, state management, versioning, and environment organization • Hands-on experience defining and operating SLIs, SLOs, and Error Budgets • Advanced troubleshooting skills in Linux, networking, DNS, HTTP, TLS, load balancers, and Kubernetes • Experience automating tasks using Go, Python, and/or Bash • Experience with GitHub Actions, ArgoCD, and GitOps practices • Experience driving critical incident response, postmortems, and root-cause analysis • Ability to make architectural decisions and communicate technical, operational, and financial trade-offs • GCP is required • AWS and/or Azure are desirable • Nice-to-haves: experience with multiple cloud providers; multi-cloud or hybrid cloud environments; multi-region architectures and Disaster Recovery; AWS EKS, RDS/Aurora and multi-account setups; Azure AKS and Landing Zones; Service Mesh (Istio, Linkerd, or Anthos Service Mesh); API Gateway and API management; Chaos Engineering; FinOps; Platform Engineering

🏖️ Benefits

• Contract (PJ) • Remote work model • Wellhub • Life insurance

Apply Now

Similar Jobs

🕒 August 5

Grupo Adriano Cobuccio

1001 - 5000

📦 Logistics

💼 Consulting

🏭 Manufacturing

Engenheiro DevOps sênior construindo pipelines CI/CD, orquestrando Kubernetes e automatizando infraestrutura cloud para uma instituição de pagamentos. Garantindo alta disponibilidade, segurança, performance e escalabilidade em ambientes de produção.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Terraform

🕒 July 28

In All Media

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

DevOps Manager leading Azure cloud and Kubernetes environments for a power management platform. Managing CI/CD, incident lifecycle, and a technical team within a remote LATAM context.

Azure

Cloud

Grafana

Kubernetes

🕒 July 21

Uberall

201 - 500

🍽️ Food & Beverage

🏨 Hospitality

🚘 Automotive

DevSecOps Engineer building, maintaining, and scaling secure cloud infrastructure and CI/CD pipelines with a security-first mindset. Collaborating with a small team to ensure platform security and compliance.

AWS

Cloud

EC2

Kubernetes

Microservices

Shell Scripting

Terraform

🕒 July 19

Metal Toad

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Cloud Engineer designing and maintaining scalable AWS infrastructure for critical applications at Metal Toad. Seeking talented professionals with strong AWS and cloud knowledge in a remote setup

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Linux

Python

TCP/IP

🕒 July 16

CIAL Dun & Bradstreet

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

SRE role managing critical incidents and observability for SaaS solutions in Latin America. Responsibilities include monitoring, incident response, and improving platform observability.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

Grafana

Prometheus

Python