Senior Platform Engineer – DevOps, Infrastructure and Platform

🕒 June 23

🇧🇷 Brazil – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 35%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of OZmap

OZmap

11 - 50 employees

☁️ SaaS

📡 Telecommunications

🤝 B2B

SaaS • Telecommunications • B2B

OZmap is a SaaS platform that provides GIS-based mapping and management for fiber-optic (FTTH) and hybrid network infrastructures, tailored for internet service providers (ISPs). It centralizes georeferenced network documentation, planning, monitoring (including OTDR and switch integration), mobile field apps, and open APIs to streamline provisioning, reduce operational costs, and speed repairs. OZmap also offers data migration, integrations with CRMs/ERPs, dashboards, and tools for commercial viability checks to support network growth, M&A and day-to-day operations.

📋 Description

• Design, operate and evolve AWS (EC2) and on-premises environments with containers (Docker), ensuring availability, security and scalability; • Operate and administer Linux production environments (systemd, kernel/network tuning, I/O, process troubleshooting); • Build and evolve CI/CD pipelines from scratch, including quality and security gates; • Develop end-to-end observability (instrumentation, exporters, PromQL, SLI/SLO, alerts); • Lead advanced troubleshooting, root cause analysis and blameless post-mortems — driving structural change afterwards, not just producing a report; • Implement automation using Infrastructure as Code; • Analyze and optimize cloud costs: rightsizing, usage analysis and proposing data-driven alternatives; • Act as a technical reference for developers and engineers, influencing architecture without relying on formal authority.

🎯 Requirements

• Required: production experience operating core primitives in AWS (~4+ years): EC2, VPC/networking, IAM and security — production operation and technical decision-making; • Linux and networking (~4+ years): server administration and production troubleshooting — disk full, OOM killer, network diagnostics; processes, memory and I/O; • CI/CD built from scratch (~3+ years): pipelines created and evolved by you (GitHub Actions, Jenkins, self-hosted runners, secrets, caching, gates); • End-to-end open-source observability (~2+ years): Prometheus, Grafana, Loki, VictoriaMetrics or equivalents — configured and operated by you, not just used. OpenTelemetry — including instrumentation, exporters, PromQL and SLI/SLO definition; • Operation under managed layers: concrete experience with nginx/HAProxy/Envoy, Linux underneath, and leading the resolution of critical incidents you have driven; • Docker in production (~3+ years): real operation of containers in critical environments — volumes, networking, resource management, graceful shutdown of services; • High autonomy: receives an ambiguous problem ("our observability is weak") and delivers end-to-end; • Ownership and proactivity: anticipates problems before they become incidents; • Clear communication and technical influence, connecting development, infrastructure and business teams; • Conducts post-mortems focused on root cause, organizational learning and continuous improvement, without a blame culture; • Maturity to self-manage while working remotely.

🏖️ Benefits

• 💻 Equipment allowance – to ensure a comfortable work setup; • 💚 Health support – because your well-being matters; • 📚 Education support – we support your continuous development journey; • 🎂 Birthday gift – because we like to celebrate together; • 🏅 Recognition for tenure – your time with us is valued; • 🗣️ Language support – to help you go beyond borders; • 🏋️ TotalPass (for employee use only); • 🌴 Paid leave after 12 months of employment; • 🎉 Online integration events and socials.

Apply Now

Similar Jobs

🕒 June 12

In All Media

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps Engineer focusing on migrating workloads from AWS to Azure for a clean energy solutions provider. Leading optimization of cloud environments and deployment workflows.

AWS

Azure

Cloud

Docker

EC2

Jenkins

Kubernetes

Terraform

🕒 June 12

Swile

201 - 500

💳 Fintech

👥 HR Tech

🤝 B2B

Senior Site Reliability Engineer at Swile providing innovative solutions in Fintech, Travel, HR, and Employee Benefits. Focused on problem-solving while enhancing the developer experience in a remote role.

🗣️🇧🇷🇵🇹 Portuguese Required

🕒 May 11

Jusbrasil

201 - 500

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Site Reliability Engineer at Jusbrasil improving the integrity and performance of product systems. Focusing on data-driven SRE practices and collaborating closely with product teams.

🗣️🇧🇷🇵🇹 Portuguese Required

ElasticSearch

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Terraform

🕒 May 8

Keyrus

1001 - 5000

🤝 B2B

💼 Consulting

SRE Engineer defining practices and improving the availability and performance of systems. Working with automation and observability strategies in a diverse and inclusive environment.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Grafana

Prometheus

Terraform

🕒 May 8

Keyrus

1001 - 5000

🤝 B2B

💼 Consulting

DevOps Analyst responsible for designing and maintaining CI/CD pipelines within Keyrus. Collaborating with teams to enhance cloud environments and implement best practices in security and reliability.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Jenkins

Kubernetes

Linux

Terraform