Senior DevOps Engineer

Job not on LinkedIn

đŸ”„ 1 minute ago

đŸ‡§đŸ‡· Brazil – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Aiphoria

Aiphoria

51 - 200 employees

Founded 2022

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

💰 $34M Series A - Aiphoria on 2025-07

Consulting ‱ Logistics ‱ Marketing

Aiphoria is a provider of AI-driven virtual employees and a proprietary platform that automates multi-channel customer-facing and back-office tasks. Its Aiphoria Pros are multilingual, multi-modal agents (voice, chat, email, calls) that handle support, sales, collections, legal, HR and marketing workflows for enterprise customers in industries such as banking, telecom and e‑commerce. The company offers cloud and on‑prem deployments, analytics and open-architecture integrations to reduce human labor, improve response times and boost customer satisfaction.

📋 Description

‱ Deploy, operate, and evolve a microservices-based platform running in Kubernetes clusters across AWS, GCP, and on-premises Rancher ‱ Operate and support GPU-based ML inference services using Triton Inference Server and vLLM on RunPod, Scaleway, and Nebius ‱ Build and maintain Docker images for all microservices and ensure a stable service lifecycle ‱ Maintain and scale development and production Kubernetes clusters ‱ Participate in deployment debugging, incident investigation, and performance troubleshooting ‱ Develop, maintain, and evolve custom Helm charts for each service ‱ Design and operate CI/CD pipelines using GitHub and GitLab for on-premises customer deployments ‱ Ensure platform compliance with SOC 2 requirements and improve security and compliance processes ‱ Manage cluster access via NetBird VPN and implement role-based access control using group policies ‱ Deploy and manage infrastructure using Terraform and Ansible IaC practices ‱ Develop and continuously improve observability systems using Grafana, Prometheus, and the ELK stack ‱ Continuously optimize infrastructure across IaC, IAM, observability, and CI/CD ‱ Work with Python, Kubernetes, Linux, Docker, GitHub CI/CD, PostgreSQL, ClickHouse, Kafka, Superset, Terraform, and Ansible

🎯 Requirements

‱ Minimum 5 years of experience in a DevOps and/or Site Reliability Engineering role ‱ Strong hands-on experience with Linux system administration ‱ Extensive experience deploying, operating, and scaling Kubernetes in both cloud and bare-metal environments ‱ Deep expertise and practical experience with at least one major cloud provider, preferably Google Cloud Platform ‱ Experience with ML inference on GPU/CPU is a strong plus ‱ Proven experience implementing SRE practices and building observability stacks using Grafana, Prometheus, and Loki ‱ Strong adherence to GitOps, Infrastructure as Code (IaC), and CI/CD principles ‱ Advanced expertise in Terraform, Ansible, and Python ‱ Ability to work in high-uncertainty environments and rapidly learn new technologies and patterns ‱ Ability to look beyond DevOps tasks and actively debug and understand the product ‱ Ability to choose technologies and architectural approaches based on long-term goals ‱ English proficiency implied by private English lessons, but no explicit language requirement stated

đŸ–ïž Benefits

‱ Award-winning AI products and cutting-edge technology stack ‱ Fully remote ‱ 21 vacation days + public holidays + 5 sick days ‱ Private English lessons via Preply ‱ Startup pace with enterprise stability ‱ Fast career progression ‱ Real ownership and direct impact of work

Apply Now

Similar Jobs

đŸ”„ 14 hours ago

KnowBe4

1001 - 5000

🔒 Cybersecurity

☁ SaaS

📚 Education

Senior Site Reliability Engineer building reliable AWS infrastructure for KnowBe4’s workforce security platform. Improving Terraform, CI/CD, observability, and distributed systems in Brazil.

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

JavaScript

Python

Ruby

Terraform

🕒 Yesterday

ALTASNET

51 - 200

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

Engenheiro DevOps SĂȘnior operando AWS EKS, Kubernetes e Helm para a Altasnet, empresa de soluçÔes de TI. Modernização de PHP e migração Percona MySQL para Amazon Aurora.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

Apache

AWS

Kubernetes

MySQL

PHP

SQL

Terraform

🕒 2 days ago

CI&T

5001 - 10000

đŸ’Œ Consulting

đŸ„ Healthcare

📣 Marketing

Senior DevOps Engineer designing scalable cloud infrastructure and CI/CD systems. Supporting CI&T’s AI-driven enterprise technology transformation and platform reliability.

AWS

Azure

Cloud

Docker

Jenkins

Kubernetes

Python

Terraform

🕒 2 days ago

Attus Procuradoria Digital

51 - 200

đŸ’Œ Consulting

⚖ Legal

đŸ€– Artificial Intelligence

DevOps sustentando e automatizando Kubernetes multi-cliente na Attus. Empresa brasileira de procuradoria digital usa IA para gestĂŁo de processos jurĂ­dicos.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Cloud

DNS

Docker

ElasticSearch

Flux

Grafana

Java

Kafka

Kubernetes

Linux

Oracle

Postgres

Prometheus

Python

Redis

Terraform

Vault

🕒 5 days ago

Verity Group

51 - 200

đŸ’Œ Consulting

đŸ„ Healthcare

đŸ›Ąïž Insurance

SRE Engineer operando cloud, Kubernetes e observabilidade para a Verity, consultoria de transformação e engenharia digital. Prevenção de incidentes, automação e evolução de ambientes resilientes.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

ElasticSearch

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Terraform