Senior Platform DevOps Engineer

🔥 3 hours ago

🌐 Colombia, Costa Rica – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gorilla Logic

Gorilla Logic

501 - 1000 employees

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

Gorilla Logic is a company renowned for its expertise in modern software and data engineering. Serving as a strategic partner rather than just a vendor, Gorilla Logic specializes in digital product design, cloud engineering, data and AI delivery, DevOps, quality assurance, and legacy modernization. With a team of skilled digital product designers, solutions architects, and Agile nearshore teams, Gorilla Logic has been instrumental in developing business-critical software applications for Fortune 500 and SMB companies for over 20 years. Their services include creating SaaS platforms, enhancing digital experiences, and providing flexible, security-focused solutions. Gorilla Logic operates with teams located in Costa Rica, Colombia, Mexico, and the United States, emphasizing collaborative partnerships to deliver cutting-edge digital engineering solutions.

📋 Description

• Manage, operate, and troubleshoot production Kubernetes environments and workloads • Build, maintain, and improve infrastructure using Terraform and Infrastructure as Code practices • Manage application deployments, upgrades, configuration changes, and complex deployment lifecycles • Work with GitOps-based deployment processes and tools such as Argo CD • Develop and maintain Python scripts and tooling for platform automation and operational workflows • Troubleshoot and resolve production incidents across infrastructure, applications, and platform services • Improve platform reliability, scalability, observability, and operational efficiency • Support Kubernetes scaling and autoscaling strategies for production workloads • Collaborate with development and platform teams to resolve infrastructure and deployment challenges • Participate in technical decisions and identify risks, reliability concerns, and improvement opportunities • Maintain high engineering and quality standards, providing technical pushback when necessary • Take ownership of platform initiatives and drive issues through resolution with minimal supervision

🎯 Requirements

• Strong hands-on experience managing Kubernetes in production environments • Experience with Kubernetes deployments, scaling, troubleshooting, and operational management • Hands-on experience with Terraform for provisioning and managing cloud infrastructure • Ability to read and write Python for scripting, automation, troubleshooting, and platform tooling • Experience with AWS cloud infrastructure, ideally including EKS or similar managed Kubernetes environments • Experience with GitOps practices and deployment tools such as Argo CD • Experience managing complex application deployment and upgrade lifecycles • Proven experience troubleshooting, triaging, and supporting production incidents • Strong understanding of infrastructure reliability, scalability, and operational best practices • Strong problem-solving skills and the ability to independently investigate complex production issues • Strong sense of ownership and accountability, with the ability to operate effectively with limited supervision • Quality-first mindset with the judgment to balance delivery speed, reliability, and long-term maintainability • Strong communication and collaboration skills • Preferred: Experience with Helm and Kubernetes package/deployment management • Preferred: Familiarity with PyTorch and Hugging Face Transformers • Preferred: Experience supporting GPU-based workloads or ML inference platforms • Preferred: Familiarity with NVIDIA Triton Inference Server • Preferred: Experience with Chainguard, distroless container images, Trivy, or container vulnerability reduction • Preferred: Experience implementing or improving Kubernetes autoscaling solutions • Preferred: Familiarity with streaming or messaging platforms such as Apache Kafka or similar technologies • Preferred: Experience with Elasticsearch or ArangoDB • Preferred: Experience troubleshooting complex service-to-service networking • Preferred: Exposure to OpenShift, IL5, FedRAMP, or similarly constrained environments • Preferred: Familiarity with AI/ML or agentic AI development environments

🏖️ Benefits

• Remote work arrangement • Full-time employment

Apply Now

Similar Jobs

🕒 September 22

Moovx

11 - 50

💼 Consulting

📣 Marketing

🏢 Enterprise

DevOps Engineer configuring AWS infrastructure, Databricks Ops, CI/CD, and observability. Supporting enterprise platform reliability for MOOVX’s Latin American operations.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

🕒 September 14

Abstra

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior DevOps Engineer designing secure cloud-native products and developer platforms for Abstra, a Latin American tech talent services company. Improving automation, observability, reliability, and service delivery.

🇨🇴 Colombia – Remote

💰 $2.3M Seed Round - Abstra on 2022-01

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Docker

EC2

Jenkins

Kubernetes

Linux

Prometheus

Python

Splunk

Terraform

Go

🕒 September 12

knowmad mood

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Consultor DevSecOps AWS gestionando seguridad, automatizaciĂłn y migraciones de datos cloud para knowmad mood. OperaciĂłn de AWS DMS, IaC y pipelines CI/CD en LATAM.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🗣️🇪🇸 Spanish Required

AWS

Cloud

EC2

Terraform

🕒 August 31

Commit

501 - 1000

🔒 Cybersecurity

Senior DevOps Engineer modernizing AWS infrastructure for the company. Migrating AWS codebases to GitHub Actions and strengthening CI/CD, observability, security, and production support.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Docker

EC2

Python

Terraform

🕒 August 26

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

Especialista DevOps remoto en GFT, consultora de transformaciĂłn digital y software. Implementando infraestructura GCP/Terraform compatible con FedRAMP para una plataforma de manufactura habilitada por IA.

🗣️🇪🇸 Spanish Required

Cloud

Google Cloud Platform

Terraform