Site Reliability Engineer – Azure, Observability, Scripting

🔥 0 minutes ago

🌐 Colombia, Brazil, +4 more countries – Remote

info

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Jalasoft

Jalasoft

1001 - 5000 employees

Founded 2003

☁️ SaaS

📚 Education

Software Development • SaaS • Education

Jalasoft is a global nearshore software development company with a strong presence across 70 cities in 13 countries. With a team of over 1000 South American-based software engineers, Jalasoft specializes in software development, quality assurance (QA), and DevOps solutions. The company focuses on staff augmentation and dedicated teams tailored to meet client needs, ensuring quality and efficiency in project delivery. Jalasoft places a strong emphasis on security, holding an ISO 27001 certification, and partners with leading technology firms like Palo Alto, NVIDIA, and Cisco to offer reliable network and data center management. In addition, Jalasoft operates Jala University, offering educational programs in technology to foster and recruit top tech talent. The company aims to drive digital transformation by providing agile, culturally aligned nearshore software solutions.

📋 Description

• Ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes • Improve system availability, monitoring, incident response, and infrastructure automation in production environments • Design and implement observability solutions, including monitoring, metrics, exporters, alerting rules, and dashboards • Define and implement service level indicators, objectives, and error budgets • Design alerting and incident response processes, author runbooks, and support on-call practices • Design and test backup, restore, and disaster recovery solutions against RPO and RTO targets • Troubleshoot Kubernetes workloads, manage resources, and perform cluster upgrades • Read and modify infrastructure as code and Azure DevOps pipelines • Support operational excellence through automation and scripting

🎯 Requirements

• 6+ years of experience • 3+ years of experience operating Kubernetes in production • Site reliability engineering or production operations for Kubernetes workloads at scale • Experience with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy • Experience with Prometheus and Grafana, including metrics, exporters, recording and alerting rules, and dashboard design • Experience defining and implementing service level indicators, objectives, and error budgets • Experience designing alerting and incident response, including runbook authoring and on-call practice • Experience designing and testing backup, restore, and disaster recovery, including validation against RPO and RTO targets • Kubernetes operations experience, including workload troubleshooting, resource management, and cluster upgrades • Ability to read and modify infrastructure as code using Terraform or Bicep and Azure DevOps pipelines • Scripting experience in Python, PowerShell, or Bash • Professional working English

🏖️ Benefits

• Remote work • 13 floating holiday • 15 vacation days per year completed • Good working environment • Equal opportunity employment without distinction based on protected characteristics

Apply Now

Similar Jobs

🕒 5 days ago

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

Senior DevOps Engineer responsible for CI/CD and scaling cloud-native platforms for enterprise applications. Collaborating in a cross-regional team delivering scalable and secure infrastructure.

🗣️🇪🇸 Spanish Required

Ansible

Docker

Jenkins

Kubernetes

Python

Shell Scripting

SQL

Vault

🕒 July 23

Stefanini LATAM

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

DevOps Engineer at Stefanini responsible for automation and CI/CD for telecom solutions. Collaborating with tech teams to ensure smooth service deployment and operations.

🗣️🇪🇸 Spanish Required

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Gradle

Grafana

Jenkins

JMeter

Kubernetes

Linux

Maven

OpenShift

Prometheus

Python

Selenium

Splunk

Terraform

🕒 July 14

DBSync

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Forward Deployment Engineer working with cloud technologies to deliver solutions for enterprise clients. Involves customer discovery, solution design, and deployment execution in Colombia.

AWS

Cloud

Java

Python

SOAP

SQL

Tableau

Go

🕒 July 11

Gorilla Logic

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer building CI/CD pipeline and automated infrastructure for cloud services at Gorilla Logic. Collaborating with teams to enhance deployment processes and manage configuration effectively.

Ansible

AWS

Chef

Cloud

Docker

Groovy

Kubernetes

Linux

Puppet

Python

Shell Scripting

Terraform

Go

🕒 July 5

Gorilla Logic

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer responsible for Azure cloud infrastructure and CI/CD pipeline management. Collaborating with engineering teams to ensure reliable and efficient delivery of cloud-native applications.

Azure

Cloud

Docker

JavaScript

Kubernetes

Node.js

Python

React

Terraform

Vault

.NET