Senior Site Reliability Operations Engineer – Finance

🔥 1 hour ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Truelogic Software

Truelogic Software

501 - 1000 employees

Founded 2004

☁️ SaaS

🤝 B2B

🏢 Enterprise

SaaS • B2B • Enterprise

Truelogic Software is a nearshore software development company specializing in agile staff augmentation services. They focus on providing custom outsourced software development with a team of highly skilled engineers from Latin America. Truelogic Software partners with both startups and Fortune 500 companies, offering solutions that align with their clients' time zones and ensuring high-quality outcomes through collaboration and responsiveness. With a presence in over 25 countries, Truelogic emphasizes remote work for better quality of life, and their engineers are experienced in various industries, delivering a wide range of successful projects globally.

📋 Description

• Lead incident response as Incident Commander, coordinating teams, communications, and service restoration • Produce executive-level incident reports and conduct root cause analyses • Drive continuous improvement following incidents • Monitor and improve observability using AWS CloudWatch, New Relic, Nagios, and SumoLogic • Reduce alert noise and observability gaps • Provide hands-on system support across Linux and Windows environments • Troubleshoot complex infrastructure issues • Manage and execute deployments through Jenkins, GitLab, or similar CI/CD platforms • Own infrastructure initiatives including migrations, upgrades, and process improvements • Enforce change management and risk assessment for production changes • Maintain documentation and standard operating procedures • Liaise between engineering teams and external vendors • Support 24/7 stability of internal IT infrastructure and mission-critical backend systems

🎯 Requirements

• 5+ years of experience in Windows and Linux environments with proven troubleshooting capabilities • Strong knowledge of AWS CloudWatch, New Relic, Nagios, and SumoLogic • Practical experience with Jenkins and GitLab CI/CD tools • Practical experience with CommVault and AWS Backup • Strong scripting skills in PowerShell, Python, or equivalent • Outstanding communication skills, especially under pressure, including executive reporting • Experience in high-paced environments and with on-call support models • Autonomous and proactive attitude; capable of managing complex tasks independently • Availability for a 1-week on-call rotation, including potential critical incident call-ins between 6:00 PM and 6:00 AM PT

🏖️ Benefits

• 100% Remote Work • Highly Competitive USD Pay • Paid Time Off • Work with Autonomy • Work with Top American Companies • Engagement activities • Work-life balance • Collaboration with a diverse, multicultural global network • Work alongside seasoned senior professionals

Apply Now

Similar Jobs

🕒 Yesterday

NDD Tech | Brasil

501 - 1000

📦 Logistics

🏭 Manufacturing

💼 Consulting

Analista SRE Pleno garantindo disponibilidade e confiabilidade da NDDPay, instituição de pagamento da NDD. Monitorando incidentes, infraestrutura Cloud, banco de dados e mudanças operacionais.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

SQL

🕒 Yesterday

Verity Group

51 - 200

💼 Consulting

🤖 Artificial Intelligence

🔒 Cybersecurity

SRE Engineer improving cloud reliability, observability, and incident response for Verity, a digital transformation and engineering consultancy. Automating resilient Kubernetes and Docker environments.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

AWS

Azure

Cloud

Docker

ElasticSearch

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Terraform

🕒 2 days ago

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Site Reliability Engineer operating monitoring platforms and improving cloud reliability for CI&T's enterprise AI transformation solutions. Driving observability, incident response, security remediation, and service-level measurement.

AWS

Azure

Cloud

Cyber Security

DNS

Google Cloud Platform

Grafana

JavaScript

Linux

Python

Splunk

Terraform

TypeScript

Go

🕒 2 days ago

Franq

51 - 200

🛡️ Insurance

💼 Consulting

💳 Fintech

DevOps/SRE administrando Kubernetes, AWS/GCP e automação de infraestrutura na Franq. Fortalecendo observabilidade, CI/CD, disponibilidade e confiabilidade de sistemas financeiros.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Apache

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Java

Kubernetes

Linux

NoSQL

Prometheus

Python

SQL

Terraform

🕒 5 days ago

SoftDesign

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

DevOps Engineer building Azure platforms, Kubernetes infrastructure, and CI/CD automation at SoftDesign. Advancing observability, DevSecOps, and developer experience.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

Cloud

Grafana

Kubernetes

Prometheus

Terraform

Vault