Site Reliability Engineer

🕒 6 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of CXM

CXM

201 - 500 funcionários

Fundada em 2015

💸 Finanças

💳 Fintech

Finance • Fintech

A CXM é uma empresa de serviços financeiros online que opera como uma plataforma de corretagem, oferecendo negociação de Forex e CFD aos clientes. A empresa impõe restrições de acesso regional para conformidade regulatória e fornece avisos de risco sobre a natureza de alto risco da negociação com alavancagem. A CXM também oferece suporte ao cliente para consultas de acesso e conformidade.

Descrição

• Own the day-to-day reliability of our .NET/C# services running on Windows. • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions. • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues. • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases. • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact. • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure. • Improve deployment safety, release automation, and rollback strategies. • Partner with developers to improve application operability, resilience, and fault isolation. • Automate operational tasks through scripting and infrastructure automation. • Create and maintain runbooks, operational documentation, and incident response procedures. • Continuously improve monitoring, alert quality, automation, and platform reliability.

🎯 Requisitos

• 3–5 years of experience in Site Reliability Engineering or related field • Strong experience debugging and supporting .NET/C# applications in production. • Hands-on experience with Windows Server environments. • Strong PowerShell scripting skills. • Experience with Python or Bash. • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools). • Experience with modern CI/CD pipelines. • Knowledge of deployment strategies, release automation, and rollback mechanisms. • Experience working with AWS. • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools. • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms. • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.

🏖️ Benefícios

• Work on mission-critical trading infrastructure that directly impacts customers. • Solve challenging reliability and scalability problems in a real-time environment. • Build world-class observability, automation, and deployment practices. • Collaborate with experienced engineers in a modern engineering culture. • Influence reliability strategy and engineering best practices across the platform.

Candidatar-se

Vagas Similares

🕒 6 dias atrás

NVIDIA

10.000+ funcionários

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

DevOps Engineer supporting NVIDIA’s Rapids project for AI and data science initiatives. Collaborating with teams to ensure high-quality software releases and infrastructure maintenance.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consultoria

📦 Logística

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Mirantis

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer handling high-performance storage for AI platforms at Mirantis. Integrating and operating storage solutions within Kubernetes environments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

NEC Software Solutions

5001 - 10000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior DevOps Engineer managing cloud infrastructure and CI/CD for mission-critical applications. Seeking someone to collaborate with teams and maintain secure, scalable solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 25

Chromatic (we're hiring!)

11 - 50

☁️ SaaS

⚡ Produtividade

🏢 Corporativo

DevOps Engineer focused on securing and stabilizing Chromatic's AWS infrastructure. Collaborating across teams to enhance security posture and developer workflows within a remote environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $155.000 - $190.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório