Site Reliability Engineer, SRE

Vaga não está no LinkedIn

🕒 4 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $145.000 - $160.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Payliance

Payliance

51 - 200 funcionários

💳 Fintech

🤝 B2B

💰 $31.000.000 Debt Financing - Payliance em 2019-12

Fintech • B2B

A Payliance é uma empresa de tecnologia de pagamentos que fornece uma plataforma de Payments-as-a-Service (PaaS), oferecendo aceitação de pagamentos (ACH, cartões de crédito/débito, pagamentos em tempo real, baseados em cheque), ferramentas de verificação/avaliação de risco e recuperação de dívidas/gestão de contas a receber. Eles atendem comerciantes, credores, agências de cobrança e provedores de BNPL, oferecendo integrações, APIs, programas de parceiros (ISOs, sistemas de gestão de empréstimos) e serviços de compliance/cobrança licenciada. A Payliance enfatiza a redução de custos de processamento, risco de fraude e melhora das taxas de recuperação; processa grandes volumes (reportado mais de $63B anualmente, 162M de transações, 40K locais de comerciantes) e apoia verticais de empréstimos (prime, subprime, acesso a salários ganhos, BNPL) e comerciantes.

Descrição

• Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues. • Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents. • Partner with application engineers to embed reliability into new feature design and deployment practices. • Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early. • Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit. • Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments. • Automate operational tasks, deployment pipelines, and disaster recovery procedures. • Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work. • Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup. • Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly. • Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains. • Capacity plan and scale database infrastructure to support transaction volume growth. • Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing. • Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements. • Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly. • Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence. • Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs. • Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM. • Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning. • Support compliance and security controls relevant to a PCI-regulated payments environment. • Participate in on-call rotation to respond to production incidents and drive swift resolution. • Define and track error budgets; use them to balance velocity and reliability investment. • Communicate status updates clearly during incidents and coordinate cross-functional response. • Maintain and improve runbooks, escalation paths, and on-call health over time. • Collaborate with platform engineering teams on architecture decisions and scalability requirements. • Share observability and reliability best practices with application teams. • Mentor engineers on SRE principles and operational excellence.

🎯 Requisitos

• 4+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role. • C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure. • AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each. • RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution. • Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK. • AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM. • SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager. • Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling. • Excellent communication skills and a collaborative, blameless engineering mindset. • Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.

🏖️ Benefícios

• Performance-based annual bonus. • Medical, Dental, and Vision insurance. • 401(k) with company match. • Generous PTO plus paid company holidays. • Company-paid life and long-term disability insurance. • Paid parental leave.

Candidatar-se

Vagas Similares

🕒 4 dias atrás

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Site Reliability Engineer focusing on building and maintaining cloud infrastructure for Cisco Meraki. Analyzing reliability, troubleshooting, and implementing solutions in a secure environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $167.700 - $245.200 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Valence

51 - 200

🤖 Inteligência Artificial

👥 RH Tech

☁️ SaaS

Senior DevOps Engineer managing AWS infrastructure for a pioneering AI coaching platform. Leading security initiatives and collaborating with multiple development teams on scalable solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

🔥 Investimento no último ano

💰 $50.000.000 Series B - Valence em 2025-09

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

NationsBenefits

1001 - 5000

🏥 Saúde

💼 Consultoria

📦 Logística

Manager, Site Reliability Engineering leading a US-based SRE team at NationsBenefits, a healthcare fintech. Driving operational excellence and mentoring engineers for high-quality service delivery.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

SS&C Technologies

10.000+ funcionários

💼 Consultoria

🛡️ Seguros

📦 Logística

Site Reliability Engineer for a leading financial services and healthcare technology company. Ensuring operational health, reliability, and availability of cloud platforms.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $120.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

SS&C Technologies

10.000+ funcionários

💼 Consultoria

🛡️ Seguros

📦 Logística

Site Reliability Engineer contributing to the operational health of FedRAMP High cloud platform at leading financial services company. This role covers advanced production support responsibilities and reliability engineering.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $130.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório