Site Reliability Engineer, SRE

Job not on LinkedIn

🔥 47 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Payliance

Payliance

51 - 200 employees

💳 Fintech

🤝 B2B

💰 $31M Debt Financing - Payliance on 2019-12

Fintech • B2B

Payliance is a payments technology company that provides a Payments-as-a-Service (PaaS) platform offering payment acceptance (ACH, credit/debit cards, real-time payments, check-based), verification/risk assessment tools, and debt recovery/accounts receivable management. They serve merchants, lenders, collection agencies, and BNPL providers, offering integrations, APIs, partner programs (ISOs, loan management systems), and compliance/licensed collection services. Payliance emphasizes reducing processing costs, fraud risk, and improving recovery rates; processes large volumes (reported $63B+ annually, 162M transactions, 40K merchant locations) and supports lending verticals (prime, subprime, earned wage access, BNPL) and merchants.

📋 Description

• Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues. • Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents. • Partner with application engineers to embed reliability into new feature design and deployment practices. • Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early. • Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit. • Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments. • Automate operational tasks, deployment pipelines, and disaster recovery procedures. • Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work. • Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup. • Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly. • Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains. • Capacity plan and scale database infrastructure to support transaction volume growth. • Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing. • Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements. • Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly. • Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence. • Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs. • Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM. • Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning. • Support compliance and security controls relevant to a PCI-regulated payments environment. • Participate in on-call rotation to respond to production incidents and drive swift resolution. • Define and track error budgets; use them to balance velocity and reliability investment. • Communicate status updates clearly during incidents and coordinate cross-functional response. • Maintain and improve runbooks, escalation paths, and on-call health over time. • Collaborate with platform engineering teams on architecture decisions and scalability requirements. • Share observability and reliability best practices with application teams. • Mentor engineers on SRE principles and operational excellence.

🎯 Requirements

• 4+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role. • C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure. • AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each. • RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution. • Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK. • AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM. • SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager. • Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling. • Excellent communication skills and a collaborative, blameless engineering mindset. • Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.

🏖️ Benefits

• Performance-based annual bonus. • Medical, Dental, and Vision insurance. • 401(k) with company match. • Generous PTO plus paid company holidays. • Company-paid life and long-term disability insurance. • Paid parental leave.

Apply Now

Similar Jobs

🔥 1 hour ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer focusing on building and maintaining cloud infrastructure for Cisco Meraki. Analyzing reliability, troubleshooting, and implementing solutions in a secure environment.

Ansible

AWS

Azure

Cloud

Docker

Grafana

Kubernetes

Linux

Python

Ruby

Splunk

Unix

Go

🔥 3 hours ago

Valence

51 - 200

🤖 Artificial Intelligence

👥 HR Tech

☁️ SaaS

Senior DevOps Engineer managing AWS infrastructure for a pioneering AI coaching platform. Leading security initiatives and collaborating with multiple development teams on scalable solutions.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $50M Series B - Valence on 2025-09

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

AWS

Firewalls

PHP

Python

Ruby

Terraform

Go

🔥 3 hours ago

NationsBenefits

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Manager, Site Reliability Engineering leading a US-based SRE team at NationsBenefits, a healthcare fintech. Driving operational excellence and mentoring engineers for high-quality service delivery.

Docker

Java

Kubernetes

MySQL

NoSQL

Python

SDLC

SQL

🔥 3 hours ago

SS&C Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

📦 Logistics

Site Reliability Engineer for a leading financial services and healthcare technology company. Ensuring operational health, reliability, and availability of cloud platforms.

AWS

Cloud

Grafana

Kubernetes

Linux

Prometheus

Python

Splunk

Go

🔥 3 hours ago

SS&C Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

📦 Logistics

Site Reliability Engineer contributing to the operational health of FedRAMP High cloud platform at leading financial services company. This role covers advanced production support responsibilities and reliability engineering.

AWS

Cloud

Grafana

Kubernetes

Linux

Prometheus

Python

Splunk

Go