
201 - 500 employees
Founded 2016
💳 Fintech
👥 B2C
🛍️ eCommerce
Fintech • B2C • eCommerce
Sezzle is a financial technology company that offers a "buy now, pay later" service, allowing consumers to purchase products and pay for them in four interest-free installments over six weeks. The Sezzle app provides users with a flexible financing alternative to traditional credit cards, enabling instant approval decisions without impacting credit scores. Sezzle partners with various top brands, including Amazon, Walmart, and Target, to offer in-app and in-store payment options. The company's mission is to empower consumers financially by providing more financial freedom and control. It is available as a mobile app, with millions of downloads and high user ratings, and works towards accessibility and inclusion on its platform.
🔥 0 minutes ago
🇺🇸 United States – Remote
💵 $400k - $600k / year
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
Founded 2016
💳 Fintech
👥 B2C
🛍️ eCommerce
Fintech • B2C • eCommerce
Sezzle is a financial technology company that offers a "buy now, pay later" service, allowing consumers to purchase products and pay for them in four interest-free installments over six weeks. The Sezzle app provides users with a flexible financing alternative to traditional credit cards, enabling instant approval decisions without impacting credit scores. Sezzle partners with various top brands, including Amazon, Walmart, and Target, to offer in-app and in-store payment options. The company's mission is to empower consumers financially by providing more financial freedom and control. It is available as a mobile app, with millions of downloads and high user ratings, and works towards accessibility and inclusion on its platform.
• Own the infrastructure vision, strategy, and multi-year roadmap: scale today's high-growth fintech platform while continually strengthening its resilience, controls, and audit-readiness. • Lead, grow, and mentor the infrastructure, platform, and SRE organization, including hiring, career development, on-call health, and building a culture of operational excellence and blameless learning. • Own reliability end-to-end: define and enforce SLOs and error budgets, mature incident management and postmortem practices, and be accountable for platform availability across the business. • Manage, and participate in, the on-call rotation, and serve as senior incident commander for high-severity events: leading recovery from major degradations and full outages through rapid, evidence-based triage, decisive action under uncertainty, and clear communication to stakeholders throughout. • Direct our AWS strategy, including account architecture, IAM and network design, multi-AZ/multi-region posture, service selection, and cost management (FinOps). You will own and defend the cloud budget. • Own the Kubernetes platform as a product: cluster architecture, upgrade strategy, workload isolation, autoscaling, progressive delivery, and the developer experience of every team that ships on it. • Own the database tier, centered on Aurora RDS (MySQL and Postgres): availability, performance, capacity, schema and migration safety practices, backup/restore verification, and encryption. • Design, implement, and continuously test disaster recovery and business continuity: defined RTO/RPO targets per system tier, regular game days and failover exercises, and DR evidence that stands up to auditor scrutiny. • Champion the AI-boosted SRE transformation: evaluate and deploy AI tooling and agents for incident triage, observability, runbook automation, and toil reduction; set standards for safe, auditable use of AI in production operations; and bring the team along through training and example. • Partner with Security and Compliance to own infrastructure's role in PCI-DSS and SOC 2: control design and operation, evidence collection, segmentation, vulnerability and patch management, and audit support, with the maturity to meet the expectations of banking partners and financial-industry examinations. • Drive infrastructure-as-code and platform automation as the default: everything reproducible, reviewed, and recoverable; nothing artisanal. • Own vendor and technology strategy for the infrastructure domain: build-vs-buy decisions, vendor risk management, contract negotiation, and third-party resilience. • Communicate crisply with executives, the board, and auditors, translating infrastructure risk, investment, and posture into business terms.
• 15+ years of combined experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial depth in infrastructure, including 5+ years leading engineering teams. • Deep, hands-on expertise with AWS: you have designed and operated production architectures across compute, networking (VPC design, Transit Gateway, PrivateLink), IAM, and multi-account organizations at scale. • Deep, hands-on expertise with Kubernetes in production: cluster lifecycle management, workload architecture, scaling, and the operational realities of running business-critical services on it (EKS experience strongly preferred). • Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): high availability, replication, failover, performance tuning, and backup/recovery you have personally verified under pressure. • Proven ownership of disaster recovery and business continuity for a production platform: you have defined RTO/RPO targets, built the capability to meet them, and run real failover tests, not just written the document. • Demonstrated AI-forward leadership: you actively use AI tooling in engineering or operations work today, have opinions grounded in practice about where it helps and where it doesn't, and have led (or are visibly leading) a team's adoption of AI-assisted workflows. • Track record of operating a 24/7, high-availability platform where downtime has direct revenue or customer impact, including mature incident command and postmortem practices. • Willingness to manage and participate in an on-call rotation, and demonstrated ability to lead recovery from a full production outage: forming and testing hypotheses from logs, metrics, and traces rather than guesswork, making the right call quickly with incomplete information, and knowing when to mitigate first and root-cause later. • Still technical, by choice: you remain a credible hands-on engineer, comfortable in a terminal, reading dashboards, and reviewing designs, and you expect to stay that way. You will lead the team and work alongside it; this is not a delegation-only role. • Experience owning significant cloud budgets and driving cost efficiency without sacrificing reliability. • Strong grounding in infrastructure-as-code (Terraform or equivalent) and modern CI/CD practices. • Demonstrated ability to hire, develop, and retain strong infrastructure and SRE talent, and to hold a high bar through growth. • Bachelor's degree in Computer Science or a similar technical field (required).
• Unlimited PTO, volunteer hours and sabbatical • Life, STD/LTD, medical, dental and vision insurance • Highly discounted LifeTime gym membership • 401k with match • Collaborative fun co-workers • The opportunity to join the fastest growing FinTech alongside a team of motivated and driven individuals
Apply Now🔥 4 minutes ago
Staff Site Reliability Engineer leading global platform reliability and observability strategy at Calix, specializing in GCP and Kubernetes infrastructure.
🇺🇸 United States – Remote
💵 $136k - $231k / year
💰 $50M Venture Round on 2009-08
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Apache
BigQuery
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Kafka
Kubernetes
Postgres
Prometheus
Python
Terraform
Go
🔥 1 hour ago
Staff Site Reliability Engineer at Datavant focuses on cloud infrastructure design and security. Collaborating with teams to enhance reliability, scalability, and operability in healthcare data solutions.
🇺🇸 United States – Remote
💵 $190k - $235k / year
💰 $40M Series B on 2020-10
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Ansible
AWS
Azure
Cloud
DNS
Terraform
🔥 2 hours ago
Staff Network Reliability Engineer managing RAN operations for satellite connectivity at Skylo. Overseeing RAN health and incident resolution in a production NTN environment.
🇺🇸 United States – Remote
💵 $150k - $162k / year
💰 $30M Venture Round - Skylo on 2025-02
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
Grafana
IoT
Kubernetes
Node.js
Prometheus
TypeScript
🔥 4 hours ago
Principal Consulting Engineer optimizing private cloud infrastructure built on VCF 9.0 for CDW. Driving standardization and automation across compute, storage, and networking management layers.
🇺🇸 United States – Remote
💵 $164k - $240.8k / year
💰 Post-IPO Equity on 2015-07
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Cloud
🔥 6 hours ago
Site Reliability Engineer automating and improving reliability at Collins Aerospace. Collaborating with teams to tackle complex technical problems in flight operations.
🇺🇸 United States – Remote
💵 $107.5k - $204.5k / year
💰 $200k Grant - RTX on 2024-11
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
Ansible
Docker
Kubernetes
Linux
SaltStack
Terraform
Unix