Principal Infrastructure Engineer

🔥 20 hours ago

🇳🇱 Netherlands – Remote

💵 $12.5k - $20.8k / month

⏰ Full Time

🔴 Lead

👷 Infrastructure Engineer

👻 Ghost score 21%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sezzle

Sezzle

201 - 500 employees

Founded 2016

💳 Fintech

👥 B2C

🛍️ eCommerce

Fintech • B2C • eCommerce

Sezzle is a financial technology company that offers a "buy now, pay later" service, allowing consumers to purchase products and pay for them in four interest-free installments over six weeks. The Sezzle app provides users with a flexible financing alternative to traditional credit cards, enabling instant approval decisions without impacting credit scores. Sezzle partners with various top brands, including Amazon, Walmart, and Target, to offer in-app and in-store payment options. The company's mission is to empower consumers financially by providing more financial freedom and control. It is available as a mobile app, with millions of downloads and high user ratings, and works towards accessibility and inclusion on its platform.

📋 Description

• Own the technical architecture and evolution of Sezzle’s core infrastructure • Identify system limits, prioritize improvements, and support increasing traffic, data volume, and workload complexity • Connect business workflows to infrastructure improvements across applications, data, and systems • Build capacity models, run load and stress tests, diagnose bottlenecks, and validate throughput, latency, saturation, and cost improvements • Design and build resilient AWS account, IAM, network, and service architectures • Build and operate the Kubernetes platform, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability • Scale and optimize Aurora RDS for MySQL and Postgres, including queries, indexes, connections, replication, failover, schema changes, and migrations • Define and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries • Participate in on-call and lead technical recovery during serious incidents and outages • Design, test, and document disaster recovery, backup, restore, and failover mechanisms • Build infrastructure-as-code and operational automation for provisioning, configuration, deployment, upgrades, and recovery • Improve observability through metrics, logs, traces, dashboards, and actionable alerts • Deliver safe infrastructure migrations with phased rollouts, validation, and rollback paths • Improve cloud cost efficiency through resource right-sizing, utilization, autoscaling, and storage optimization • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and repetitive operations • Write architecture proposals, evaluate technology tradeoffs through prototypes and benchmarks, and document system behavior and failure modes • Report to engineering leadership and collaborate with application engineering, Security, and Compliance teams

🎯 Requirements

• Bachelor's degree in Computer Science or a similar technical field (required) • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting • Deep expertise with RDS/Aurora for MySQL and/or Postgres, including query performance, indexing, connection management, replication, high availability, failover, and backup/recovery • Track record of personally delivering infrastructure scaling improvements • Strong coding and automation skills using Golang, Python, or similar languages • Infrastructure-as-code experience with Terraform or equivalent • Strong fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes • Experience operating a 24/7 high-availability platform with hands-on incident response and postmortem remediation • Willingness to participate in an on-call rotation • Experience implementing and testing disaster recovery against defined recovery objectives • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices • Active use of AI tooling in engineering or operations • Ability to carry ambiguous technical problems through production delivery and collaborate across engineering disciplines • Preferred: EKS, fintech/payments/banking, multi-region architectures, chaos engineering, Prometheus, Grafana, Loki, Tempo, internal platform capabilities, self-service tooling, and AI-assisted operational automation

Apply Now