Principal Infrastructure Engineer

🔥 13 hours ago

🇦🇷 Argentina – Remote

💵 $12.5k - $20.8k / month

⏰ Full Time

🔴 Lead

👷 Infrastructure Engineer

👻 Ghost score 21%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sezzle

Sezzle

201 - 500 employees

Founded 2016

💳 Fintech

👥 B2C

🛍️ eCommerce

Fintech • B2C • eCommerce

Sezzle is a financial technology company that offers a "buy now, pay later" service, allowing consumers to purchase products and pay for them in four interest-free installments over six weeks. The Sezzle app provides users with a flexible financing alternative to traditional credit cards, enabling instant approval decisions without impacting credit scores. Sezzle partners with various top brands, including Amazon, Walmart, and Target, to offer in-app and in-store payment options. The company's mission is to empower consumers financially by providing more financial freedom and control. It is available as a mobile app, with millions of downloads and high user ratings, and works towards accessibility and inclusion on its platform.

📋 Description

• Design, build, operate, and scale the platform powering Sezzle, a fintech company providing interest-free installment payments and shopping technology • Own complex infrastructure initiatives from architecture and prototyping through implementation, production rollout, and ongoing operation • Identify system limits and implement improvements for traffic, data volume, workload complexity, throughput, latency, resilience, and cost efficiency • Design resilient AWS account, IAM, network, and service architectures, including multi-AZ or multi-region capabilities • Build and operate the Kubernetes platform, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, capacity planning, failover, schema changes, and migrations • Define and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries • Participate in on-call rotation and lead technical recovery during serious incidents and outages • Design, test, and document disaster-recovery backup, restore, and failover mechanisms • Build infrastructure-as-code and operational automation for provisioning, configuration, deployments, upgrades, and recovery • Improve observability through metrics, logs, traces, dashboards, and actionable alerts • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths • Improve cloud cost efficiency through resource right-sizing, utilization improvements, autoscaling, and storage tuning • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and toil reduction • Write architecture proposals, evaluate technology tradeoffs through prototypes and benchmarks, review shared-infrastructure changes, and document system operation and failure • Work with application engineers, Security, Compliance, and engineering leadership to connect infrastructure decisions to customer and business outcomes

🎯 Requirements

• Bachelor's degree in Computer Science or a similar technical field (required) • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines • Deep production expertise with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity • Deep production expertise with Kubernetes; EKS experience strongly preferred • Deep expertise with RDS/Aurora MySQL and/or Postgres at scale • Track record of personally delivering infrastructure scaling improvements • Strong coding and automation skills using Golang, Python, or similar languages • Infrastructure-as-code experience with Terraform or equivalent • Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes • Experience operating a 24/7 high-availability platform • Willingness and ability to participate in an on-call rotation and handle production incidents • Experience implementing and testing disaster recovery against defined recovery objectives • Experience with observability, load testing, capacity planning, and safe CI/CD practices • Active use of AI tooling in engineering or operations • Ability to carry ambiguous technical problems through production delivery and collaborate across engineering disciplines • Preferred: fintech, payments, or banking experience; multi-region architectures; chaos engineering; Prometheus, Grafana, Loki, Tempo; internal platform and self-service tooling; AI-assisted incident investigation or operational automation

🏖️ Benefits

• Competitive gross monthly compensation of $12,500–$20,800 USD based on location and experience • Remote work in Argentina / Latin America • Open-source-focused engineering environment

Apply Now