A complete application in a minute — tailored resume and cover letter, ready to send.
Sezzle is seeking a Principal Infrastructure Engineer to design, build, operate, and scale the platform that powers Sezzle. You will own complex initiatives from architecture through production, driving reliability, performance, and cost efficiency at scale.
The role requires deep AWS and Kubernetes expertise, Aurora databases, and hands-on leadership to reduce toil and improve automation. You will participate in on-call rotations, implement disaster recovery, and deliver measurable improvements
Willingness to participate in an on-call rotation and demonstrated ability to recover production systems under pressure using evidence-based triage, safe mitigation, and clear technical communicationExperience implementing and testing disaster recovery against defined recovery objectives, including restoring data and validating service recoveryYou stay accountable for the outcome: you verify that changes work in production and that fixes remain effective as the platform scalesStrong systems fundamentals: Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes, with the ability to debug problems across infrastructure and application boundariesYou’re not bound by convention - your success—and much of the fun—lies in developing new ways to do thingsYou earn trust - you listen attentively, speak candidly, and treat others respectfullyStrong coding and automation skills, using Golang, Python, or similar languages to build production tooling and eliminate operational toil, alongside infrastructure-as-code experience with Terraform or equivalentDeep expertise with Kubernetes in production: cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting business-critical workloads. EKS experience is strongly preferredBachelor’s degree in Computer Science or a similar technical field (required)For this principal development role, with 12+ years of experienceYou build and ship: you turn architecture into working code, tested infrastructure, and safe production changesExperience operating a 24/ 7, high-availability platform where downtime has direct customer or revenue impact, including hands-on incident response and postmortem remediationYou earn trust: you listen carefully, communicate clearly, challenge technical decisions respectfully, and follow through on commitmentsPractical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial hands-on depth designing and operating production infrastructure at scaleYou measure results: you can explain the impact of your work in availability, latency, capacity, recovery time, cost, or hours of toil removedYou go deep: you investigate how systems behave under load and failure, follow the evidence, and fix underlying causesYou have relentlessly high standards - many people may think your standards are unreasonably high. You are continually raising the bar and driving those around you to deliver great results. You make sure that defects do not get sent down the line and that problems are fixed so they stay fixedDeep expertise with AWS: production experience across compute, IAM, multi-account architectures, and networking, including VPC design and private connectivityA track record of personally delivering infrastructure scaling improvements: identifying constraints, measuring baseline behavior, implementing changes, and demonstrating gains in capacity, latency, reliability, or cost efficiencyAbility to carry ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines to resolve system-wide constraintsYou have backbone; disagree, then commit - you can respectfully challenge decisions when you disagree, even when doing so is uncomfortable or exhausting. You have conviction and are tenacious. You do not compromise for the sake of social cohesion. Once a decision is determined, you commit whollyYou need action - speed matters in business. Many decisions and actions are reversible and do not need extensive study. We value calculated risk-takingActive use of AI tooling in engineering or operations, with practical judgment about its limitations and how to verify generated code, recommendations, and operational actionsYou deliver results - you focus on the key inputs and deliver them with the right quality and in a timely fashion. Despite setbacks, you rise to the occasion and never settleDeep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): query performance, indexing, connection management, replication, high availability, failover, and verified backup and recoveryYou value simplicity: you choose systems that are understandable, operable, and appropriate to the problem, and reduce unnecessary complexityExperience in fintech, payments, or banking, operating infrastructure with demanding reliability, security, and audit requirementsExperience with multi-region architectures, chaos engineering, and failure testing, including the consistency and recovery tradeoffs of distributed data systemsExperience building internal platform capabilities and self-service tooling, including deployment automation, progressive delivery, and reusable infrastructure componentsExperience building AI-assisted incident investigation or operational automation with restricted access, auditable execution, and clear human review pointsProficiency with Prometheus, Grafana, Loki, Tempo, or comparable observability systems