Principal Infrastructure Engineer

Sezzle

Paris

Sur place

EUR 150 000 - 190 000

Plein temps

Il y a 11 jours
Générateur de candidature

A complete application in a minute — tailored resume and cover letter, ready to send.

Passez les filtres ATS

Avantages offerts par ce poste

Remote-friendly company
Generous parental & family leave
Equity ownership

Résumé du poste

Sezzle is seeking a Principal Infrastructure Engineer to design, build, operate, and scale the platform that powers Sezzle. You will own complex initiatives from architecture through production, driving reliability, performance, and cost efficiency at scale.

The role requires deep AWS and Kubernetes expertise, Aurora databases, and hands-on leadership to reduce toil and improve automation. You will participate in on-call rotations, implement disaster recovery, and deliver measurable improvements

Qualifications

  • Bachelor's degree in Computer Science or equivalent required.
  • 12+ years of experience in infrastructure/engineering roles.
  • Deep expertise with AWS, Kubernetes, and relational databases.

Responsabilités

  • Own technical initiatives from architecture to production rollout.
  • Design, build, operate, and scale Sezzle's platform infrastructure.
  • Lead disaster recovery, on-call, and incident response.
  • Develop AI-assisted tooling for incident investigation and automation.

Connaissances

Golang
Python
Kubernetes
AWS
Observability

Formation

Bachelor's degree in Computer Science

Outils

Terraform
Aurora (MySQL/Postgres)
CI/CD (GitLab)

Description du poste

  • We are seeking an exceptional Principal Infrastructure Engineer to design, build, operate, and scale the platform that powers Sezzle
  • Your focus will be the hardest infrastructure problems: increasing throughput, reducing latency, removing capacity bottlenecks, strengthening resilience, and making production operations more automated and predictable as the business grows
  • You will own complex technical initiatives from architecture and prototyping through implementation, production rollout, and ongoing operation
  • Your impact will come from the systems you build, the problems you solve, and measurable improvements in reliability, performance, and cost efficiency
  • Our stack runs on AWS, with workloads orchestrated on Kubernetes and data anchored in Aurora RDS (MySQL and Postgres)
  • You should know these technologies deeply and be comfortable moving between cloud architecture, networking, cluster internals, database performance, and application behavior to understand how the entire system scales
  • You will write code and infrastructure-as-code, debug production systems, and deliver changes that hold up under real traffic and failure conditions
  • Operational ownership is part of the job. You will participate in the on-call rotation and take a hands-on role in recovering from major incidents, including full outages
  • . We need someone who can form and test hypotheses using logs, metrics, and traces, make sound mitigation decisions with incomplete information, and turn incident findings into lasting engineering fixes
  • You will also build and apply AI-assisted infrastructure and SRE tooling for incident investigation, capacity analysis, runbook automation, and toil reduction
  • You will evaluate these tools through practical results and apply appropriate access controls, validation, and auditability to their use in production
  • This role reports to engineering leadership and works closely with application engineers, Security, and Compliance
  • You will develop a deep understanding of how Sezzle’s business operates and how customer journeys, transaction patterns, and product decisions shape infrastructure demand and behavior
  • You will use that understanding to identify and deliver improvements across teams and technical domains, connecting infrastructure decisions to better customer outcomes and business performance
  • Own the technical architecture and evolution of core infrastructure: identify system limits, prioritize technical improvements, and implement changes that support increasing traffic, data volume, and workload complexity
  • Connect business understanding to improvements across the system: learn how key business workflows behave in production, trace their impact across applications, data, and infrastructure, and partner across teams to improve performance, reliability, and cost efficiency beyond any single service or team’s scope
  • Engineer for scale and performance: build capacity models, run load and stress tests, diagnose bottlenecks across compute, networking, Kubernetes, and databases, and validate improvements against throughput, latency, saturation, and cost per workload
  • Design and build AWS infrastructure: implement resilient account, IAM, network, and service architectures; address service quotas, fault isolation, and multi-AZ or multi-region requirements as workloads grow
  • Build and operate the Kubernetes platform: improve cluster architecture, lifecycle automation, workload isolation, resource allocation, autoscaling, safe upgrades, and deployment reliability
  • Scale and optimize Aurora RDS for MySQL and Postgres: tune queries and indexes, address connection and replication bottlenecks, plan capacity, improve failover behavior, and implement safe schema changes and database migrations with application engineers
  • Improve reliability through engineering: define and instrument service-level objectives and error budgets with service owners; implement failure isolation, backpressure, load shedding, and safe retry behavior where needed to prevent cascading failures
  • Participate in on-call and drive technical recovery during serious incidents: use evidence-based triage, execute mitigations, communicate findings, and implement corrective actions from postmortems
  • Implement and test disaster recovery: design backup, restore, and failover mechanisms against agreed recovery time and recovery point objectives; run recovery exercises and document measured results
  • Build infrastructure-as-code and operational automation: make provisioning, configuration, deployments, upgrades, and recovery reproducible, reviewed, testable, and recoverable. Eliminate recurring manual work through code
  • Build observability that makes production diagnosable: improve metrics, logs, traces, dashboards, and actionable alerts, with visibility into service health, scaling limits, and customer impact
  • Deliver safe infrastructure migrations: design phased rollouts, compatibility checks, validation, and rollback paths for changes to shared production systems
  • Improve cloud cost efficiency through technical changes: right-size resources, improve utilization, tune autoscaling and storage, and quantify savings while maintaining reliability and performance targets
  • Build and evaluate AI-assisted operational tooling: apply AI to investigation, runbooks, anomaly analysis, and repetitive operations, with bounded permissions, reviewable actions, and measurable improvements in accuracy or toil
  • Make technical decisions clear and executable: write architecture proposals, evaluate technology tradeoffs through prototypes and benchmarks, review changes affecting shared infrastructure, and document how systems operate and fail
  • Sezzle’s Technology Stack:
  • Languages: Golang, Python
  • Database: MySQL, Postgres
  • DevOps and Cloud: AWS, Kubernetes
  • Version Control: Git
  • CI/CD: Gitlab
  • Open Source: Sezzle is focused on using open source, and we build what we can before buying!
Benefits
  • Comprehensive Benefit Plans
  • Generous Parental & Family Leave
  • Competitive 401k Match
  • Paid Time Off & Volunteer Time Off
  • Ownership Through Equity
  • 100% of Donations to Charity Matched
  • Remote Friendly Company
  • Highly Discounted Fitness Membership

Willingness to participate in an on-call rotation and demonstrated ability to recover production systems under pressure using evidence-based triage, safe mitigation, and clear technical communicationExperience implementing and testing disaster recovery against defined recovery objectives, including restoring data and validating service recoveryYou stay accountable for the outcome: you verify that changes work in production and that fixes remain effective as the platform scalesStrong systems fundamentals: Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes, with the ability to debug problems across infrastructure and application boundariesYou’re not bound by convention - your success—and much of the fun—lies in developing new ways to do thingsYou earn trust - you listen attentively, speak candidly, and treat others respectfullyStrong coding and automation skills, using Golang, Python, or similar languages to build production tooling and eliminate operational toil, alongside infrastructure-as-code experience with Terraform or equivalentDeep expertise with Kubernetes in production: cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting business-critical workloads. EKS experience is strongly preferredBachelor’s degree in Computer Science or a similar technical field (required)For this principal development role, with 12+ years of experienceYou build and ship: you turn architecture into working code, tested infrastructure, and safe production changesExperience operating a 24/ 7, high-availability platform where downtime has direct customer or revenue impact, including hands-on incident response and postmortem remediationYou earn trust: you listen carefully, communicate clearly, challenge technical decisions respectfully, and follow through on commitmentsPractical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial hands-on depth designing and operating production infrastructure at scaleYou measure results: you can explain the impact of your work in availability, latency, capacity, recovery time, cost, or hours of toil removedYou go deep: you investigate how systems behave under load and failure, follow the evidence, and fix underlying causesYou have relentlessly high standards - many people may think your standards are unreasonably high. You are continually raising the bar and driving those around you to deliver great results. You make sure that defects do not get sent down the line and that problems are fixed so they stay fixedDeep expertise with AWS: production experience across compute, IAM, multi-account architectures, and networking, including VPC design and private connectivityA track record of personally delivering infrastructure scaling improvements: identifying constraints, measuring baseline behavior, implementing changes, and demonstrating gains in capacity, latency, reliability, or cost efficiencyAbility to carry ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines to resolve system-wide constraintsYou have backbone; disagree, then commit - you can respectfully challenge decisions when you disagree, even when doing so is uncomfortable or exhausting. You have conviction and are tenacious. You do not compromise for the sake of social cohesion. Once a decision is determined, you commit whollyYou need action - speed matters in business. Many decisions and actions are reversible and do not need extensive study. We value calculated risk-takingActive use of AI tooling in engineering or operations, with practical judgment about its limitations and how to verify generated code, recommendations, and operational actionsYou deliver results - you focus on the key inputs and deliver them with the right quality and in a timely fashion. Despite setbacks, you rise to the occasion and never settleDeep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): query performance, indexing, connection management, replication, high availability, failover, and verified backup and recoveryYou value simplicity: you choose systems that are understandable, operable, and appropriate to the problem, and reduce unnecessary complexityExperience in fintech, payments, or banking, operating infrastructure with demanding reliability, security, and audit requirementsExperience with multi-region architectures, chaos engineering, and failure testing, including the consistency and recovery tradeoffs of distributed data systemsExperience building internal platform capabilities and self-service tooling, including deployment automation, progressive delivery, and reusable infrastructure componentsExperience building AI-assisted incident investigation or operational automation with restricted access, auditable execution, and clear human review pointsProficiency with Prometheus, Grafana, Loki, Tempo, or comparable observability systems

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Principal Infrastructure Engineer
Principal Infrastructure Engineer

Sezzle • France

Hybride
EUR 121 000 - 201 000
Principal Infrastructure Engineer
Principal Infrastructure Engineer

Sezzle • Paris

Sur place
EUR 121 000 - 202 000
DevOps / Infra Lead
DevOps / Infra Lead

Payflows • Paris

Sur place
EUR 90 000 - 120 000
Sr. Devops Engineer II
Sr. Devops Engineer II

DoubleVerify • Paris

Sur place
EUR 90 000 - 125 000
Principal Infra Engineer: Scale, Reliability, AI Ops
Principal Infra Engineer: Scale, Reliability, AI Ops

Sezzle • France

Hybride
EUR 121 000 - 201 000
Lead Software Engineer - Solution & Data Architecture - B2B SaaS Fintech
Lead Software Engineer - Solution & Data Architecture - B2B SaaS Fintech

Landytech • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Private medical insurance
Pension plan with NEST
+1
Principal Infra Engineer — Scale, AWS & Kubernetes
Principal Infra Engineer — Scale, AWS & Kubernetes

Sezzle • Paris

Sur place
EUR 121 000 - 202 000
Incident Operations Lead
Incident Operations Lead

Jobgether • France

Sur place
EUR 120 000 - 190 000
Stock options
Health benefits
One-time USD 500 home-office setup
+1
Cloud-Scale Infrastructure Architect
Cloud-Scale Infrastructure Architect

Sezzle • Paris

Sur place
EUR 150 000 - 190 000
Remote-friendly company
Generous parental & family leave
Equity ownership
Senior Software Engineer - Clearing
Senior Software Engineer - Clearing

Jobgether SRL • France

À distance
EUR 90 000 - 150 000
Competitive salary
Stock options
Home-office setup allowance (USD 500)
+2