Principal Infrastructure Engineer

United States Digital Space LLC

United States

Remote

USD 140,000 - 232,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a Principal Infrastructure Engineer to design, build, operate, and scale its cloud platform. You will own complex initiatives spanning architecture, automation, and cost efficiency in a high-traffic AWS/Kubernetes environment.

You’ll work across engineering teams to improve reliability, performance, and resilience, while leading incident response and disaster recovery efforts.

Qualifications

  • Own the technical architecture and evolution of core infrastructure.
  • Deep expertise with AWS, Kubernetes, and Aurora RDS (MySQL and Postgres).
  • Experience writing code and infrastructure-as-code for production systems.

Responsibilities

  • Own the technical architecture and evolution of core infrastructure: identify system limits, prioritize improvements, implement changes.
  • Connect business understanding to improvements across the system: learn production workflows and partner across teams.
  • Engineer for scale and performance: build capacity models, run load tests, diagnose bottlenecks, validate improvements.
  • Design and build AWS infrastructure: implement resilient account, IAM, network, services; handle quotas and multi-AZ/region.
  • Build and operate the Kubernetes platform: improve cluster architecture, automation, autoscaling, deployments.
  • Scale and optimize Aurora RDS for MySQL and Postgres: tune queries, address bottlenecks, plan capacity, implement migrations.
  • Improve reliability through engineering: define SLOs, implement failure isolation, backpressure, safe retries.
  • Participate in on-call and drive technical recovery during incidents: triage, mitigations, postmortems.
  • Implement and test disaster recovery: backup, restore, failover, recovery objectives; run drills.
  • Build infrastructure-as-code and automation: reproducible provisioning, deployments, upgrades, recovery.
  • Build observability: metrics, logs, traces, dashboards, alerts; visibility into health and impact.
  • Deliver safe infrastructure migrations: phased rollouts, validation, rollback paths.
  • Improve cloud cost efficiency through changes: right-size resources, autoscale tuning, cost metrics.
  • Build and evaluate AI-assisted operational tooling: bounded permissions, measurable toil reductions.
  • Make technical decisions clear and executable: architecture proposals, benchmarks, documentation.

Skills

AWS
Kubernetes
Aurora RDS
IaC
Observability

Tools

Terraform
CloudFormation

Job description

About the company:

With a mission to financially empower the next generation, the company is revolutionizing the shopping experience beyond payments, blending cutting-edge tech with seamless, interest-free installment plans that make shopping smarter and more accessible. We’re not just transforming payments; we’re redefining how people discover, interact with, and purchase the things they love while driving real impact on merchant sales through increased conversions and higher order values. As we continue to shape the future of fintech and retail, we’re building an innovative, dynamic team passionate about creating more than just a transaction but a truly unique shopping journey. If you’re excited about pushing boundaries in tech and delivering a game-changing experience for consumers and merchants alike, come join us at the company and help create the future of shopping!

Compensation:

For this principal development role, with 12+ years of experience, the compensation range is $12,500 - $20,800 USD based on location and experience level per month and in gross amount. This range acknowledges the extensive expertise, leadership capabilities, and significant contributions expected at this level, offering a competitive salary to reflect the value of advanced skills and experience

About the Role:

We are seeking an exceptional Principal Infrastructure Engineer to design, build, operate, and scale the platform that powers the company. Your focus will be the hardest infrastructure problems: increasing throughput, reducing latency, removing capacity bottlenecks, strengthening resilience, and making production operations more automated and predictable as the business grows.

You will own complex technical initiatives from architecture and prototyping through implementation, production rollout, and ongoing operation. Your impact will come from the systems you build, the problems you solve, and measurable improvements in reliability, performance, and cost efficiency.

Our stack runs on AWS, with workloads orchestrated on Kubernetes and data anchored in Aurora RDS (MySQL and Postgres). You should know these technologies deeply and be comfortable moving between cloud architecture, networking, cluster internals, database performance, and application behavior to understand how the entire system scales. You will write code and infrastructure-as-code, debug production systems, and deliver changes that hold up under real traffic and failure conditions.

Operational ownership is part of the job. You will participate in the on-call rotation and take a hands-on role in recovering from major incidents, including full outages. We need someone who can form and test hypotheses using logs, metrics, and traces, make sound mitigation decisions with incomplete information, and turn incident findings into lasting engineering fixes.

You will also build and apply AI-assisted infrastructure and SRE tooling for incident investigation, capacity analysis, runbook automation, and toil reduction. You will evaluate these tools through practical results and apply appropriate access controls, validation, and auditability to their use in production.

This role reports to engineering leadership and works closely with application engineers, Security, and Compliance. You will develop a deep understanding of how the company's business operates and how customer journeys, transaction patterns, and product decisions shape infrastructure demand and behavior. You will use that understanding to identify and deliver improvements across teams and technical domains, connecting infrastructure decisions to better customer outcomes and business performance.

What you’ll do:
  • Own the technical architecture and evolution of core infrastructure: identify system limits, prioritize technical improvements, and implement changes that support increasing traffic, data volume, and workload complexity.
  • Connect business understanding to improvements across the system: learn how key business workflows behave in production, trace their impact across applications, data, and infrastructure, and partner across teams to improve performance, reliability, and cost efficiency beyond any single service or team's scope.
  • Engineer for scale and performance: build capacity models, run load and stress tests, diagnose bottlenecks across compute, networking, Kubernetes, and databases, and validate improvements against throughput, latency, saturation, and cost per workload.
  • Design and build AWS infrastructure: implement resilient account, IAM, network, and service architectures; address service quotas, fault isolation, and multi-AZ or multi-region requirements as workloads grow.
  • Build and operate the Kubernetes platform: improve cluster architecture, lifecycle automation, workload isolation, resource allocation, autoscaling, safe upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres: tune queries and indexes, address connection and replication bottlenecks, plan capacity, improve failover behavior, and implement safe schema changes and database migrations with application engineers.
  • Improve reliability through engineering: define and instrument service-level objectives and error budgets with service owners; implement failure isolation, backpressure, load shedding, and safe retry behavior where needed to prevent cascading failures.
  • Participate in on-call and drive technical recovery during serious incidents: use evidence-based triage, execute mitigations, communicate findings, and implement corrective actions from postmortems.
  • Implement and test disaster recovery: design backup, restore, and failover mechanisms against agreed recovery time and recovery point objectives; run recovery exercises and document measured results.
  • Build infrastructure-as-code and operational automation: make provisioning, configuration, deployments, upgrades, and recovery reproducible, reviewed, testable, and recoverable. Eliminate recurring manual work through code.
  • Build observability that makes production diagnosable: improve metrics, logs, traces, dashboards, and actionable alerts, with visibility into service health, scaling limits, and customer impact.
  • Deliver safe infrastructure migrations: design phased rollouts, compatibility checks, validation, and rollback paths for changes to shared production systems.
  • Improve cloud cost efficiency through technical changes: right-size resources, improve utilization, tune autoscaling and storage, and quantify savings while maintaining reliability and performance targets.
  • Build and evaluate AI-assisted operational tooling: apply AI to investigation, runbooks, anomaly analysis, and repetitive operations, with bounded permissions, reviewable actions, and measurable improvements in accuracy or toil.
  • Make technical decisions clear and executable: write architecture proposals, evaluate technology tradeoffs through prototypes and benchmarks, review changes affecting shared infrastructure, and document how systems operate and fail.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer/AWS_Hybrid (NYC)
Senior DevOps Engineer/AWS_Hybrid (NYC)

PulseRise Technologies • New York (NY)

On-site
USD 160,000 - 200,000
Infrastructure Engineer
Infrastructure Engineer

Eden Prescott • San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity
Health benefits
Significant equity and benefits
Principal Software Engineer-Infrastructure
Principal Software Engineer-Infrastructure

NV Energy • Portland (OR)

On-site
USD 150,000 - 190,000
Principal / Staff / Senior Infrastructure Engineer
Principal / Staff / Senior Infrastructure Engineer

Allspice, Inc. • Boston (MA), Northern (KY)

On-site
USD 170,000 - 230,000
Flexible work
Health benefits
Generous PTO
+3
Staff Engineer - Distributed Systems
Staff Engineer - Distributed Systems

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Principal and Senior Principal Engineer, AWS Compute and ML Services
Principal and Senior Principal Engineer, AWS Compute and ML Services

Amazon Inc. • New York (NY)

On-site
USD 220,000 - 298,000
Senior Platform Engineer
Senior Platform Engineer

Epsilon ASI • Denver (CO), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Level AI • Auckland (CA)

On-site
USD 140,000 - 210,000
Infrastructure/Platform Engineer
Infrastructure/Platform Engineer

Stealth AI Startup • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Principal Cloud Engineer
Principal Cloud Engineer

NextGenEnergyJobs • San Jose (CA), Northern (KY)

On-site
USD 180,000 - 280,000