An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Docker is seeking a Staff Infrastructure Engineer for the Billing Platform to ensure reliable, scalable operations. You will design AI-agent-assisted deployment workflows, implement IaC patterns in Terraform on AWS, and improve billing accuracy through observability and incident response.
You'll collaborate with software engineers to translate constraints into robust infrastructure, mentor teammates, and sometimes participate in on-call rotations outside business hours as needed.
DockerDocker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.
_______________________________________________________________________
We’re building AI-native development practices into how this team works at a foundational level. That means infrastructure design needs to account for a new kind of collaborator: AI agents that generate, deploy, and operate software. The Staff Infrastructure Engineer on this team will keep systems running, define what safe, observable, AI-assisted infrastructure operations look like in practice, and set the standard for how the broader engineering organization follows.
The Billing Platform Engineering team owns the systems that make Docker's commercial model real. You'll work on problems like:
How do we design infrastructure that makes AI-generated deployments safe to ship and easy to roll back?
How do we instrument billing systems so that failures — billing miscalculations, entitlement gaps, payment errors — are detected immediately and unambiguously?
How do we build infrastructure that scales with usage-based billing workloads without manual intervention?
How do we make the developer experience on this team faster and more reliable — local environments, CI/CD pipelines, deployment tooling?