Senior / Lead Infrastructure & Operations Engineer - (AI Startup)
We’re partnering with a well‑funded, stealth‑stage AI startup at the intersection of frontier AI research and fundamental physics. They’re building AI systems that can discover new physics at scale.
This is a foundational infrastructure role: you’ll own the platform every other engineering team ships through.
The Opportunity
Our client is looking for a senior infrastructure engineer to own the path from commit to production for a codebase where most commits are written by AI agents. Human review doesn’t scale at that velocity; CI does. What CI enforces is the architecture.
You’ll decide what CI enforces and build it, across a cloud estate spanning GCP and specialist GPU providers, on a network you design and a failover you rehearse. This is being built from a thin starting point, not inherited from a mature platform.
What You’ll Do as Senior / Lead Infrastructure & Operations Engineer
- Own end‑to‑end CI/CD: pipelines, staged environments, promotion with real gates. The artifact that passes tests is the artifact that runs.
- Run the cloud estate as code across GCP and specialist GPU providers: org structure, IAM, quotas, and provisioning so engineers don’t file tickets for basic needs. Repo and permission management counts as infrastructure here.
- Design and operate the underlying network: VPCs/subnets across providers, interconnects, DNS, and cross‑cloud failover that is actually rehearsed.
- Make onboarding a system: treat time‑to‑productivity as an infrastructure property. Build the “paved road” with environments, docs‑as‑code, and defaults that make the right thing the easy thing.
- Lead incidents and reduce toil with code, preferring self‑service abstractions over tickets.
What You’ll Bring as Senior / Lead Infrastructure & Operations Engineer
- 5+ years building and operating production infrastructure at companies known for engineering rigor (e.g., Stripe, Cloudflare, Datadog, Snowflake, Databricks, Google, Netflix, or comparable).
- Deep fluency with infrastructure as code (Terraform, Pulumi, or similar), CI/CD systems, Kubernetes, and at least one major cloud (GCP preferred; AWS acceptable).
- Experience building CI/CD from 0→1, not just maintaining a mature system. You can explain mechanically how you’d cut a build or registry bill; container build and registry economics are something you reason about before they show up on an invoice.
- Hands‑on experience running a multi‑account cloud estate as code, ideally with multi‑cloud networking and a rehearsed failover.
- A track record of leading incidents and systematically reducing toil with automation and self‑service platforms.
Nice to Have
- Built CI/CD or release engineering from scratch at a fast‑growing company.
- Strong FinOps instinct: cloud cost at account and architecture level (commitments, egress, idle spend, unit economics).
- Experience with specialist GPU‑cloud providers (e.g., Modal, CoreWeave, or equivalents) and the account/quota/network realities of running across them alongside a major cloud.
- Production observability with OpenTelemetry, Prometheus, Grafana, or similar.
- Experience supporting machine‑generated or unusually high‑velocity commit patterns.