Senior / Lead Infrastructure & Operations Engineer

Austin Werner

Boston (MA)

On-site

USD 120,000 - 150,000

Full time

13 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Austin Werner is seeking a Senior / Lead Infrastructure & Operations Engineer for an AI startup. You’ll own the platform from commit to production, shaping CI/CD pipelines across GCP and specialist GPU providers.

You will design networks, manage multi-account cloud estate as code, and lead incidents with automation to reduce toil, building self‑service tooling and paved road environments.

This foundational role requires hands-on, security-conscious engineers who can operate across clouds.

Qualifications

  • 5+ years building and operating production infrastructure at engineering‑driven companies
  • Deep fluency with infrastructure as code (Terraform, Pulumi, or similar), CI/CD systems, Kubernetes, and at least one major cloud (GCP preferred; AWS acceptable)
  • Hands-on experience running multi-account cloud estates as code across providers

Responsibilities

  • Own end‑to‑end CI/CD pipelines, environments, and gated promotions
  • Manage cloud estate as code across GCP and specialist GPU providers, including IAM and provisioning
  • Design and operate underlying network: VPCs/subnets, interconnects, DNS, and cross‑cloud failover
  • Make onboarding repeatable with environments, docs‑as‑code, and sensible defaults
  • Lead incidents and reduce toil with self‑service tooling

Skills

Infrastructure as code
CI/CD
Kubernetes
Cloud platforms (GCP)
Multi-cloud networking

Tools

Terraform
Pulumi
GCP
AWS

Job description

Senior / Lead Infrastructure & Operations Engineer - (AI Startup)

We’re partnering with a well‑funded, stealth‑stage AI startup at the intersection of frontier AI research and fundamental physics. They’re building AI systems that can discover new physics at scale.

This is a foundational infrastructure role: you’ll own the platform every other engineering team ships through.

The Opportunity

Our client is looking for a senior infrastructure engineer to own the path from commit to production for a codebase where most commits are written by AI agents. Human review doesn’t scale at that velocity; CI does. What CI enforces is the architecture.

You’ll decide what CI enforces and build it, across a cloud estate spanning GCP and specialist GPU providers, on a network you design and a failover you rehearse. This is being built from a thin starting point, not inherited from a mature platform.

What You’ll Do as Senior / Lead Infrastructure & Operations Engineer
  • Own end‑to‑end CI/CD: pipelines, staged environments, promotion with real gates. The artifact that passes tests is the artifact that runs.
  • Run the cloud estate as code across GCP and specialist GPU providers: org structure, IAM, quotas, and provisioning so engineers don’t file tickets for basic needs. Repo and permission management counts as infrastructure here.
  • Design and operate the underlying network: VPCs/subnets across providers, interconnects, DNS, and cross‑cloud failover that is actually rehearsed.
  • Make onboarding a system: treat time‑to‑productivity as an infrastructure property. Build the “paved road” with environments, docs‑as‑code, and defaults that make the right thing the easy thing.
  • Lead incidents and reduce toil with code, preferring self‑service abstractions over tickets.
What You’ll Bring as Senior / Lead Infrastructure & Operations Engineer
  • 5+ years building and operating production infrastructure at companies known for engineering rigor (e.g., Stripe, Cloudflare, Datadog, Snowflake, Databricks, Google, Netflix, or comparable).
  • Deep fluency with infrastructure as code (Terraform, Pulumi, or similar), CI/CD systems, Kubernetes, and at least one major cloud (GCP preferred; AWS acceptable).
  • Experience building CI/CD from 0→1, not just maintaining a mature system. You can explain mechanically how you’d cut a build or registry bill; container build and registry economics are something you reason about before they show up on an invoice.
  • Hands‑on experience running a multi‑account cloud estate as code, ideally with multi‑cloud networking and a rehearsed failover.
  • A track record of leading incidents and systematically reducing toil with automation and self‑service platforms.
Nice to Have
  • Built CI/CD or release engineering from scratch at a fast‑growing company.
  • Strong FinOps instinct: cloud cost at account and architecture level (commitments, egress, idle spend, unit economics).
  • Experience with specialist GPU‑cloud providers (e.g., Modal, CoreWeave, or equivalents) and the account/quota/network realities of running across them alongside a major cloud.
  • Production observability with OpenTelemetry, Prometheus, Grafana, or similar.
  • Experience supporting machine‑generated or unusually high‑velocity commit patterns.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Psi • Boston (MA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Meaningful equity
Competitive compensation
Benefits
Senior DevOps Engineer
Senior DevOps Engineer

Clera • New York (NY)

On-site
USD 160,000 - 200,000
Equity participation
On-site role in New York
Startup growth environment
Director of Infrastructure Engineering
Director of Infrastructure Engineering

Appsierra Group • United States

On-site
USD 350,000 - 500,000
Equity compensation eligibility
Performance-based bonuses
Health insurance reimbursement up to 1
+3
Infrastructure Engineer — Seed-Stage AI Lab
Infrastructure Engineer — Seed-Stage AI Lab

Aionia • San Francisco (CA)

On-site
USD 185,000 - 235,000
Visa Sponsorship
Equity Options
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Colossus Technologies Group • Boston (MA)

Hybrid
USD 180,000 - 220,000
Health & wellness benefit
Competitive equity package
Hybrid work flexibility
Senior DevOps Engineer/AWS_Hybrid (NYC)
Senior DevOps Engineer/AWS_Hybrid (NYC)

PulseRise Technologies • New York (NY)

Hybrid
USD 160,000 - 200,000
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities
Senior / Lead / Principal Platform Engineer
Senior / Lead / Principal Platform Engineer

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Opportunities for career growth in a high-growth AI company
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000