Principal Platform Engineer — Kubernetes & Cloud Infrastructure

Ombud

Denver (CO)

On-site

USD 130,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dormont Manufacturing Co seeks a Principal Platform Engineer to lead the technical direction of their cloud infrastructure in Denver. This senior individual-contributor role requires 8+ years of experience in platform engineering, with expertise in production Kubernetes and AWS.

The role includes shaping the company’s infrastructure strategy, managing costs, and ensuring reliability. Successful candidates will possess strong communication skills and a proven track record of impactful architectural decisions.

Qualifications

  • 8+ years of platform, infrastructure, SRE, or DevOps experience.
  • At least 3+ years operating production Kubernetes at scale.
  • Track record of improvements in reliability, cost, or developer velocity.

Responsibilities

  • Own the technical direction for the cloud infrastructure.
  • Manage production Kubernetes clusters and AWS infrastructure.
  • Lead architecture for self-service infrastructure roadmap.

Skills

Platform, infrastructure, SRE, or DevOps experience
Production Kubernetes expertise
Deep AWS expertise
Production fluency with Terraform
Strong written communication

Tools

Terraform
Docker
Linux
CI/CD systems

Job description

  • Location: Denver, CO (hybrid — Tue/Wed/Thu in office)
  • Reports to: CEO
The role

Ombud’s platform runs production AI workloads for enterprise customers, and we’re scaling toward a self‑service motion where customers onboard, ingest content, and operate the product without manual implementation. That requires an infrastructure foundation that can handle multi‑tenant scale, high reliability, and the unique demands of generative AI workloads — without ballooning the AWS bill.

We’re hiring a Principal Platform Engineer to own that foundation. This is a senior individual‑contributor role with broad architectural authority. You will not have direct reports. You will set the technical direction for our cloud infrastructure, partner with engineering on production scaling decisions, and operate the platform with the discipline a SOC 2 / ISO 27001 customer base requires.

What you’ll own
  • Production Kubernetes (EKS) clusters: capacity planning, node group strategy, gen‑AI workload isolation, blast‑radius containment.
  • AWS infrastructure end‑to‑end: RDS, DMS, Kafka (MSK), ECR, networking, IAM, multi‑region deployments (including Ireland for EU data residency).
  • Infrastructure‑as‑code in Terraform — modules, environments, drift management, peer review.
  • CI/CD pipelines (Jenkins, GitHub Actions, or your recommended replacement) — fast, reliable, secure builds for backend and frontend services.
  • Observability: Grafana dashboards, Prometheus metrics, log pipelines, on‑call alerting, SLO definition.
  • Cost optimization. AWS spend is one of our top three variable costs. Reducing it by 20% is a tangible objective for this seat.
  • Security posture: secrets management (Consul/Vault), IAM hygiene, vulnerability patching, support for SOC 2 and ISO 27001 audit cycles.
  • Architecture leadership on the self‑service infrastructure roadmap: how we onboard a customer without human intervention and scale to 10x our current tenant count.
  • Documentation and runbooks that let the rest of the engineering team operate the platform when you’re unavailable.
Must‑haves
  • 8+ years of platform, infrastructure, SRE, or DevOps experience, with at least 3+ years operating production Kubernetes at scale.
  • Deep AWS expertise across compute, storage, networking, data services, and IAM.
  • Production fluency with Terraform, Docker, Linux, and CI/CD systems.
  • Track record of architectural decisions that materially improved reliability, cost, or developer velocity — with specific, measurable outcomes you can point to.
  • Comfort operating as a senior IC who sets technical direction across teams without formal authority.
  • Strong written communication — runbooks, architecture decision records, post‑incident reviews.
  • Willingness to be in‑office Tuesday through Thursday in Denver.
Nice‑to‑haves
  • Production experience supporting generative AI or ML workloads (GPU node groups, vector databases, model serving).
  • Experience with Qdrant, Pinecone, Weaviate, or other vector stores in production.
  • PostgreSQL operational depth — replication, performance tuning, backup/restore.
  • Experience scaling a multi‑tenant SaaS platform from ~100 customers to ~1,000.
  • SOC 2 Type II and ISO 27001 audit experience.
  • Familiarity with event‑driven architectures (Kafka, Kinesis, or equivalent).
What success looks like
First 30 days
  • Complete a written audit of our current infrastructure: what we have, where the risks are, what’s costing us money.
  • Establish on‑call rotation participation and respond to your first production incident.
  • Identify the top three architectural debt items.
First 60 days
  • Deliver first architectural recommendation with implementation plan — typically cost optimization or scaling bottleneck.
  • Refresh and own the observability stack.
  • Document the production runbook for the rest of the engineering team.
First 90 days
  • Ship a measurable improvement: cost reduction, reliability uplift, deployment velocity, or scale headroom.
  • Deliver the multi‑tenant scale roadmap for the self‑service motion.
  • Establish quarterly architecture review cadence with the engineering team.
Why Ombud

You’ll own the platform that runs production AI for some of the largest enterprise software companies in the world. The infrastructure decisions you make will directly enable our 2026 strategy of moving from response management to autonomous revenue execution. You’ll work with a small, senior engineering team that ships fast and trusts each other.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Full Stack AI Engineer
Senior Full Stack AI Engineer

Ombud • Denver (CO)

Hybrid
USD 120,000 - 150,000
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies • Palo Alto (CA)

On-site
USD 190,000 - 270,000
Platform Engineer
Platform Engineer

AirOps • San Francisco (CA), New York (NY)

On-site
USD 130,000 - 160,000
Equity in a fast-growing startup
Competitive benefits package
Flexible time off policy
+2
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies Inc • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Platform Engineer
Platform Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2
Senior Platform Engineer (Cloud Platform)
Senior Platform Engineer (Cloud Platform)

Amplitude • San Francisco (CA)

Hybrid
USD 130,000 - 170,000
Founding Infrastructure Engineer
Founding Infrastructure Engineer

Modern Relay • United States

On-site
USD 140,000 - 200,000
Principal Software Engineer (Infra/Platform)
Principal Software Engineer (Infra/Platform)

Gradial • Seattle (WA)

On-site
USD 180,000 - 235,000
Performance-based bonus
Equity awards
Medical, dental & vision
+6
Infrastructure Engineer
Infrastructure Engineer

Overland AI • Seattle (WA)

On-site
USD 130,000 - 225,000
Competitive salary: $130K – $225K annually
Equity compensation
Best-in-class healthcare, dental, and vision plans
+3
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)
Senior / Lead / Principal Platform Engineer (DevOps / Cloud Infrastructure)

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Long-term career growth opportunities