Senior Principal DevOps Engineer — AI Infra & Kubernetes

Upscale AI

United States

Hybrid

USD 263,000 - 284,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Upscale AI is building high‑performance AI infrastructure. We’re seeking an experienced SRE/DevOps leader to own reliability, deployment, and the platform infrastructure behind Orchestrator and AI Fabric environments.

You will operate Kubernetes clusters across on‑prem and cloud, design CI/CD pipelines, manage Terraform‑driven infrastructure, and handle secret rotations. You’ll also implement observability stacks and contribute to tooling that improves operational efficiency.

Qualifications

  • 8–14 years in SRE, DevOps, or infrastructure engineering roles supporting production distributed systems.
  • Deep Kubernetes expertise across cluster administration, networking, storage, RBAC, and troubleshooting in cloud and bare‑metal.
  • Strong Terraform skills with multi‑environment, multi‑provider infrastructure at scale.
  • Hands‑on experience building and operating CI/CD pipelines end‑to‑end.
  • Production experience with observability tools: Prometheus, Grafana, Loki, Splunk, Datadog, or comparable.
  • Scripting in Python, Bash, or Go; Linux systems internals and networking knowledge.
  • Experience managing TLS/mTLS certificates and Vault or equivalents in production.
  • Comfort working across AWS/GCP and on‑prem environments.

Responsibilities

  • Own reliability, deployment, and operational infrastructure for Orchestrator and AI Fabric environments.
  • Build and maintain Kubernetes clusters across on‑prem and cloud; design CI/CD pipelines with automated testing gates.
  • Manage Terraform‑driven infrastructure with proper state management and drift detection.
  • Operate the full observability stack (Prometheus, Grafana, Loki, Splunk, Datadog, Timestream) for end‑to‑end visibility.
  • Automate secret and certificate rotations (TLS/MTLS, Vault) and credential lifecycle for multi‑tenant deployments.
  • Lead incident response: runbooks, root cause analysis, post‑incident reviews with real fixes.
  • Support customer deployments at site level; handle edge appliances and varying network constraints.
  • Develop internal tooling and analytics to reduce toil and speed up engineering.
  • Plan capacity, optimize costs, and forecast infrastructure needs as deployments scale.

Skills

Kubernetes
Terraform
CI/CD
Prometheus
Grafana
Loki
Datadog
Timestream
Splunk
Python
Bash
Go
Linux
TLS/MTLS
AWS
GCP
Troubleshooting

Tools

Rancher
Helm
GitHub Actions
Vault

Job description

Upscale AI is building high‑performance AI infrastructure. We’re seeking an experienced SRE/DevOps leader to own reliability, deployment, and the platform infrastructure behind Orchestrator and AI Fabric environments.

You will operate Kubernetes clusters across on‑prem and cloud, design CI/CD pipelines, manage Terraform‑driven infrastructure, and handle secret rotations. You’ll also implement observability stacks and contribute to tooling that improves operational efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra SRE Lead — Kubernetes, CI/CD & Observability
AI Infra SRE Lead — Kubernetes, CI/CD & Observability

The Consensus • United States

On-site
USD 263,000 - 284,000
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud

StackAI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote work available
Office near Salesforce Park
Principal Platform Engineer — AI-Driven Infra Leader
Principal Platform Engineer — AI-Driven Infra Leader

LinkedIn • California (MO)

Hybrid
USD 226,000 - 369,000
Senior Staff DevOps Engineer – Orchestration
Senior Staff DevOps Engineer – Orchestration

Upscale AI • United States

Hybrid
USD 263,000 - 284,000
AI Infra Platform Lead — Kubernetes, Terraform & Ansible
AI Infra Platform Lead — Kubernetes, Terraform & Ansible

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+1
Senior AI Platform Engineer – Kubernetes & Infra Lead
Senior AI Platform Engineer – Kubernetes & Infra Lead

CB Smart Recruit • Los Angeles (CA)

On-site
USD 200,000 - 300,000
Competitive sign-on bonus
Comprehensive benefits package
Significant technical ownership
Senior DevOps Engineer: Cloud, Kubernetes & CI/CD
Senior DevOps Engineer: Cloud, Kubernetes & CI/CD

StackAI • New York (NY)

Hybrid
USD 130,000 - 210,000
Hybrid work model
Remote days
Office in SF
Sr AI Infra Lead - K8s & Automation
Sr AI Infra Lead - K8s & Automation

Seekr • Austin (TX)

Hybrid
USD 140,000 - 200,000
Equity RSUs
Unlimited PTO + 14 holidays
Hybrid work - Reston, VA & Austin, TX
+4
AI Infra Engineer: Kubernetes & LLM Ops
AI Infra Engineer: Kubernetes & LLM Ops

Alignity Solutions • New York (NY)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000