Senior Staff DevOps Engineer – Orchestration

Upscale AI

United States

On-site

USD 263,000 - 284,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Upscale AI is building high‑performance AI infrastructure. We’re seeking an experienced SRE/DevOps leader to own reliability, deployment, and the platform infrastructure behind Orchestrator and AI Fabric environments.

You will operate Kubernetes clusters across on‑prem and cloud, design CI/CD pipelines, manage Terraform‑driven infrastructure, and handle secret rotations. You’ll also implement observability stacks and contribute to tooling that improves operational efficiency.

Qualifications

  • 8–14 years in SRE, DevOps, or infrastructure engineering roles supporting production distributed systems.
  • Deep Kubernetes expertise across cluster administration, networking, storage, RBAC, and troubleshooting in cloud and bare‑metal.
  • Strong Terraform skills with multi‑environment, multi‑provider infrastructure at scale.
  • Hands‑on experience building and operating CI/CD pipelines end‑to‑end.
  • Production experience with observability tools: Prometheus, Grafana, Loki, Splunk, Datadog, or comparable.
  • Scripting in Python, Bash, or Go; Linux systems internals and networking knowledge.
  • Experience managing TLS/mTLS certificates and Vault or equivalents in production.
  • Comfort working across AWS/GCP and on‑prem environments.

Responsibilities

  • Own reliability, deployment, and operational infrastructure for Orchestrator and AI Fabric environments.
  • Build and maintain Kubernetes clusters across on‑prem and cloud; design CI/CD pipelines with automated testing gates.
  • Manage Terraform‑driven infrastructure with proper state management and drift detection.
  • Operate the full observability stack (Prometheus, Grafana, Loki, Splunk, Datadog, Timestream) for end‑to‑end visibility.
  • Automate secret and certificate rotations (TLS/MTLS, Vault) and credential lifecycle for multi‑tenant deployments.
  • Lead incident response: runbooks, root cause analysis, post‑incident reviews with real fixes.
  • Support customer deployments at site level; handle edge appliances and varying network constraints.
  • Develop internal tooling and analytics to reduce toil and speed up engineering.
  • Plan capacity, optimize costs, and forecast infrastructure needs as deployments scale.

Skills

Kubernetes
Terraform
CI/CD
Prometheus
Grafana
Loki
Datadog
Timestream
Splunk
Python
Bash
Go
Linux
TLS/MTLS
AWS
GCP
Troubleshooting

Tools

Rancher
Helm
GitHub Actions
Vault

Job description

Why join Upscale AI

Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.

We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high‑performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real‑world impact.

If you’re looking to do high‑impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.

About the role

Own the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.

You will build and maintain Kubernetes clusters across on‑prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform‑driven infrastructure‑as‑code, and handle secret and certificate rotations.

You will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end‑to‑end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.

Beyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast‑moving team.

What you'll work on
  • Kubernetes cluster lifecycle across hybrid environments: provision, upgrade, scale, and harden clusters running on bare‑metal (on‑prem customer datacenters) and cloud (AWS/GCP) using kubeadm, Rancher, or equivalent tooling
  • CI/CD pipeline design and ownership: build and maintain pipelines (GitHub Actions or equivalent) that deliver Go microservices, React UI, Helm charts, and edge appliance images from commit to production with automated testing gates
  • Infrastructure‑as‑code: manage all cloud and on‑prem infrastructure through Terraform modules with proper state management, drift detection, and PR‑based review workflows
  • Observability stack operations: deploy, tune, and maintain Prometheus (metrics), Grafana (dashboards), Loki (logs), Splunk and Datadog (enterprise monitoring), and Timestream (time‑series analytics) — build the dashboards and alerts that give the team real‑time visibility into platform and customer‑site health
  • Secret and certificate management: automate mTLS certificate rotation across hub‑to‑edge communication channels, manage Vault or equivalent secret stores, and handle credential lifecycle for multi‑tenant deployments
  • Incident response and debugging: own the runbooks, triage production issues across the distributed hub/edge architecture, perform root cause analysis, and drive post‑incident reviews that result in real fixes — not just documents
  • Customer site operations: support deployment, upgrade, and troubleshooting of edge appliances running in customer datacenters with varying network constraints and access patterns
  • Internal tooling: build CLI tools, deployment automation, environment provisioners, and operational dashboards that reduce toil and make the engineering team faster
  • Capacity planning and cost optimization: monitor resource utilization across clusters and cloud accounts, right‑size workloads, and forecast infrastructure needs as site count grows
What you bring
  • 8–14 years in SRE, DevOps, or infrastructure engineering roles supporting production distributed systems
  • Deep Kubernetes expertise: cluster administration, networking (CNI, ingress, service mesh), storage (PV/PVC, CSI drivers), RBAC, and troubleshooting pod/node‑level issues in both cloud and bare‑metal environments
  • Strong Terraform skills with experience managing multi‑environment, multi‑provider infrastructure at scale
  • Hands‑on experience building and operating CI/CD pipelines end‑to‑end — not just configuring someone else’s templates
  • Production experience with at least three of: Prometheus, Grafana, Loki, Splunk, Datadog, Timestream, or comparable observability tools
  • Solid scripting and automation skills in Python, Bash, or Go
  • Working knowledge of Linux systems internals: networking (iptables, DNS, TCP debugging), storage, process management, and performance analysis
  • Experience managing TLS/mTLS certificates, secret rotation, and Vault or equivalent in production
  • Comfort working across cloud (AWS/GCP) and on‑prem environments with different constraints and access models
  • Strong debugging instincts — you can follow a problem from a user report through load balancers, ingress, service mesh, application logs, and database queries to root cause
Nice to have
  • Experience supporting network infrastructure or datacenter automation platforms
  • Helm chart authoring and management for complex multi‑service applications
  • Bare‑metal Kubernetes provisioning (not just managed EKS/GKE)
  • eBPF‑based observability tools (Cilium, Pixie, Hubble)
  • Experience operating Kafka, ClickHouse, ArangoDB, or Redis in production
  • Familiarity with SONiC, network switch management, or ZTP workflows
  • On‑call experience with a structured incident management process (PagerDuty, Opsgenie)
  • SOC 2, FedRAMP, or equivalent compliance experience for infrastructure

$263,000 - $284,000 a year

Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.

Equal Opportunity

Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.

Accessibility & Accommodations

We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us athiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Support Senior Staff Engineer
Technical Support Senior Staff Engineer

The Consensus • United States

On-site
USD 200,000 - 216,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Technical Support Principal Engineer – AI Network
Technical Support Principal Engineer – AI Network

The Consensus • United States

On-site
USD 248,000 - 269,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BetterUp • Austin (TX)

Hybrid
USD 147,000 - 185,000
Access to BetterUp coaching
Medical, dental, and vision insurance
Flexible paid time off
+2
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 319,000
Equity
Bonus
Health insurance
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Access to BetterUp coaching
Competitive compensation plan
Medical, dental, and vision insurance
+3
Software Engineer, Delivery / CD
Software Engineer, Delivery / CD

Slope • San Francisco (CA)

On-site
USD 230,000 - 490,000
Senior Engineer, Platform Infrastructure
Senior Engineer, Platform Infrastructure

Vts • New York (NY)

On-site
USD 160,000 - 200,000
Competitive salary
Comprehensive health benefits
401(k) plan
+2
Infrastructure Engineer
Infrastructure Engineer

Overland AI • Seattle (WA)

On-site
USD 130,000 - 225,000
Competitive salary: $130K – $225K annually
Equity compensation
Best-in-class healthcare, dental, and vision plans
+3
Deployment Engineering Manager, Enterprise
Deployment Engineering Manager, Enterprise

Scale AI • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2