Staff DevOps Engineer: Industrial AI Infra at Scale

Nexxa.ai

Sunnyvale (CA)

On-site

USD 200,000 - 260,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Competitive compensation

Job summary

Nexxa.ai in Sunnyvale, CA is hiring a Senior/Staff DevOps Engineer to own and scale the infrastructure powering AI and industrial systems. You will run GPU-backed training and inference clusters, design reproducible environments, and partner with AI/data teams to keep production workloads fast and reliable.

You will build pipelines, manage Kubernetes platforms, and drive observability, security, and reliability across cloud and on-prem deployments.

Qualifications

  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles.
  • Deep hands-on experience with Cloud platforms (AWS, GCP, or Azure) at production scale.
  • Kubernetes in production, including GPU workload scheduling.
  • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent).
  • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD).
  • Strong track record designing and operating observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry).
  • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus.
  • Excellent scripting/programming skills (Python, Go, or Bash).
  • Proven ability to independently scope and lead infrastructure projects from design through production rollout.
  • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause.

Responsibilities

  • Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end.
  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams.
  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments.
  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling.
  • Partner with data and AI teams to support the infrastructure behind data warehouses and lakehouse architectures, feature stores, embedding indices, and retrieval pipelines.
  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems.
  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations.
  • Design for security and compliance across cloud infrastructure, secrets management, and access control.
  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity.
  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy.
  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org.

Skills

Kubernetes in production
Cloud platforms (AWS/GCP/Azure)
Infrastructure-as-code (Terraform/Pulu
CI/CD tooling
Observability stacks
Automation scripting (Python/Go/Bash)
Incident management

Tools

Terraform
Pulumi
GitHub Actions
GitLab CI
CircleCI
Jenkins
ArgoCD
Prometheus
Grafana
Datadog

Job description

Nexxa.ai in Sunnyvale, CA is hiring a Senior/Staff DevOps Engineer to own and scale the infrastructure powering AI and industrial systems. You will run GPU-backed training and inference clusters, design reproducible environments, and partner with AI/data teams to keep production workloads fast and reliable.

You will build pipelines, manage Kubernetes platforms, and drive observability, security, and reliability across cloud and on-prem deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff DevOps Engineer - AI Infra for Production at Scale
Staff DevOps Engineer - AI Infra for Production at Scale

Nexxa.AI • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.ai • Sunnyvale (CA)

On-site
USD 200,000 - 260,000
Equity
Competitive compensation
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.AI • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Backend AI Engineer — Scalable ML Infra & APIs
Backend AI Engineer — Scalable ML Infra & APIs

Nexxa.ai • Sunnyvale (CA)

On-site
USD 140,000 - 220,000
Backend AI Engineer — Scalable ML Infra & APIs
Backend AI Engineer — Scalable ML Infra & APIs

Nexxa.AI • San Francisco (CA)

On-site
USD 140,000 - 230,000
Staff Cloud-Native Engineer for AI Infrastructure
Staff Cloud-Native Engineer for AI Infrastructure

Nscale • Houston (TX), Northern (KY)

Hybrid
USD 220,000 - 265,000
Medical benefits
Dental benefits
Flexible PTO
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Lead QA Engineer - Autonomous AI Agents for Industry
Lead QA Engineer - Autonomous AI Agents for Industry

Nexxa.AI • Sunnyvale (CA)

On-site
USD 140,000 - 200,000
Competitive compensation
Equity package
Senior AI Infra Engineer - Scalable Cloud Platform
Senior AI Infra Engineer - Scalable Cloud Platform

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits