Staff DevOps Engineer - AI Infra for Production at Scale

Nexxa.AI

San Francisco (CA)

On-site

USD 170,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity

Job summary

Nexxa.ai is seeking a Senior/Staff DevOps Engineer to own and operate the infrastructure that powers production AI workloads at scale. You will ensure reliable GPU-backed training, inference, and data pipelines across cloud and on-prem environments.

Collaborate with AI, data, and product teams to design scalable, observable systems, implement IaC, and drive reliability practices with strong incident response and security focus.

Qualifications

  • 6+ years of experience in DevOps, SRE, Platform Engineering, or infra-focused roles.
  • Deep hands-on experience with cloud platforms (AWS, GCP, or Azure) at production scale.
  • Kubernetes in production with GPU workload scheduling.
  • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent).
  • CI/CD systems (GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD).
  • Strong observability stack experience (Prometheus, Grafana, Datadog, OpenTelemetry).
  • Experience supporting ML/AI infrastructure — training clusters and model serving.
  • Excellent scripting/programming skills (Python, Go, or Bash).
  • Proven ability to lead infra projects end-to-end; strong incident management.

Responsibilities

  • Own and evolve Nexxa's core infrastructure end-to-end.
  • Design and operate CI/CD pipelines for AI, data, and product teams.
  • Build and maintain infrastructure-as-code for cloud and on-prem/edge deployments.
  • Architect and manage Kubernetes-based platforms for training, inference, and workloads.
  • Collaborate with data/AI teams to support data warehouses, feature stores, and model infra.
  • Define observability practices: metrics, logging, tracing, and alerting.
  • Establish reliability practices: SLOs/SLIs, incident response, postmortems, on-call rotations.
  • Design for security and compliance across cloud infra and access control.
  • Balance cost, latency, reliability, and developer velocity; contribute to platform roadmap.
  • Mentor engineers and raise the bar for operational excellence.

Job description

Nexxa.ai is seeking a Senior/Staff DevOps Engineer to own and operate the infrastructure that powers production AI workloads at scale. You will ensure reliable GPU-backed training, inference, and data pipelines across cloud and on-prem environments.

Collaborate with AI, data, and product teams to design scalable, observable systems, implement IaC, and drive reliability practices with strong incident response and security focus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff DevOps Engineer - AI/ML Infra, Kubernetes & Scale
Staff DevOps Engineer - AI/ML Infra, Kubernetes & Scale

Engg • San Francisco (CA)

On-site
USD 170,000 - 210,000
Staff DevOps Engineer: Industrial AI Infra at Scale
Staff DevOps Engineer: Industrial AI Infra at Scale

Nexxa.ai • Sunnyvale (CA)

On-site
USD 200,000 - 260,000
Equity
Competitive compensation
Staff DevOps Engineer
Staff DevOps Engineer

Engg • San Francisco (CA)

On-site
USD 170,000 - 210,000
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.ai • Sunnyvale (CA)

On-site
USD 200,000 - 260,000
Equity
Competitive compensation
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.AI • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Backend AI Platform Engineer – Scalable ML Infra
Backend AI Platform Engineer – Scalable ML Infra

Engg • San Francisco (CA)

On-site
USD 180,000 - 240,000
Backend AI Engineer — Scalable ML Infra & APIs
Backend AI Engineer — Scalable ML Infra & APIs

Nexxa.ai • Sunnyvale (CA)

On-site
USD 140,000 - 220,000
Backend AI Engineer — Scalable ML Infra & APIs
Backend AI Engineer — Scalable ML Infra & APIs

Nexxa.AI • San Francisco (CA)

On-site
USD 140,000 - 230,000
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior MLOps Engineer — AI Infra & Scale Leader
Senior MLOps Engineer — AI Infra & Scale Leader

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package