Staff DevOps Engineer

Engg

San Francisco (CA)

On-site

USD 170,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nexxa.ai is seeking a Senior/Staff DevOps Engineer to own and operate the production infrastructure powering AI and industrial workloads. You will design and run CI/CD pipelines, build IaC for cloud and edge deployments, and manage Kubernetes platforms with GPU scheduling.

You will collaborate with AI, data, and product teams to ensure reliability, security, and performance at scale, mentoring engineers and driving platform strategy across the organization.

Qualifications

  • 6+ years in DevOps/SRE/Platform or infra‑focused roles
  • Production scale cloud platforms (AWS/GCP/Azure)
  • Kubernetes in production with GPU workloads
  • Infrastructure‑as‑code tooling (Terraform, Pulumi)
  • CI/CD systems (GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
  • Observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry)
  • ML/AI infra experience (training clusters, model serving, data pipelines)
  • Scripting/programming (Python, Go, Bash) for automation

Responsibilities

  • Own and evolve core infrastructure — compute, networking, storage, deployments
  • Design and operate CI/CD pipelines for AI, data, and product teams
  • Build and maintain infrastructure‑as‑code for reproducible environments across cloud and edge
  • Architect and manage Kubernetes platforms for training, inference, workloads
  • Support data warehouses, feature stores, and model training/inference infra
  • Define observability practices — metrics, logging, tracing, alerting
  • Establish reliability practices: SLOs/SLIs, incident response, postmortems
  • Design for security and compliance across cloud infra and secrets
  • Collaborate on infra roadmap and mentor engineers on best practices

Skills

Cloud platforms
SRE mindset
Automation
Python/Go/Bash

Tools

Terraform
Pulumi
Kubernetes
GitHub Actions
GitLab CI
CircleCI
Jenkins
ArgoCD
Prometheus
Grafana
Datadog
OpenTelemetry
Snowflake
BigQuery
Redshift
Databricks

Job description

ABOUT THE ROLE

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments. This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

WHAT YOU'LL DO
  • Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end
  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
  • Partner with data and AI teams to support the infrastructure behind:
    • Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)
    • Feature stores, embedding indices, and retrieval pipelines
    • Model training, evaluation, and serving infrastructure
  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity
  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy
  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org
REQUIRED QUALIFICATIONS
  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
  • Deep hands-on experience with:
    • Cloud platforms (AWS, GCP, or Azure) at production scale
    • Kubernetes in production, including GPU workload scheduling
    • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)
    • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
    • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
    • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus
    • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
    • Proven ability to independently scope and lead infrastructure projects from design through production rollout
    • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause
PREFERRED QUALIFICATIONS
  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)
  • History of building internal developer platforms or self-service infrastructure tooling
  • Experience scaling infrastructure teams or setting technical direction at a Staff level
WHAT SUCCESS LOOKS LIKE
  • You can own ambiguous, high-stakes infrastructure problems end-to-end
  • Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today
  • You bring strong technical judgment on tradeoffs between reliability, cost, and speed
  • You raise the bar for operational rigor and engineering discipline across the team
  • You help define what's next for the platform, not just execute what's known
WHY JOIN NEXXA.AI
  • Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies
  • Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement
  • Professional Growth: Benefit from significant opportunities for career development and advancement
  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.AI • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.ai • Sunnyvale (CA)

On-site
USD 200,000 - 260,000
Equity
Competitive compensation
Security & Infrastructure Engineer
Security & Infrastructure Engineer

Nexxa.AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Security & Infrastructure Engineer
Security & Infrastructure Engineer

Nexxa.ai • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Innovative Environment
Collaborative Culture
Professional Growth
+1
Backend AI Engineer
Backend AI Engineer

Nexxa.ai • Sunnyvale (CA)

On-site
USD 140,000 - 220,000
Backend AI Engineer
Backend AI Engineer

Engg • San Francisco (CA)

On-site
USD 180,000 - 240,000
Backend AI Engineer
Backend AI Engineer

Nexxa.AI • San Francisco (CA)

On-site
USD 140,000 - 230,000
Security & Infrastructure Engineer
Security & Infrastructure Engineer

Engg • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff DevOps Engineer - AI/ML Infra, Kubernetes & Scale
Staff DevOps Engineer - AI/ML Infra, Kubernetes & Scale

Engg • San Francisco (CA)

On-site
USD 170,000 - 210,000
Staff DevOps Engineer - AI Infra for Production at Scale
Staff DevOps Engineer - AI Infra for Production at Scale

Nexxa.AI • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity