Member of Technical Staff, DevOps

Reactor

San Francisco (CA)

On-site

USD 100,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage

Job summary

Reactor is looking for a DevOps/SRE engineer in San Francisco to enhance the reliability and observability of their AI platform. The position requires running production Kubernetes clusters, strong CI/CD pipeline experience, and a solid understanding of GitOps and infrastructure as code. Successful candidates will triage production issues, manage secret infrastructure, and define SLOs. Benefits include a competitive salary, equity, and health coverage, with opportunities for relocation support.

Qualifications

  • Proven experience in running production Kubernetes clusters.
  • Strong background in CI/CD practices and tools.
  • Expert in GitOps and reconciliations.
  • Fluency in Infrastructure as Code with Terraform.
  • Experience in setting up observability and defining SLOs.
  • Proficiency in secret management practices.
  • Experience in incident response with effective postmortems.
  • Ability to write scripts in Go, Python, or Bash.

Responsibilities

  • Own and evolve CI/CD pipelines and deployment lifecycles.
  • Maintain observability stack for all services.
  • Define SLOs and manage incident response.
  • Handle infrastructure as code across different cloud providers.
  • Operate secret management and ensure deployment safety.

Skills

Kubernetes management
CI/CD pipeline development
GitOps practices
Infrastructure as Code
Observability and monitoring
Incident response
Coding proficiency in Go, Python, or Bash
Secret management

Tools

Terraform
Helm

Job description

Description

We're looking for a DevOps / SRE engineer to own the reliability, delivery, and observability of our AI platform. You'll be the person who ensures models get from a developer's branch to production without anyone losing sleep — and when something does go wrong at 2am, you'll be the one who knows where to look.

Department: Engineering

Location: San Francisco

We run production across multiple Kubernetes clusters, cloud providers, and regions. Our deployment pipeline is fully automated through CI/CD and GitOps, our infrastructure is managed as code, and our observability stack gives us full visibility across every service and GPU workload. This role is about making all of that faster, more reliable, and easier to operate as we scale.

What You'll Do
  • Own and evolve our CI/CD pipelines: dynamic pipeline generation across a monorepo of Go services, Python model containers, and Helm charts
  • Operate and improve our GitOps deployment lifecycle: Helm releases, Kustomizations, and image automation across multiple clusters
  • Build and maintain our observability stack: distributed tracing, metrics, dashboards, and alerting across all services and GPU workloads
  • Define and track SLOs for core platform services, including session latency, model cold start time, and streaming reliability
  • Run incident response: triage production issues, write postmortems, build runbooks, and drive reliability improvements
  • Manage infrastructure-as-code across multiple cloud providers and regions: plan/apply workflows, state management, drift detection
  • Operate secret management: encrypted secrets, external secret syncing, certificate automation
  • Improve deployment safety: canary rollouts, health checks, startup probes, rollback automation
  • Manage authentication infrastructure: OIDC federation for CI, workload identity for cloud services, cross-cloud credential management
  • Participate in on-call rotation and build the tooling that makes on-call less painful
What We're Looking For
  • You've run production Kubernetes clusters and been on-call for them. You've debugged node scheduling failures, OOM kills, and mysterious pod evictions at 3am
  • Strong CI/CD experience: you've built and maintained pipelines for monorepos, not just single-service repos
  • GitOps experience: you understand reconciliation loops, drift detection, and why image automation matters
  • Infrastructure-as-code fluency with Terraform or similar across multiple environments and cloud accounts
  • You know observability beyond just "set up dashboards". You've defined SLOs, built alerting that doesn't page on noise, and used traces to debug cross-service latency issues.
  • Comfortable with secret management patterns (KMS, encrypted configs, external secret operators). You've thought about credential rotation and zero-trust.
  • Incident response experience: you've triaged production outages, written postmortems that actually led to improvements, and built runbooks that other engineers could follow
  • You write code, not just YAML. Proficiency in Go, Python, or Bash for building tooling, automation, and pipeline scripts
Nice to Have
  • Experience with GPU workloads on Kubernetes: device plugins, GPU-aware scheduling, GPU monitoring
  • Multi-cloud operations beyond a single provider
  • Real-time or streaming workloads: low-latency systems where p99 matters more than average
  • Experience with Helm chart authoring and managing complex value layering across environments
  • Familiarity with real-time media or relay infrastructure
  • FinOps experience: GPU cost optimization, spot/preemptible instance management
What We're Not Looking For
  • Engineers who treat infrastructure-as-code as "click around in the console and import later"
  • SREs who've only monitored systems but never built the deployment pipelines that ship to them
  • Candidates whose CI/CD experience is limited to GitHub Actions for a single-service repo
  • People who write alerts that fire every day and then get ignored
Logistics

We are based in-person in San Francisco. We are also hiring for this role in Europe for on-call coverage and timezone distribution.

Benefits
  • Competitive salary and meaningful early equity
  • Visa sponsorship and relocation support
  • Generous health, dental, and vision coverage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Infrastructure & SRE Engineer — Secure AI Platform, Equity
Infrastructure & SRE Engineer — Secure AI Platform, Equity

Crosscheck Staffing • San Francisco (CA)

On-site
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior Site Reliability Engineer -AI Infrastructure Operations
Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
SRE / DevOps Engineer
SRE / DevOps Engineer

HeadHR • Town of Poland (NY)

On-site
USD 120,000 - 150,000
AI Platform DevOps & SRE Lead
AI Platform DevOps & SRE Lead

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Emergentlabsinc • San Francisco (CA)

On-site
USD 110,000 - 150,000
401(k)
Health, dental, and vision insurance
Unlimited Paid Time Off
+1