Staff SRE: ML & Real-Time Analytics Reliability Leader

Stratitech Services LLC

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

StratITech is hiring for a Staff-level reliability role focused on building and governing the next generation of high-scale ML infrastructure in San Francisco. You will own production reliability for ML and real-time analytics, drive CI/CD strategy, and set robust observability standards across teams.

The role requires hands-on leadership in Kubernetes, IaC, and incident response, with direct access to engineering leadership.

Qualifications

  • Deep experience operating Linux infrastructure in production.
  • Proven impact as a Staff SRE, Senior SRE, or senior DevOps/Platform Engineer.
  • Experience supporting complex, data-intensive or ML-driven systems in production.
  • Strong hands-on experience with Docker and Kubernetes.
  • Strong scripting ability (Bash and/or Python).
  • CI/CD ownership experience (GitHub Actions, ArgoCD, or similar).
  • Experience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
  • Ability to debug systemic failures across infrastructure, deployments, and workloads.
  • Clear communicator across engineering and data teams.
  • Engineers who evolved from infrastructure foundations into reliability leaders will thrive here.

Responsibilities

  • Production reliability for ML and real-time analytics workloads.
  • CI/CD strategy, deployment automation, and rollback design.
  • Observability frameworks (SLOs, alerting, monitoring, incident response).
  • Infrastructure-as-Code and Kubernetes environments.
  • Capacity planning and performance optimization.
  • Post-incident reviews driving long-term reliability improvements.
  • Reliability standards across teams, not just a single service.
  • Partner with engineering and data science teams to ensure ML workloads are production-ready.

Skills

Linux
Staff SRE
Docker
Kubernetes
Scripting Bash/Python
CI/CD ownership
Observability stacks
Cross-team collaboration
Production ML/real-time

Tools

Docker
Kubernetes

Job description

StratITech is hiring for a Staff-level reliability role focused on building and governing the next generation of high-scale ML infrastructure in San Francisco. You will own production reliability for ML and real-time analytics, drive CI/CD strategy, and set robust observability standards across teams.

The role requires hands-on leadership in Kubernetes, IaC, and incident response, with direct access to engineering leadership.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stratitech Services LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Competitive compensation with equity
100% medical, dental, and vision coverage
Generous PTO including Winter Break
+2
Staff SRE: AI-Driven Reliability & Platform Architect
Staff SRE: AI-Driven Reliability & Platform Architect

Devopsroles • Northern (KY)

Remote
USD 150,000 - 225,000
Equity
Benefits program
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Staff SRE: AI-Driven Reliability Architect (Hybrid)
Staff SRE: AI-Driven Reliability Architect (Hybrid)

EarnIn Bfwf • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Staff Site Reliability Engineer: Lead Reliability at Global Scale
Staff Site Reliability Engineer: Lead Reliability at Global Scale

Attentive • United States

On-site
USD 180,000 - 240,000
Health benefits
Equity
Flexible work
Site Reliability Engineer — ML Infra & Observability
Site Reliability Engineer — ML Infra & Observability

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Staff ML Platform Engineer: Scalable AI Infra & Systems
Staff ML Platform Engineer: Scalable AI Infra & Systems

Stripe • Seattle (WA)

On-site
USD 224,000 - 336,000
Senior SRE: AI-Driven Reliability & Incident Leadership
Senior SRE: AI-Driven Reliability & Incident Leadership

Salesforce • San Francisco (CA)

On-site
USD 149,000 - 246,000