Staff Site Reliability Engineer

Stratitech Services LLC

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

StratITech is hiring for a Staff-level reliability role focused on building and governing the next generation of high-scale ML infrastructure in San Francisco. You will own production reliability for ML and real-time analytics, drive CI/CD strategy, and set robust observability standards across teams.

The role requires hands-on leadership in Kubernetes, IaC, and incident response, with direct access to engineering leadership.

Qualifications

  • Deep experience operating Linux infrastructure in production.
  • Proven impact as a Staff SRE, Senior SRE, or senior DevOps/Platform Engineer.
  • Experience supporting complex, data-intensive or ML-driven systems in production.
  • Strong hands-on experience with Docker and Kubernetes.
  • Strong scripting ability (Bash and/or Python).
  • CI/CD ownership experience (GitHub Actions, ArgoCD, or similar).
  • Experience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
  • Ability to debug systemic failures across infrastructure, deployments, and workloads.
  • Clear communicator across engineering and data teams.
  • Engineers who evolved from infrastructure foundations into reliability leaders will thrive here.

Responsibilities

  • Production reliability for ML and real-time analytics workloads.
  • CI/CD strategy, deployment automation, and rollback design.
  • Observability frameworks (SLOs, alerting, monitoring, incident response).
  • Infrastructure-as-Code and Kubernetes environments.
  • Capacity planning and performance optimization.
  • Post-incident reviews driving long-term reliability improvements.
  • Reliability standards across teams, not just a single service.
  • Partner with engineering and data science teams to ensure ML workloads are production-ready.

Skills

Linux
Staff SRE
Docker
Kubernetes
Scripting Bash/Python
CI/CD ownership
Observability stacks
Cross-team collaboration
Production ML/real-time

Tools

Docker
Kubernetes

Job description

You will define how mission-critical machine learning and real-time analytics systems operate in production — influencing reliability strategy, deployment standards, and infrastructure architecture across engineering.

This team operates in a highly collaborative, in-person engineering environment in SOMA. Infrastructure, ML, and engineering leaders work side by side to design, build, and operate complex systems in real time. The pace is fast, the feedback loops are tight, and decisions happen quickly.

If you’ve grown from Linux systems → DevOps → Staff-level SRE , and you now think in terms of systemic risk, scalability, and long-term reliability strategy — this role gives you direct influence and visibility.

This role is intentionally in-person because:

  • Reliability decisions happen at architectural depth — not over Slack threads
  • ML, data, and infrastructure teams collaborate continuously in real time
  • Post-incident reviews, system design debates, and performance tuning sessions are hands-on and high impact

You will have direct access to engineering leadership and decision-makers

The infrastructure you’re operating is mission-critical and evolving quickly

If you value deep technical collaboration, tight feedback loops, and being at the center of high-scale ML systems — this environment is built for that.

What You’ll Own
  • Production reliability for ML and real-time analytics workloads
  • CI/CD strategy, deployment automation, and rollback design
  • Observability frameworks (SLOs, alerting, monitoring, incident response)
  • Infrastructure-as-Code and Kubernetes environments
  • Capacity planning and performance optimization
  • Post-incident reviews that drive measurable, long-term reliability improvements
  • Reliability standards across teams — not just within a single service
  • You’ll partner directly with engineering and data science teams to ensure ML workloads are production-ready and reliable by design.
What We’re Looking For
  • Deep experience operating Linux infrastructure and networking in production environments
  • Proven impact as a Staff SRE, Senior SRE, or senior-level DevOps/Platform Engineer supporting distributed systems
  • Experience supporting complex, data-intensive or ML-driven systems in production
  • Strong hands-on experience with Docker and Kubernetes
  • Strong scripting ability (Bash and/or Python)
  • CI/CD ownership experience (GitHub Actions, ArgoCD, or similar)
  • Experience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry)
  • Ability to debug systemic failures across infrastructure, deployments, and workloads
  • Clear communicator who works effectively across engineering and data teams
  • Engineers who have evolved from infrastructure foundations into strategic reliability leaders will thrive here.
These Skills Are a Plus
  • Experience operating ML platforms at scale (training + inference)
  • AWS or cloud-managed services experience
  • Exposure to data platforms such as Spark, Airflow, or Kafka
  • Experience in SOC 2 or regulated environments
Why This Opportunity
  • Staff-level ownership of mission-critical ML infrastructure
  • Direct influence over reliability standards across engineering
  • High-visibility role with architectural impact
  • Collaborative engineering culture designed for speed and depth

If you're a Staff-level reliability engineer who wants real ownership and architectural influence — let’s start the conversation.

StratITech is partnering with our San Francisco client to build the next generation of high-scale ML infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff SRE: ML & Real-Time Analytics Reliability Leader
Staff SRE: ML & Real-Time Analytics Reliability Leader

Stratitech Services LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Wand AI • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • Dunwoody (GA)

On-site
USD 200,000 - 225,000
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage