Turn this role into an interview — a resume and cover letter built around what this employer wants.
StratITech is hiring for a Staff-level reliability role focused on building and governing the next generation of high-scale ML infrastructure in San Francisco. You will own production reliability for ML and real-time analytics, drive CI/CD strategy, and set robust observability standards across teams.
The role requires hands-on leadership in Kubernetes, IaC, and incident response, with direct access to engineering leadership.
You will define how mission-critical machine learning and real-time analytics systems operate in production — influencing reliability strategy, deployment standards, and infrastructure architecture across engineering.
This team operates in a highly collaborative, in-person engineering environment in SOMA. Infrastructure, ML, and engineering leaders work side by side to design, build, and operate complex systems in real time. The pace is fast, the feedback loops are tight, and decisions happen quickly.
If you’ve grown from Linux systems → DevOps → Staff-level SRE , and you now think in terms of systemic risk, scalability, and long-term reliability strategy — this role gives you direct influence and visibility.
This role is intentionally in-person because:
You will have direct access to engineering leadership and decision-makers
The infrastructure you’re operating is mission-critical and evolving quickly
If you value deep technical collaboration, tight feedback loops, and being at the center of high-scale ML systems — this environment is built for that.
If you're a Staff-level reliability engineer who wants real ownership and architectural influence — let’s start the conversation.
StratITech is partnering with our San Francisco client to build the next generation of high-scale ML infrastructure.