Staff Software Engineer, Quality & Reliability Platform

BuildOps, Inc.

San Francisco (CA)

On-site

USD 230,000 - 320,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

BuildOps is seeking a Staff Software Engineer to set and drive the company-wide technical strategy for building, shipping, and operating reliable software. This is a high-impact, cross-functional role spanning engineering, product, infrastructure, and customer-facing teams to identify systemic risks and establish architectural standards for safe delivery by default.

You will influence design, validation, observability, resilience, and engineering practices while mentoring engineers and leading

Qualifications

  • Significant software engineering experience at Staff level or equivalent scope.
  • Proven track record leading multi-team initiatives to improve production reliability and delivery.
  • Strong systems thinking connecting architecture, data integrity, and operational behavior.
  • Experience designing and operating distributed systems in AWS or similar cloud environments.
  • Strong programming skills in TypeScript, Java, or other relevant languages.
  • Experience with observability, resilience engineering, CI/CD, and release safety.
  • Ability to define useful engineering measures that improve outcomes.
  • Proven success influencing architecture and practices across teams.
  • Excellent written and verbal communication with technical stakeholders.
  • Practical approach balancing long-term goals with immediate reliability.

Responsibilities

  • Define and drive BuildOps' technical strategy for engineering quality, production reliability, and safe software delivery.
  • Identify systemic sources of customer-impacting failures and lead cross-team initiatives that address root causes.
  • Establish architectural principles, engineering standards, and paved roads for reliable design and safe delivery by default.
  • Partner with engineering teams during system and product design to improve resilience, operability, testability, and failure isolation before implementation begins.
  • Build or guide development of shared platform capabilities for release safety, automated validation, production feedback, environment management, and developer self-service.
  • Advance BuildOps' observability strategy so teams can understand system behavior, detect regressions, diagnose failures, and invest in reliability.
  • Improve validation of interactions across services, data boundaries, financial workflows, and other business-critical systems.
  • Define meaningful measures of quality and reliability, then use them to identify priorities and demonstrate improvements in outcomes.
  • Lead technical programs spanning multiple teams, aligning stakeholders and driving decisions without direct authority.
  • Mentor engineers and technical leaders, raising the organization’s ability to reason about risk, reliability, and quality.

Skills

Staff scope
System design
Distributed systems
AWS
TypeScript/Java
Observability
CI/CD
Release safety
Incident learning
Leadership across teams

Job description

BuildOps is looking for a Staff Software Engineer to set and drive our company-wide technical strategy for building, shipping, and operating reliable software.

This is a high-impact, cross-functional role for an engineer who sees quality and reliability as properties of the entire system-not as a final testing phase or the responsibility of a separate team. You will work across engineering, product, infrastructure, and customer-facing organizations to identify systemic risks, establish architectural standards, and create platform capabilities that make safe, dependable delivery the default.

You will influence how teams design systems, validate changes, observe production behavior, manage risk, and learn from failures. Your work will span architecture, developer tooling, release safety, observability, resilience, testability, and engineering practices. The specific solutions will evolve as you identify the highest-leverage opportunities.

Product engineering teams at BuildOps own the quality and testing of their features. Your role is not to test their work for them. You will establish the technical direction, paved roads, systems, and standards that enable every team to deliver reliable software with greater confidence and speed.

What You Will Do
  • Define and drive BuildOps' technical strategy for engineering quality, production reliability, and safe software delivery.
  • Identify systemic sources of customer-impacting failures and lead cross-team initiatives that address root causes rather than individual symptoms.
  • Establish architectural principles, engineering standards, and paved roads that make reliable system design and safe delivery easier by default.
  • Partner with engineering teams during system and product design to improve resilience, operability, testability, and failure isolation before implementation begins.
  • Build or guide the development of shared platform capabilities for release safety, automated validation, production feedback, test data, environment management, and developer self-service.
  • Advance BuildOps' observability strategy so teams can understand system behavior, detect regressions quickly, diagnose failures, and make informed reliability investments.
  • Improve how we validate interactions across services, data boundaries, financial workflows, and other business-critical systems.
  • Define meaningful measures of quality and reliability, then use them to identify priorities and demonstrate improvements in customer and engineering outcomes.
  • Lead technical programs that span multiple teams and organizations, aligning stakeholders and driving decisions without relying on direct authority.
  • Mentor engineers and technical leaders, raising the organization's ability to reason about risk, reliability, and quality throughout the software lifecycle.
  • Evaluate BuildOps' existing practices and technology objectively, evolving or replacing them when they no longer meet our needs.
What Success Looks Like
  • Fewer customer-impacting defects and recurring classes of production failures.
  • Greater release confidence and a lower change-failure rate.
  • Faster detection, diagnosis, and recovery when failures occur.
  • Shorter, more reliable feedback loops for engineers making changes.
  • Clearer ownership and better visibility into the health of critical systems and workflows.
  • Increased engineering velocity without sacrificing safety or reliability.
  • Broad adoption of shared practices and platform capabilities without creating a centralized quality bottleneck.
What We Look For
  • Significant software engineering experience, including operating at Staff or equivalent scope on ambiguous, cross-cutting technical problems.
  • A track record of leading multi-team initiatives that improved production reliability, software delivery, platform capabilities, or engineering effectiveness.
  • Strong systems thinking and the ability to connect architecture, data integrity, operational behavior, developer workflows, and customer impact.
  • Experience designing and operating distributed systems in a cloud environment such as AWS.
  • Strong software design and programming skills in TypeScript, Java, or another relevant language.
  • Experience with several of the following: observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning.
  • The ability to define useful engineering measures while avoiding metrics that reward activity without improving outcomes.
  • Demonstrated success influencing architecture and engineering practices across teams that do not report to you.
  • Strong written and verbal communication, including the ability to explain technical risks, tradeoffs, and strategy to engineering, product, and business stakeholders.
  • A practical approach that balances long-term d
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Platform Engineer - Quality, Reliability & Safe Delivery
Staff Platform Engineer - Quality, Reliability & Safe Delivery

BuildOps, Inc. • San Francisco (CA)

On-site
USD 230,000 - 320,000
Director Production Engineering
Director Production Engineering

Vista Applied Solutions Group Inc • Durham (NC)

On-site
USD 130,000 - 160,000
Staff Software Engineer - Platform Architect (Full-Stack)
Staff Software Engineer - Platform Architect (Full-Stack)

MetaProp • San Francisco (CA)

Hybrid
USD 197,000 - 262,000
Generous equity grant
Hybrid work schedule
Lunch provided for in-office days
+1
Staff Software Engineer
Staff Software Engineer

BuildOps, Inc. • San Francisco (CA)

On-site
USD 197,000 - 262,000
Generous equity grant
Hybrid work schedule
Lunch provided for in-office days
+1
Staff Software Engineer
Staff Software Engineer

Eton Solution • Bellevue (WA)

On-site
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Synergy • Chicago (IL)

On-site
USD 100,000 - 150,000
Quality Engineer
Quality Engineer

Better Impact • Winston-Salem (NC)

On-site
USD 120,000 - 160,000
Medical, dental & vision
401(k)
Unlimited vacation + 9 holidays
+2
Staff Quality Engineer
Staff Quality Engineer

VoltForce • Oakland (CA)

On-site
USD 140,000 - 210,000
Platform Engineer - Site Reliability
Platform Engineer - Site Reliability

Auto Hauler Exchange • Rochester (MI)

On-site
USD 120,000 - 160,000
Director of Engineering – Quality & Reliability
Director of Engineering – Quality & Reliability

Jobtailor • New York (NY)

On-site
USD 180,000 - 260,000