Director - Application Site-Reliability Engineering

Caris Life Sciences

Irving (TX)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Caris Life Sciences seeks an experienced Director of Application Reliability and Production Operations to lead the on-site/hybrid team responsible for the clinical software portfolio. You will define SLOs, manage on-call incidents, and drive automation with strong governance under FDA and SOX controls.

You will hire and develop the App-SRE team, coordinate deployments, and ensure audit-ready evidence across change processes. This role emphasizes AI-assisted practice and observability maturity.

Qualifications

  • Bachelor's degree or equivalent practical experience in a technical field.
  • 10+ years in SRE/DevOps/production operations.
  • 4+ years in people-management within SRE/production-operations.
  • Hands-on incident response experience at Tier 1 or critical systems.
  • Experience defining SLOs, error budgets, and alerting in production.
  • Experience building/ maturing production-support and on-call functions.
  • Regulatory framework experience (FDA, SOX, CAP/CLIA, HIPAA) with audit-ready documentation.
  • Experience applying AI-assisted practices to operations or engineering.

Responsibilities

  • Own the production-support model for the clinical application portfolio, including on-call rotation and incident response playbooks.
  • Define, audit, and enforce SLOs, error budgets, and availability targets with leadership; monitor health metrics.
  • Lead incident response as commander or senior responder; coordinate cross-functional teams and drive remediation.
  • Develop runbooks and automation; ensure safe execution within documented procedures.
  • Drive automation-first operations; reduce toil and favor self-service tooling.
  • Coordinate deployments with engineering; validate health and execute rollback decisions.
  • Grow and mentor the App-SRE team; set performance expectations and career paths.
  • Represent App-SRE in leadership forums and compliance audits; communicate risk clearly.
  • Lead AI-first practices for runbook automation and incident analysis.

Skills

SRE leadership
Incident response
SLOs & error budgets
On-call management
Regulatory compliance
AI in operations

Education

Bachelor's degree in CS/Engineering

Tools

CI/CD tooling
Observability platforms
Automation tooling

Job description

At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives. We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting‑edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do. But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose. Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.

Position Summary

Caris Life Sciences is one of the largest precision‑oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double‑digit rates. Behind every case is a matched molecular, imaging, and clinical‑outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it reliable is this role's charter.

Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on‑call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements.

The Director hires and develops the team, sets standards and selects tooling, establishes production‑access governance and segregation‑of‑duties controls with the information‑security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here. Frontier AI coding assistants are standard‑issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation.

Characterization tests and golden‑master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application‑level observability and risk‑based release governance aligned with FDA Computer Software Assurance guidance.

Modernization of established systems is active engineering work, not deferred maintenance. The infrastructure organization owns the platform and observability runtime; this role owns application‑layer reliability on top of it. The operating model is automation‑first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.

Location: Irving, Texas (Dallas–Fort Worth), on‑site/hybrid, minimum three days per week on campus; co‑located with Caris's laboratory and clinical operations.

Job Responsibilities
  • Own the production‑support model for the clinical application portfolio: the on‑call rotation, escalation procedures, and incident‑response playbooks.
  • Establish and audit production‑access and segregation‑of‑duties controls with engineering, information‑security, quality, and infrastructure partners; keep the evidence audit‑ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged‑action logs.
  • Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational‑health metrics such as mean time to detect, mean time to recover, and on‑call burden; intervene while an error budget is burning, not after it is spent.
  • Lead incident response for high‑severity production events as incident commander or senior technical responder; coordinate cross‑functional teams, run post‑incident reviews, and drive systemic remediation to closure.
  • Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.
  • Drive the automation‑first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self‑service tooling over ticket‑driven request work.
  • Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.
  • Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit‑ready by‑product of release.
  • Hire and develop the App‑SRE team; set performance expectations, on‑call responsibilities, and career growth frameworks.
  • Own application‑layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.
  • Represent App‑SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.
  • Run the function AI‑first: make AI‑assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.
Required Qualifications
  • Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.
  • 10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.
  • 4+ years of direct experience in a people‑management or team‑lead role within an SRE or production‑operations function.
  • Hands‑on experience leading incident response for Tier 1 or business‑critical production systems, including serving as incident commander or senior technical responder.
  • Experience defining and implementing SLOs, error budgets, and associated alerting and on‑call workflows in a production environment.
  • Experience building or significantly maturing a production‑support, on‑call, or SRE function, including runbook development and on‑call‑rotation design.
  • Experience operating production systems under a formal regulatory or financial‑controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.
  • A record of applying AI‑assisted practice to operations or engineering work, personally or through a team.
Preferred Qualifications
  • Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.
  • Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access‑control, change‑management, and monitoring controls.
  • Working knowledge of modern cloud‑native observability at the application‑instrumentation layer, including open standards for telemetry and tracing and application‑performance‑monitoring platforms.
  • Experience with deployment pipelines, release‑management workflows, and rollback procedures in a continuous‑delivery environment.
  • Ability to operate as a player‑coach, contributing directly to technical work while building and leading a team.
  • Track record of reducing operational toil through automation programs in an SRE or production‑operations organization.
  • Experience presenting operational strategy and risk posture to senior or executive audiences.
Physical Demands

Ability to sit, stand, and work at a computer for extended periods.

Training

All job‑specific, safety, and compliance training is assigned based on the job functions associated with this employee.

Other

This role serves in the senior escalation tier of the production on‑call rotation, with after‑hours response to high‑severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on‑sites, and leadership reviews.

Conditions of Employment

Individual must successfully complete pre‑employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification. This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.

Caris Life Sciences is an equal opportunity employer.

All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.

Caris Life Sciences is a leading innovator in molecular science and artificial intelligence focused on fulfilling the promise of precision medicine through quality and innovation. Caris is committed to quality and excellence at our state‑of‑the‑art laboratories. Learn more about our tissue lab and the advanced technologies that are helping improve the lives of cancer patients.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director - Application Site-Reliability Engineering
Director - Application Site-Reliability Engineering

Caris MPI, Inc. • Irving (TX)

Hybrid
USD 160,000 - 210,000
Software Engineer
Software Engineer

Caris Life Sciences • Irving (TX), Northern (KY)

On-site
USD 110,000 - 140,000
Associate Software Engineer
Associate Software Engineer

Caris Life Sciences • Irving (TX), Northern (KY)

Hybrid
USD 70,000 - 95,000
Senior Software Engineer - Report Generation
Senior Software Engineer - Report Generation

Caris Life Sciences • Irving (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Product Manager
Senior Product Manager

Caris Life Sciences • Irving (TX), Northern (KY)

On-site
USD 140,000 - 190,000
Product Manager
Product Manager

Caris Life Sciences • Irving (TX)

Hybrid
USD 110,000 - 150,000
Identity Resolution Engineer - MS1 / Core Database
Identity Resolution Engineer - MS1 / Core Database

Caris Life Sciences • Irving (TX)

Hybrid
USD 120,000 - 180,000
VP, Early Detection Informatics & Data Science
VP, Early Detection Informatics & Data Science

Caris Life Sciences • Tempe (AZ)

On-site
USD 150,000 - 200,000
Senior Equipment Engineering Technician
Senior Equipment Engineering Technician

Caris Life Sciences • Phoenix (AZ)

On-site
USD 85,000 - 110,000
VP, Data Science — AI-Driven Precision Medicine Leader
VP, Data Science — AI-Driven Precision Medicine Leader

Caris Life Sciences • Tempe (AZ)

On-site
USD 250,000 - 350,000