Director - Application Site-Reliability Engineering

Scorpion Therapeutics

Irving (TX)

On-site

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Scorpion Therapeutics in Irving, TX seeks an experienced Site/Platform Reliability Engineer to own the production-support model across the clinical application portfolio, including on-call rotation, escalation, and incident-response playbooks.

You will establish and audit access controls, maintain SOX ITGC and FDA-ready evidence, and define SLOs and budgets with a focus on minimum downtime and rapid recovery. The role emphasizes AI-assisted operations and dev-focused automation.

Qualifications

  • BS in CS/SE/IS (or equivalent) + 10+ years SRE/DevOps/platform/production ops.
  • 4+ years people-management/team-lead experience.
  • Hands-on incident response leadership for Tier 1/business-critical systems.
  • Experience defining/implementing SLOs/error budgets and on-call workflows.
  • Experience building/maturing production support/on-call/SRE (runbooks, rotation design).
  • Experience operating under regulated/financial-controls frameworks (e.g., SOX ITGCs, HIPAA, FDA/CAP/CLIA) with audit-capable documentation.
  • Track record applying AI-assisted practice to operations/engineering.

Responsibilities

  • Own the production-support model for the clinical application portfolio (on-call rotation, escalation procedures, incident-response playbooks).
  • Establish/audit production-access and segregation-of-duties controls; maintain SOX ITGC and applicable FDA audit-ready evidence (access grants, role changes, privileged-action logs).
  • Define and manage clinical SLOs, error budgets, and availability targets; track MTTD/MTTR and on-call burden; intervene before error budgets are spent.
  • Lead incident response for high-severity events as incident commander/senior technical responder; coordinate teams, run post-incident reviews, drive remediation to closure.
  • Develop runbooks and approved operational automation; enable safe, documented standard interventions.
  • Drive automation-first operations (reduce toil; favor self-service tooling over ticket requests).
  • Coordinate deployments, verify health, and own rollback decisions.
  • Produce audit-ready change/deployment evidence via CI/CD pipelines.
  • Hire/develop the App-SRE team; set expectations and career growth frameworks.
  • Own application-layer observability (dashboards, alerts, SLO monitors) and represent App-SRE in leadership/compliance forums.
  • Run the function AI-first (AI-assisted runbooks, incident analysis, operational tooling).

Skills

SRE Leadership
Incident Response
DevOps / Platform
Audit & Compliance
AI-assisted Ops

Education

BS in CS/SE/IS

Tools

CI/CD tooling
Observability tools
Cloud platforms

Job description

Job Responsibilities:
  • Own the production-support model for the clinical application portfolio (on-call rotation, escalation procedures, incident-response playbooks).
  • Establish/audit production-access and segregation-of-duties controls; maintain SOX ITGC and applicable FDA audit-ready evidence (access grants, role changes, privileged-action logs).
  • Define and manage clinical SLOs, error budgets, and availability targets; track MTTD/MTTR and on-call burden; intervene before error budgets are spent.
  • Lead incident response for high-severity events as incident commander/senior technical responder; coordinate teams, run post-incident reviews, drive remediation to closure.
  • Develop runbooks and approved operational automation; enable safe, documented standard interventions.
  • Drive automation-first operations (reduce toil; favor self-service tooling over ticket requests).
  • Coordinate deployments, verify health, and own rollback decisions.
  • Produce audit-ready change/deployment evidence via CI/CD pipelines.
  • Hire/develop the App-SRE team; set expectations and career growth frameworks.
  • Own application-layer observability (dashboards, alerts, SLO monitors) and represent App-SRE in leadership/compliance forums.
  • Run the function AI-first (AI-assisted runbooks, incident analysis, operational tooling).
Required Qualifications:
  • BS in CS/SE/IS (or equivalent) + 10+ years SRE/DevOps/platform/production ops.
  • 4+ years people-management/team-lead experience.
  • Hands-on incident response leadership for Tier 1/business-critical systems.
  • Experience defining/implementing SLOs/error budgets and on-call workflows.
  • Experience building/maturing production support/on-call/SRE (runbooks, rotation design).
  • Experience operating under regulated/financial-controls frameworks (e.g., SOX ITGCs, HIPAA, FDA/CAP/CLIA) with audit-capable documentation.
  • Track record applying AI-assisted practice to operations/engineering.
Preferred Qualifications:
  • Clinical diagnostics/lab/molecular pathology/digital health domain experience.
  • Direct SOX ITGC audit or CAP/CLIA inspection support with evidence gathering.
  • Cloud-native observability knowledge (telemetry/tracing/APM).
  • Deployment/release/rollback experience in continuous delivery.
  • Player-coach ability; automation-toil reduction track record; experience presenting strategy/risk to executives.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site-Reliability Engineer, Application Operations
Site-Reliability Engineer, Application Operations

Scorpion Therapeutics • Irving (TX)

On-site
USD 120,000 - 180,000
Director of App SRE & Reliability Operations
Director of App SRE & Reliability Operations

Scorpion Therapeutics • Irving (TX)

On-site
USD 140,000 - 210,000
Director - Application Site-Reliability Engineering
Director - Application Site-Reliability Engineering

Caris Life Sciences • Irving (TX)

Hybrid
USD 180,000 - 240,000
Director of AI‑Driven App-SRE & Reliability
Director of AI‑Driven App-SRE & Reliability

Caris Life Sciences • Irving (TX)

Hybrid
USD 180,000 - 240,000
Director - Application Site-Reliability Engineering
Director - Application Site-Reliability Engineering

Caris MPI, Inc. • Irving (TX)

Hybrid
USD 160,000 - 210,000
Director of App Reliability & AI-Driven Operations
Director of App Reliability & AI-Driven Operations

Caris MPI, Inc. • Irving (TX)

Hybrid
USD 160,000 - 210,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000