Get more replies from employers
Send a job-specific resume in minutes.
Claritas Rx is seeking an experienced Sr. Manager of Site Reliability Engineering to lead our SRE/DevOps function.
You will own the reliability, availability, and scalability of our AWS-hosted SaaS platform, partnering with Software Engineering, Product Management, and Security to ensure regulatory compliance and operational excellence. You will lead a blended team of onshore/offshore engineers, drive incident response, and shape the architectural direction to support growth and reliability
Claritas Rx brings clarity to the challenges of specialty biopharmaceutical products in the marketplace. In today's highly complex specialty networks, our mission is to illuminate the patient experience beyond the clinical trial. Claritas Rx leverages a proprietary technology platform and deep manufacturer expertise to automate and integrate channel, commercial, and clinical data and help biopharmaceutical companies generate actionable business insights. Our work uncovers the real-world variables impacting patient access, duration of therapy, and other metrics key to commercial success, making a real impact on patient healthcare.
We are seeking an experienced and hands‑on Sr. Manager of Site Reliability Engineering to lead our SRE and DevOps function. Reporting to the Sr. Director of Software Engineering, you will be responsible for the reliability, availability, scalability, and operational excellence of our AWS-hosted SaaS platform—a system that handles sensitive patient and commercial data for some of the world's leading biopharmaceutical companies. You will lead a small, high‑impact team of full‑time SRE/DevOps engineers and offshore contractors, setting technical direction, establishing operational standards, and rolling up your sleeves to solve hard infrastructure and reliability problems alongside your team. This is a player‑coach role: you are expected to lead by example, contribute directly to infrastructure design and tooling, and grow a team that the broader engineering organization depends on. You will partner closely with Software Engineering, Product Management, and Security to ensure our platform operates with the reliability and compliance posture our customers and regulators require.
Own the reliability, availability, performance, and scalability of Claritas Rx's AWS-hosted SaaS platform, with accountability for SLA/SLO commitments made to customers. Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all production services; use error budgets to drive engineering prioritization conversations. Lead incident response and on‑call operations: triage, coordinate resolution, communicate to stakeholders, and conduct thorough post‑incident reviews with actionable corrective actions. Drive a proactive reliability culture — identifying risks before they become incidents through load testing, chaos engineering, and systematic failure mode analysis.
Architect, build, and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform, or equivalent), ensuring environments are reproducible, auditable, and secure. Oversee and continuously improve CI/CD pipelines (GitHub Actions) to enable rapid, safe, and consistent delivery of application and infrastructure changes across environments. Manage and optimize core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront. Ensure robust observability across the stack — centralizing logs, metrics, traces, and alerts using CloudWatch, Sentry, and related tooling — so the team can detect and respond to issues quickly. Manage platform capacity planning, cost optimization, and cloud spend governance.
Ensure all infrastructure design and operational practices meet HIPAA, SOC 2, and HITRUST requirements, given the PHI our platform processes. Partner with the Security function on vulnerability management, infrastructure hardening, secrets management, and access control. Maintain and regularly test disaster recovery (DR) and business continuity plans, including defined RTO/RPO targets for all production systems. Support audit readiness and evidence collection for compliance certifications.
Lead, mentor, and grow a blended team of full‑time SREs/DevOps engineers and offshore contractors, fostering a culture of ownership, continuous improvement, and operational excellence. Manage distributed team dynamics effectively — establishing clear communication rhythms, documentation standards, and hand‑off protocols to ensure offshore resources are productive and well-integrated. Conduct regular 1:1s, set clear goals and development plans for direct reports, and advocate for your team's growth and recognition. Build and maintain a healthy on‑call rotation with appropriate tooling, runbooks, and escalation paths to protect team sustainability.
Collaborate with Software Engineering teams to embed reliability practices into the SDLC — including production readiness reviews, deployment standards, and shared observability tooling. Work with Product Management and Engineering leadership to balance feature delivery velocity against operational risk and technical debt. Contribute to architecture decisions across the platform, provi