Senior Site Reliability Engineer
Dexcom is seeking a motivated and experienced Senior Site Reliability Engineer to architect, build, and operate the resilient, scalable, and secure cloud infrastructure powering our R&D Platform serving millions of customers every day. The role is crucial in ensuring rapid, safe, and compliant delivery of life-changing medical technologies.
As a senior technical leader within the SRE team, you will provide strategic guidance and technical oversight, partner with engineering, platform, and architecture groups to drive organizational reliability maturity, and lead initiatives in automation, observability, and incident management while fostering a culture of operational excellence and continuous improvement.
Responsibilities
- Architect and evolve Dexcom’s observability ecosystem, defining standards for metrics, logging, tracing, and SLO/SLA-driven reliability.
- Design, build, and operate highly available cloud infrastructure on Google Cloud Platform (GCP), focusing on performance, scalability, and security.
- Lead Kubernetes platform operations, improving cluster reliability, multi-tenant architecture, and deployment patterns.
- Diagnose and resolve complex failures across cloud infrastructure, CI/CD pipelines, policy engines, and microservices.
- Set the direction for Infrastructure as Code (IaC), defining best practices with Terraform, Pulumi, or Crossplane for automated provisioning.
- Drive automation strategy to eliminate toil, build self-service capabilities, and operationalize guardrails for compliance and cost efficiency.
- Lead major incident response and conduct deep post-incident reviews to implement remediations that prevent recurring failure categories.
- Mentor engineers and influence cross-functional practices to help teams adopt operational discipline and cloud-native best practices.
- Partner with developer teams to optimize capacity strategies and ensure the seamless delivery of high-quality solutions.
Qualifications
- Problem-solver and innovator: Proven ability to solve complex failures across distributed systems, navigating technical debt to drive long-term systematic fixes.
- Technical maestro: Expert-level knowledge of GCP and deep Kubernetes operational mastery.
- Visionary leader with a portfolio of well-orchestrated reliability initiatives.
- Observability strategist: Extensive experience designing metrics pipelines and SLO frameworks that provide actionable insights and reduce MTTR.
- Automation advocate: Advanced proficiency in Python, Go, or Bash, with a track record of building maintainable tooling that eliminates manual toil.
- Great communicator with exceptional skills in articulating complex technical concepts to peers and stakeholders.
- Analytical architect: Hands-on experience with modern declarative ecosystems like Pulumi, Crossplane, or similar tools.
- Collaborative mentor capable of influencing architectural decisions and growing junior engineers.
- Compliance conscious: Experience operating in regulated environments (HIPAA, ISO, or medical device) is a significant plus.
- Agile mindset: Ability to deal with ambiguity and efficiently change plans in a fast-paced R&D environment.
- Bachelor’s degree in Computer Science or a related field.
- 5–8+ years of experience in SRE, DevOps, or Cloud Engineering, operating mission-critical production systems.
- Relevant certifications such as CKA, CKAD, or GCP Professional Cloud Engineer are highly preferred.
Travel Requirement
Potential for occasional domestic or international travel.
Benefits
- A front-row seat to life-changing CGM technology; learn about our brave #dexcomwarriors community.
- A full and comprehensive benefits program.
- Growth opportunities on a global scale.
- Access to career development through in-house learning programs and/or qualified tuition reimbursement.
- An exciting and innovative, industry-leading organization committed to our employees, customers, and the communities we serve.