The Company
Dexcom Corporation (NASDAQ DXCM) is a pioneer and global leader in continuous glucose monitoring (CGM). Dexcom began as a small company with a big dream: to forever change how diabetes is managed, unlocking information and insights that drive better health outcomes. We are 25 years later, having pioneered an industry, and we’re just getting started. Our vision extends beyond diabetes to empower people to take control of health, providing personalized, actionable insights that solve important health challenges. We continue to improve human health and aspire to become a leading consumer health technology company while developing solutions for serious health conditions.
Meet the Team
Dexcom is seeking a motivated and experienced Senior Site Reliability Engineer to architect, build, and operate the resilient, scalable, and secure cloud infrastructure that powers our R&D Platform serving millions of customers every day. This role is crucial in ensuring rapid, safe, and compliant delivery of life‑changing medical technologies.
As a senior technical leader within the SRE team, you will provide strategic guidance and technical oversight, partnering with engineering, platform, and architecture groups to drive organizational reliability maturity. You will lead initiatives in automation, observability, and incident management while fostering a culture of operational excellence and continuous improvement. This position offers a unique opportunity to shape the future of Dexcom’s evolving cloud reliability strategy in a fast‑paced, collaborative environment.
Where You Come In
- Architect and evolve Dexcom’s observability ecosystem, defining standards for metrics, logging, tracing, and SLO/SLA-driven reliability.
- Design, build, and operate highly available cloud infrastructure on Google Cloud Platform (GCP), focusing on performance, scalability, and security.
- Lead Kubernetes platform operations, improving cluster reliability, multi‑tenant architecture, and deployment patterns.
- Diagnose and resolve complex failures across cloud infrastructure, CI/CD pipelines, policy engines, and microservices.
- Set the direction for Infrastructure as Code (IaC), defining best practices with Terraform, Pulumi, or Crossplane for automated provisioning.
- Drive automation strategy to eliminate toil, build self‑service capabilities, and operationalize guardrails for compliance and cost efficiency.
- Lead major incident response and conduct deep post‑incident reviews to implement remediations that prevent recurring failure categories.
- Mentor engineers and influence cross‑functional practices to help teams adopt operational discipline and cloud‑native best practices.
- Partner with developer teams to optimize capacity strategies and ensure the seamless delivery of high‑quality solutions.
What Makes You Successful
- Problem‑Solver & Innovator: Proven ability to solve complex failures across distributed systems, navigating technical debt to drive long‑term systematic fixes.
- Technical Maestro: Expert‑level knowledge of GCP and deep Kubernetes operational mastery, setting the stage for resilient cloud‑native architectures.
- Visionary Leader: Portfolio highlights well‑orchestrated reliability initiatives, reflecting a flair for turning technical strategy into operational success.
- Observability Strategist: Extensive experience designing metrics pipelines and SLO frameworks that provide actionable insights and reduce MTTR.
- Automation Advocate: Advanced proficiency in Python, Go, or Bash, with a track record of building maintainable tooling that eliminates manual toil.
- Great Communicator: Exceptional skills in articulating complex technical concepts to both engineering peers and stakeholders, ensuring alignment on reliability goals.
- Analytical Architect: Build, test, and maintain infrastructure code that ensures scalable and repeatable platform provisioning. Hands‑on experience with modern declarative ecosystems like Pulumi, Crossplane, or similar tools is highly preferred.
- Collaborative Mentor: Strong ability to influence architectural decisions across teams while motivating and growing junior engineers.
- Compliance Conscious: Experience operating in regulated environments (HIPAA, ISO, or medical device) is a significant plus.
- Agile Mindset: Ability to deal with ambiguity and efficiently change plans to meet evolving business needs in a fast‑paced R&D environment.
Travel Required
Potential for occasional domestic or international travel.
What You’ll Get
- A front‑row seat to life‑changing CGM technology and exposure to our brave #dexcomwarriors community.
- A full and comprehensive benefits program.
- Growth opportunities on a global scale.
- Access to career development through in‑house learning programs and/or qualified tuition reimbursement.
- An exciting and innovative, industry‑leading organization committed to our employees, customers, and the communities we serve.
Experience And Education
- Typically requires a bachelor’s degree in Computer Science or a related field.
- 5–8+ years of experience in SRE, DevOps, or Cloud Engineering, operating mission‑critical production systems.
- Relevant certifications such as CKA, CKAD, or GCP Professional Cloud Engineer are highly preferred.