Site Reliability Engineer III

Openkrill

United States

Remote

USD 140,000 - 190,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Vida is seeking a remote Site Reliability Engineer to join the Enablement Team. You will modernize and scale our infrastructure across Terraform, Kubernetes, and cloud services, while shaping SRE practices for a growing enterprise client base.

You will own automation, monitoring, and onboarding for production systems, collaborating with multiple teams and external partners to maintain security and reliability.

Qualifications

  • Bachelor's degree in a related field is required.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering with production ownership.
  • Deep hands-on Terraform experience across environments.
  • Strong knowledge of GCP including GKE, Cloud SQL, IAM, networking, and cost management.
  • Production Kubernetes experience including autoscaling and resource management.
  • Hands-on experience building monitoring, alerting, and dashboards (Datadog/Cloud Monitoring).
  • Proficiency in Python for tooling and automation.
  • Able to explain infrastructure decisions to non-specialists.

Responsibilities

  • Consolidate Terraform patterns into a clean, documented structure.
  • Normalize environments and automate builds with GitHub Actions; add drift detection and alerting.
  • Apply patches and upgrades to Cloud SQL and runtimes.
  • Right-size compute and DB workloads for growth and multi-cluster considerations.
  • Evaluate Kubernetes architecture and potential multi-cluster setups.
  • Improve observability with Datadog and Cloud Monitoring; ensure early issue detection.
  • Design observability access for contractors while protecting PHI.
  • Retire legacy infrastructure and create repeatable runbooks and on-call procedures.
  • Prepare infrastructure readiness for enterprise launches in January.
  • Additional responsibilities as needed.

Skills

Terraform
Kubernetes
GCP
Python
Datadog
CI/CD
GitHub Actions
Django

Education

Bachelor's degree

Tools

Cloud Monitoring
Airflow

Job description

ABOUT US

At Vida, we help people get better- and we're helping the healthcare system get better, too.

Vida is a virtual, personalized obesity care provider that uses evidence-based treatment to help patients manage obesity and related conditions like diabetes, high blood pressure, anxiety and depression. Vida's team of Obesity Medicine-Certified Physicians, Registered Dietitians, Expert Coaches and Licensed Therapists takes a whole-person approach to care, helping people lose weight, reduce stress and improve their overall health.

By combining advanced technology with top-notch healthcare providers, Vida is breaking down the barriers that have historically kept people from getting the best care. It's trusted by Fortune 100 companies, major national payers and large providers to enable their employees to live their healthiest lives.

Vida has been operating and growing for years, and our infrastructure reflects that. We run on GCP with a production GKE cluster hosting around 50 workloads, from Django applications to scheduled Airflow jobs. Our data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Our infrastructure is defined in two Terraform repositories, one for core GCP infrastructure and one for our data platform, and both have grown across many contributors over time. Until now, our infrastructure has been managed by backend engineers with deep infrastructure experience, and this role adds our first dedicated SRE to that group.

You’ll be Vida’s first dedicated Site Reliability Engineer. You’ll join the Enablement Team, which owns the platform and tooling our Engineering Teams build on. You’ll report to the Engineering Manager and work closely with the team’s Lead Engineer, who sets technical direction and will mentor you. This is a fully remote role with no time zone restrictions.

You’ll modernize, consolidate, and scale our infrastructure as Vida takes on a wave of new enterprise contracts starting January 1. You’ll also help shape what SRE looks like at Vida going forward.

Repsonsibilities:
  • Consolidate our Terraform, which has grown into inconsistent patterns across our infrastructure and data repositories, into a clean, well-documented structure the whole team can work in. Establish conventions for state management, module structure, code review, and CI checks.
  • Normalize environments, improve build and deploy automation in GitHub Actions, and add drift detection and alerting.
  • Apply overdue patches and upgrades across our Cloud SQL databases and application runtimes.
  • Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services.
  • Evaluate our Kubernetes architecture as we grow, including whether and when to move to a multi-cluster setup.
  • Improve monitoring and observability in Datadog and Cloud Monitoring so we catch issues before they become incidents.
  • Design observability access for contractors and external partners that gives them the visibility they need while keeping protected health information out of view.
  • Retire legacy infrastructure and tooling that has been replaced but not yet decommissioned.
  • Build repeatable operational processes, including runbooks, an on-call rotation, and escalation documentation.
  • Support infrastructure readiness for Vida's January 1 enterprise launches.
  • Additional responsibilities as needed.
Qualifications:
  • Bachelor's degree at a minimum.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
  • Deep hands-on Terraform experience, including structuring modules and managing state across environments.
  • Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
  • Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
  • Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
  • Proficiency in Python for tooling and automation.
  • Comfortable working across multiple teams and disciplines, and explaining infrastructure decisions to non-specialists.
Preferred:
  • Experience as an early or first SRE hire.
  • Experience refactoring or consolidating a large, organically grown Terraform codebase.
  • Experience improving observability from a less mature baseline.
  • Experience in a HIPAA-regulated or other compliance-driven environment.
  • CI/CD experience with GitHub Actions.
  • Experience running Django applications or Airflow in production on Kubernetes.
  • Experience designing or migrating to multi-cluster Kubernetes architectures.

Vida is proud to be an Equal Employment Opportunity and Affiliation Action employer.

Diversity is more than a commitment at Vida—it is the foundation of what we do. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, gender, gender identity or expression, sexual orientation, marital status, national origin, genetics, disability, age, or Veteran status. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law.

We seek to recruit, develop and retain the most talented people from a diverse candidate pool. We don’t just accept differences — we celebrate them, we support them, and we thrive on them for the benefit of our employees, our platform and those we serve. Vida is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.

We do not accept unsolicited assistance from any headhunters or recruitment firms for any of our job openings. All resumes or profiles submitted by search firms to any employee at Vida in any form without a valid, signed search agreement in place for the specific position will be deemed the sole property of Vida. No fee will be paid in the event the candidate is hired by Vida as a result of the unsolicited referral.

**Vida is authorized to do business in many, but not all, states. If you are not located in or able to work from a state where Vida is registered, you will not be eligible for employment. Please speak with your recruiter to learn more about where Vida is registered.

Please note: Applicants must be authorized to work in the U.S. as Vida is unable to sponsor work visas for any position.

All Vida Employees must reside in/be able to work from the U.S.- international work is prohibited. Job postings at Vida will remain open through end of year, until filled.

#LI-remote

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III
Site Reliability Engineer III

Vida-Health • United States

On-site
USD 175,000 - 185,000
Site Reliability Engineer III
Site Reliability Engineer III

Vida Health • New York (NY)

On-site
USD 175,000 - 185,000
Senior Director, Engineering Operations & Security
Senior Director, Engineering Operations & Security

Vida-Health • United States

Remote
USD 240,000 - 250,000
Senior Data Engineer- Data Platform
Senior Data Engineer- Data Platform

Vida • United States

Remote
USD 120,000 - 180,000
Backend Software Engineer II- Data Platform
Backend Software Engineer II- Data Platform

Openkrill • United States

Remote
USD 90,000 - 130,000
Senior Data Engineer- Data Platform
Senior Data Engineer- Data Platform

Openkrill • United States

Remote
USD 120,000 - 190,000
Backend Software Engineer II- Data Platform
Backend Software Engineer II- Data Platform

Vida • United States

Remote
USD 110,000 - 150,000
Backend Software Engineer II- Data Platform
Backend Software Engineer II- Data Platform

Vida Health • United States

On-site
USD 120,000 - 130,000
Site Reliability Engineer III (Fully Remote)
Site Reliability Engineer III (Fully Remote)

Vida • United States

Remote
USD 130,000 - 180,000
Principal AI/ML Engineering Lead
Principal AI/ML Engineering Lead

Vida-Health • United States

Remote
USD 250,000 - 275,000