Senior Site Reliability Engineer

Datavant Corporation

Northern (KY)

On-site

USD 168,000 - 200,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Datavant is seeking a Senior Site Reliability Engineer to join our Data & ML Platform team. You will help build and operate a resilient, observable, and scalable platform enabling mission‑critical data and ML workloads across the organization.

You will design reliable, cloud‑native infrastructure on AWS, advance platform observability with Datadog, and drive CI/CD for data pipelines and ML workloads using GitHub Actions and Terraform.

Qualifications

  • 6+ years in SRE, platform engineering, or DevOps roles supporting data‑intensive or ML‑powered applications.
  • AI‑native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands‑on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Snowflake experience is a plus.
  • Deep understanding of cloud‑native infrastructure on AWS (or similar), including VPCs, IAM, event‑driven patterns, and serverless compute.
  • Proven expertise with observability tools (especially Datadog) and architecting platform‑wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, Terraform, and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault‑tolerant systems.
  • Excellent communication and collaboration skills; able to work effectively across teams.

Responsibilities

  • Operate and improve Databricks and Snowflake platforms lifecycle.
  • Design for reliability: architect resilient, scalable, and secure infrastructure across cloud environments.
  • Advance observability: build platform‑wide monitoring, alerting, and logging infrastructure; define SLOs/SLAs.
  • Drive CI/CD for data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
  • Enable data flow across platforms: inter‑ and intra‑cloud data movement across Snowflake, S3, Delta Lake, and Kafka.
  • Champion event‑driven architectures: use EventBridge, SNS/SQS, and Lambda for scalable data systems.
  • Collaborate across teams: be the SRE and platform partner for analytics, data science, and product teams.
  • Contribute to strategy: influence engineering decisions on data platform architecture and ML enablement.

Skills

SRE experience
Platform engineering
CI/CD
Python scripting
Shell scripting
Communication
Cross-team collaboration
Observability

Tools

Databricks
Snowflake
AWS
GitHub Actions
Terraform
Datadog
EventBridge
Lambda

Job description

Senior Site Reliability Engineer

7531 Remote - United States Full-time regular

Datavant is the data collaboration platform trusted for healthcare. Guided by our mission to make the world’s health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies. From fulfilling a single patient’s request for their medical records to powering the AI revolution in healthcare, Datavanters are building the future of how data is connected and used to improve health. By joining Datavant today, you’re stepping onto a driven and highly collaborative team that is passionate about creating transformative change in healthcare.

What We’re Looking For

We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission‑critical data and ML workloads across our organization.

What You Will Do
  • Operate and Improve Databricks and Snowflake: Own Databricks & Snowflake platforms lifecycle—including automation, workspace governance, job orchestration, and cost optimization.
  • Design for Reliability: Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.
  • Advance Observability: Build and maintain platform‑wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.
  • Drive CI/CD for Data & ML: Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
  • Enable Data Flow Across Platforms: Build patterns and tooling to support inter‑ and intra‑cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
  • Champion Event‑Driven Architectures: Leverage cloud‑native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
  • Collaborate Across Teams: Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.
  • Contribute to Strategy: Influence engineering‑wide decisions on data platform architecture, ML enablement, and data product strategy.
What You Need to Succeed
  • 6+ years in SRE, platform engineering, or DevOps roles supporting data‑intensive or ML‑powered applications.
  • AI‑native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands‑on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
  • Deep understanding of cloud‑native infrastructure on AWS (or similar), including VPCs, IAM, event‑driven patterns, and serverless compute.
  • Proven expertise with observability tools (especially Datadog) and architecting platform‑wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, infrastructure‑as‑code (Terraform), and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault‑tolerant systems.
  • Excellent communication and collaboration skills; able to work effectively across teams.
What Helps You Stand Out
  • DevSecOps mindset: Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.
  • Experience with ML infrastructure tooling such as MLflow, Feature Stores, and GPU workload orchestration.
  • Strong experience in both Databricks and Snowflake in a large scale production lakehouse with cross‑warehouse interoperability, e.g. Iceberg v3, Glue, etc.
  • Background in compliance‑aware architecture (e.g., HIPAA, SOC 2) or regulated industries.
  • Familiarity with multi‑cloud or hybrid cloud data environments; experience with Azure.
  • Contributions to open‑source infrastructure, SRE, or observability tools.

We are committed to building a diverse team of Datavanters who are all responsible for stewarding a high‑performance culture in which all Datavanters belong and thrive. We are proud to be an Equal Employment Opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, disability, veteran status, or other legally protected status.

We are committed to working with and providing reasonable accommodations to individuals with physical and mental disabilities.

At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.

The estimated total cash compensation range for this role is: $168,000 - $200,000 USD

To ensure the safety of patients and staff, many of our clients require post-offer health screenings and proof and/or completion of various vaccinations such as the flu shot, Tdap, COVID-19, etc. Any requests to be exempted from these requirements will be reviewed by Datavant Human Resources and determined on a case‑by‑case basis. Depending on the state in which you will be working, exemptions may be available on the basis of disability, medical contraindications to the vaccine or any of its components, pregnancy or pregnancy‑related medical conditions, and/or religion.

This job is not eligible for employment sponsorship.

Datavant is committed to a work environment free from job discrimination. We are proud to be an Equal Employment Opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, disability, veteran status, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • United States

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant2 • United States

Hybrid
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Columbus (OH)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Richmond (VA)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Frankfort (KY)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Hartford (CT)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Jackson (MS)

Hybrid
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Nashville (TN)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Annapolis (MD)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • San Juan (PR)

On-site
USD 168,000 - 200,000