Senior Data Platform SRE: Cloud, Databricks & Observability

Datavant

Jefferson City (MO)

On-site

USD 168,000 - 200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Datavant is seeking a Senior Site Reliability Engineer to join our Data & ML Platform team. You will build a resilient, observable data platform supporting ML workloads across a hybrid cloud environment, partnering with data scientists and engineers to deliver secure, scalable solutions.

You will own platform reliability, implement Datadog-based observability, and drive CI/CD for data pipelines and ML components using GitHub Actions and Terraform.

Qualifications

  • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
  • AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
  • Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.
  • Proven expertise with observability tools (Datadog) and architecting platform-wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, infrastructure-as-code (Terraform), and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault-tolerant systems.
  • Excellent communication and collaboration skills; able to work effectively across teams.

Responsibilities

  • Operate and improve Databricks and Snowflake platforms lifecycle including automation, workspace governance, job orchestration, and cost optimization.
  • Design for reliability: architect resilient, scalable, and secure infrastructure across cloud environments.
  • Advance observability: build platform-wide monitoring, alerting, and logging infrastructure with Datadog and other tools; define SLOs/SLAs.
  • Drive CI/CD for Data & ML: automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
  • Enable data flow across platforms: build patterns and tooling for inter- and intra-cloud data movement (Snowflake, S3, Delta Lake, Kafka).
  • Champion event-driven architectures: use EventBridge, SNS/SQS, and Lambda to build scalable data systems.
  • Collaborate across teams: act as SRE and platform partner for analytics, data science, and product use cases.
  • Contribute to strategy: influence data platform architecture, ML enablement, and data product strategy.

Skills

SRE experience
Cloud infrastructure
Databricks experience
Datadog observability
GitHub Actions
Terraform IaC
Python scripting
Shell scripting
Communication skills

Tools

Databricks
Snowflake
AWS
Terraform
GitHub Actions
Datadog
Lambda
EventBridge

Job description

Datavant is seeking a Senior Site Reliability Engineer to join our Data & ML Platform team. You will build a resilient, observable data platform supporting ML workloads across a hybrid cloud environment, partnering with data scientists and engineers to deliver secure, scalable solutions.

You will own platform reliability, implement Datadog-based observability, and drive CI/CD for data pipelines and ML components using GitHub Actions and Terraform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - Data & ML Platform, Cloud Reliability
Senior SRE - Data & ML Platform, Cloud Reliability

Datavant • United States

On-site
USD 168,000 - 200,000
Senior SRE: Data & ML Platform, Cloud & Observability
Senior SRE: Data & ML Platform, Cloud & Observability

Datavant • Indianapolis (IN)

Hybrid
USD 168,000 - 200,000
Senior SRE: Data & ML Platform, Scalable Cloud
Senior SRE: Data & ML Platform, Scalable Cloud

Datavant • Baton Rouge (LA)

On-site
USD 168,000 - 200,000
Senior Data Platform SRE — Scale & Observability
Senior Data Platform SRE — Scale & Observability

Datavant • Tallahassee (FL)

On-site
USD 168,000 - 200,000
Senior SRE – Data & ML Platform
Senior SRE – Data & ML Platform

Datavant • Washington

On-site
USD 168,000 - 200,000
Senior SRE — Data & ML Platform Leader
Senior SRE — Data & ML Platform Leader

Datavant • Springfield (IL)

On-site
USD 168,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • United States

On-site
USD 168,000 - 200,000
Senior Databricks AI Platform SRE
Senior Databricks AI Platform SRE

Central Business Solutions, Inc • Alpharetta (GA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer II — Scale, Automate & Observe
Site Reliability Engineer II — Scale, Automate & Observe

Worky • Denver (CO)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
401k matching
Employee Stock Purchase Plan
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Datavant • Springfield (IL)

On-site
USD 168,000 - 200,000