Senior SRE - Data & ML Platform: Scale & Resilience

Datavant

Phoenix (AZ)

On-site

USD 168,000 - 200,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Datavant is seeking a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll build and operate a resilient, observable, and scalable platform for data and ML workloads across the organization.

This role suits someone with a strong SRE mindset, cloud infrastructure experience, and data platform expertise. You’ll work with Data Scientists, Analysts, and App Engineering to keep systems secure, self-service, and production-grade.

Qualifications

  • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive apps.
  • AI-native working style with tools like Claude Code, Cursor, Copilot.
  • Hands-on Databricks experience, including clusters and CI/CD integration.
  • Deep understanding of cloud-native infra on AWS or similar.
  • Observability expertise and platform-wide logging/monitoring design.
  • Strong CI/CD with GitHub Actions and Terraform for data systems.
  • Shell scripting and Python proficiency.
  • Experience building highly available, fault-tolerant systems.

Responsibilities

  • Operate and improve Databricks and Snowflake platforms lifecycle, including governance and cost optimization.
  • Design for reliability across cloud environments with failover and autoscaling.
  • Advance observability with monitoring, alerting, and logging; define SLOs/SLAs.
  • Drive CI/CD for data and ML workflows using GitHub Actions and IaC tooling.
  • Enable data flow across platforms including Snowflake, S3, Delta Lake, Kafka.
  • Champion event-driven architectures using EventBridge, SNS/SQS, and Lambda.
  • Collaborate across teams as platform partner for analytics, data science, and product use cases.
  • Contribute to strategy on data platform architecture and ML enablement.

Skills

SRE mindset
Cloud infrastructure
Observability
CI/CD tooling
Python scripting
Shell scripting
Communication

Tools

Databricks
Snowflake
AWS
GitHub Actions
Terraform
Datadog
EventBridge
Lambda
S3
Kafka

Job description

Datavant is seeking a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll build and operate a resilient, observable, and scalable platform for data and ML workloads across the organization.

This role suits someone with a strong SRE mindset, cloud infrastructure experience, and data platform expertise. You’ll work with Data Scientists, Analysts, and App Engineering to keep systems secure, self-service, and production-grade.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE - Data & ML Platform, Cloud Reliability
Senior SRE - Data & ML Platform, Cloud Reliability

Datavant • United States

On-site
USD 168,000 - 200,000
Senior Data & ML Platform SRE
Senior Data & ML Platform SRE

Datavant2 • United States

Hybrid
USD 168,000 - 200,000
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Senior Platform SRE: Scale & Reliability
Senior Platform SRE: Scale & Reliability

United States Digital Space LLC • United States

Hybrid
USD 103,000 - 162,000
Health insurance
Vacation and RTT
Mental health and coaching
+7
Senior SRE: Scale Reliability Leader (Hybrid)
Senior SRE: Scale Reliability Leader (Hybrid)

Plenful • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Healthcare Coverage
401(k) with Company Match
Equity
+5
Senior SRE: AI & Data Platform—Remote (ET)
Senior SRE: AI & Data Platform—Remote (ET)

Systematix group • Maryland

On-site
USD 150,000 - 210,000
Remote contract
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE: Scale Reliability for Health AI Platform
Senior SRE: Scale Reliability for Health AI Platform

RXinsider LTD. • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Healthcare Coverage
401(k) Match
Equity
+5
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Staff Data Platform SRE — Scale, Resilience & AI Infra
Staff Data Platform SRE — Scale, Resilience & AI Infra

Lightspeed • United States

Hybrid
USD 150,000 - 190,000
Flexible paid time off
Equity options
Pension contributions
+3