Senior Platform Reliability Engineer (Databricks)

United States Digital Space LLC

United States

Remote

USD 180,000 - 270,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

shield.ai is seeking a Sr. Staff Platform / Data Reliability Engineer to make the Databricks platform reliable, secure, scalable for enterprise use.

This role focuses on the platform layer above ingestion, governance, and production workloads, enabling multiple domains and regulated data handling. You will partner with data engineers and Cloud & Infra teams to define CI/CD standards, platform policies, and incident response, mentoring engineers and documenting standards to scale beyond the

Qualifications

  • 12+ years of experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure.
  • Hands-on experience with Databricks or a closely related cloud data platform in production environments.
  • Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines.
  • Strong understanding of platform operations concepts such as observability, monitoring, alerting, incident management, and reliability engineering.
  • Experience with compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms.
  • Ability to work effectively in a regulated or security-sensitive environment with strong expectations around access control, auditability, and operational discipline.
  • Strong collaboration skills and comfort partnering with cloud/infrastructure, security, data engineering, and analytics stakeholders.

Responsibilities

  • Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns for data jobs and platform services.
  • Define and maintain CI/CD and promotion standards for Databricks assets, including workflows, jobs, notebooks, code packages, infrastructure configuration, and environment promotion from dev to prod.
  • Design and maintain platform standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and production execution reliability.
  • Establish reusable operational templates and enablement patterns for new domains onboarding to Databricks, including logging conventions, job tagging, metadata capture, and support handoff expectations.
  • Partner with the Senior Data Engineer to ensure ingestion and medallion patterns are implemented in a way that is observable, recoverable, cost-aware, and secure in production.
  • Work with the cloud/infrastructure team to align Databricks configuration and usage patterns with broader enterprise cloud standards, especially where commercial and future government-hosted environments are involved.
  • Help enforce technical controls for data segregation, access boundaries, and operational compliance in a highly regulated environment.
  • Track and improve platform health metrics such as job success rates, incident trends, data pipeline reliability, cost efficiency, and environment drift.
  • Document platform standards, operational expectations, and support models so the Databricks platform can scale beyond a small founding team.
  • Mentor internal engineers who are growing into platform responsibilities, helping expand Databricks operational knowledge within the team.

Skills

Data platform engineering
Observability
Reliability engineering
Security/compliance
Collaboration

Tools

Databricks
Delta Lake
Unity Catalog
Workflows
Terraform

Job description

shield.ai is seeking a Sr. Staff Platform / Data Reliability Engineer to make the Databricks platform reliable, secure, scalable for enterprise use.

This role focuses on the platform layer above ingestion, governance, and production workloads, enabling multiple domains and regulated data handling. You will partner with data engineers and Cloud & Infra teams to define CI/CD standards, platform policies, and incident response, mentoring engineers and documenting standards to scale beyond the

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Databricks Platform Reliability Lead
Senior Databricks Platform Reliability Lead

Shield AI • United States

Remote
USD 180,000 - 270,000
Bonus
Equity
Benefits
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Shield AI • United States

Remote
USD 180,000 - 270,000
Bonus
Equity
Benefits
Data & AI Platform Reliability Leader
Data & AI Platform Reliability Leader

Amerilife Group, LLC • Town of Florida (NY)

Hybrid
USD 160,000 - 177,000
PTO
Medical, dental, vision
Retirement savings
+2
Senior Platform Engineer — Azure Databricks Expert
Senior Platform Engineer — Azure Databricks Expert

Floor & Decor • United States

On-site
USD 140,000 - 190,000
Bonus opportunities
401k with company match
Employee Stock Purchase Plan
+5
Senior Staff Production Engineer, AI-Driven Reliability
Senior Staff Production Engineer, AI-Driven Reliability

Cacheflow • San Francisco (CA)

On-site
USD 228,000 - 315,000
Senior GovCloud & AI Platform Reliability Engineer
Senior GovCloud & AI Platform Reliability Engineer

Databricks Inc. • McLean (VA)

On-site
USD 130,000 - 179,000
Comprehensive benefits
Equity options
Annual performance bonus
Remote Databricks Platform Engineer
Remote Databricks Platform Engineer

Agility Technologies Inc • United States

Remote
USD 140,000 - 210,000
401(k) matching
Dental insurance
Health insurance
+3
Remote Databricks Platform Engineer - CI/CD & Automation
Remote Databricks Platform Engineer - CI/CD & Automation

Fortis Industries, Inc. DBA - LTS, Inc. • Atmore (AL)

On-site
USD 120,000 - 180,000
Senior Data Engineer (Databricks)
Senior Data Engineer (Databricks)

Elitmind Sp. z o.o. • Town of Poland (NY)

On-site
USD 120,000 - 180,000
Senior SRE - Data & ML Platform, Cloud Reliability
Senior SRE - Data & ML Platform, Cloud Reliability

Datavant • United States

On-site
USD 168,000 - 200,000