Staff Software Engineer, Data Warehouse

Jobtailor

California (MO)

On-site

USD 140,000 - 210,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Office 3–4 days per week

Job summary

Jobtailor is seeking an experienced data platform engineer to own and operate our end-to-end data warehouse infrastructure. You will design CDC pipelines, manage data lake architectures, and optimize query and storage costs while sustaining HIPAA and SOC 2 compliance.

You will work with Debezium, Kafka/Redpanda, and StarRocks or similar engines, building dbt models and orchestration with Airflow or Dagster in a cloud environment.

Qualifications

  • 6+ years of software engineering experience building or operating data platforms at scale.
  • Experience with CDC, Debezium or equivalent.
  • Experience with streaming technologies (Kafka/Redpanda).
  • Experience with data lake formats (Iceberg/Delta Lake/Hudi).

Responsibilities

  • Own data warehouse platform end-to-end.
  • Design and operate low-latency CDC pipelines.
  • Architect a data lake on object storage.
  • Build and maintain CI/CD and observability tooling.

Skills

CDC Pipeline Development
SQL Proficiency
Data Platform Architecture
Cost/Latency Trade-offs

Tools

Debezium
Kafka
Redpanda
Iceberg
Delta Lake
Hudi
StarRocks
dbt
Airflow
Dagster

Job description

  • Own Commure’s data warehouse platform end-to-end, including CDC pipelines, data lake, query layer, transformation layer, and analytics-facing tooling
  • Design and operate low-latency, high-fidelity CDC pipelines using Debezium and Kafka, Redpanda, or an equivalent streaming backbone
  • Architect a data lake on object storage using Iceberg, Delta Lake, or Hudi with Parquet
  • Enable both batch and streaming workloads with separation of storage from compute
  • Run and scale StarRocks or adjacent MPP/lakehouse engines as the query and serving layer
  • Design schemas, materialized views, ingestion patterns, and performance/cost trade-offs
  • Build the dbt transformation layer, including modeling standards, tests, documentation, and a semantic layer
  • Establish orchestration with Airflow, Dagster, or similar tools
  • Build CI/CD, observability, and data-quality tooling
  • Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability to meet HIPAA and SOC 2 standards
  • Set schema contracts, ingestion patterns, and self-service tooling for product and analytics teams
  • Make architectural decisions, write critical production code, and establish engineering patterns
Requirements
  • 6+ years of software engineering experience, with significant time building or operating data platforms at scale
  • Experience with CDC, including Debezium or equivalent
  • Experience with streaming technologies such as Kafka or Redpanda
  • Experience with data lake formats including Iceberg, Delta Lake, or Hudi
  • Experience with MPP or lakehouse query engines such as StarRocks, ClickHouse, Trino, Snowflake, or Databricks
  • Experience with dbt
  • Fluent in SQL, schema design, and query optimization
  • Ability to reason about cost and latency trade-offs on large datasets
  • Experience running production data infrastructure, including orchestration, observability, on-call, data quality, and incident response
  • Preferred: direct production experience with Debezium, StarRocks, and dbt
  • Preferred: experience with semantic layers or data catalogs/lineage
  • Preferred: experience with HIPAA-regulated data, PHI handling, de-identification, and access governance
  • Preferred: experience powering AI/ML workloads
  • Preferred: experience across AWS, GCP, and Azure; infrastructure-as-code; and Kubernetes controllers
  • Must be comfortable working in an office 3–4 days per week
  • Must answer whether sponsorship to work in the US is required
Core Competencies

Demonstrates expertise in building and operating data platforms, including CDC pipelines, data lakes, and query engines. Proficient in SQL, schema design, and optimizing performance while ensuring compliance with HIPAA and SOC 2 standards.

Highest-signal resume keywords
  • CDC Pipeline Development
  • Data Lake Architecture
  • SQL Proficiency
  • Data Quality Management
  • Cloud Infrastructure Experience
ATS Optimization Keywords
Hard Skills
  • Debezium
  • Kafka
  • Redpanda
  • Iceberg
  • Delta Lake
  • Hudi
  • StarRocks
  • Dbt
  • SQL
  • Schema Design
Industry Keywords
  • HIPAA
  • SOC 2
  • PHI Handling
  • Data Governance
  • AI/ML Workloads
Tools & Technologies
  • Airflow
  • Dagster
  • CI/CD
  • Observability Tools
  • Data Quality Tooling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Data Engineer
Senior Staff Data Engineer

Jobtailor • Town of Texas (WI)

On-site
USD 150,000 - 200,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

commure • United States

On-site
USD 180,000 - 240,000
Senior Data Platform Engineer
Senior Data Platform Engineer

Jobtailor • Lehi (UT)

On-site
USD 140,000 - 190,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

Athelas • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

Commure • Northern (KY)

Hybrid
USD 150,000 - 210,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

Commure • Mountain View (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

Monograph • San Francisco (CA)

On-site
USD 190,000 - 270,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Naperville (IL)

On-site
USD 120,000 - 160,000
Staff Data Engineer
Staff Data Engineer

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, Data Warehouse
Staff Software Engineer, Data Warehouse

Socket.dev • Mountain View (CA)

On-site
USD 210,000 - 260,000