- Own Commure’s data warehouse platform end-to-end, including CDC pipelines, data lake, query layer, transformation layer, and analytics-facing tooling
- Design and operate low-latency, high-fidelity CDC pipelines using Debezium and Kafka, Redpanda, or an equivalent streaming backbone
- Architect a data lake on object storage using Iceberg, Delta Lake, or Hudi with Parquet
- Enable both batch and streaming workloads with separation of storage from compute
- Run and scale StarRocks or adjacent MPP/lakehouse engines as the query and serving layer
- Design schemas, materialized views, ingestion patterns, and performance/cost trade-offs
- Build the dbt transformation layer, including modeling standards, tests, documentation, and a semantic layer
- Establish orchestration with Airflow, Dagster, or similar tools
- Build CI/CD, observability, and data-quality tooling
- Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability to meet HIPAA and SOC 2 standards
- Set schema contracts, ingestion patterns, and self-service tooling for product and analytics teams
- Make architectural decisions, write critical production code, and establish engineering patterns
Requirements
- 6+ years of software engineering experience, with significant time building or operating data platforms at scale
- Experience with CDC, including Debezium or equivalent
- Experience with streaming technologies such as Kafka or Redpanda
- Experience with data lake formats including Iceberg, Delta Lake, or Hudi
- Experience with MPP or lakehouse query engines such as StarRocks, ClickHouse, Trino, Snowflake, or Databricks
- Experience with dbt
- Fluent in SQL, schema design, and query optimization
- Ability to reason about cost and latency trade-offs on large datasets
- Experience running production data infrastructure, including orchestration, observability, on-call, data quality, and incident response
- Preferred: direct production experience with Debezium, StarRocks, and dbt
- Preferred: experience with semantic layers or data catalogs/lineage
- Preferred: experience with HIPAA-regulated data, PHI handling, de-identification, and access governance
- Preferred: experience powering AI/ML workloads
- Preferred: experience across AWS, GCP, and Azure; infrastructure-as-code; and Kubernetes controllers
- Must be comfortable working in an office 3–4 days per week
- Must answer whether sponsorship to work in the US is required
Core Competencies
Demonstrates expertise in building and operating data platforms, including CDC pipelines, data lakes, and query engines. Proficient in SQL, schema design, and optimizing performance while ensuring compliance with HIPAA and SOC 2 standards.
Highest-signal resume keywords
- CDC Pipeline Development
- Data Lake Architecture
- SQL Proficiency
- Data Quality Management
- Cloud Infrastructure Experience
ATS Optimization Keywords
Hard Skills
- Debezium
- Kafka
- Redpanda
- Iceberg
- Delta Lake
- Hudi
- StarRocks
- Dbt
- SQL
- Schema Design
Industry Keywords
- HIPAA
- SOC 2
- PHI Handling
- Data Governance
- AI/ML Workloads
Tools & Technologies
- Airflow
- Dagster
- CI/CD
- Observability Tools
- Data Quality Tooling