Data Architect | Experience Range: 8-13 Years | Location: Mumbai, Bangalore, Trivandrum
Role Overview
The Data Architect is the senior technical owner of the platform's design. You will define and evolve the architectural blueprint— the canonical data model, the ingestion framework, the transformation patterns, the FHIR serialization layer, the bidirectional flow with FHIR-Repository, and the governance and observability frameworks that hold them together. You will set the standards that every other role implements against. This role works in close partnership with the customer's existing Health Data Engine technical owners, who carry deep operational knowledge of healthcare data at scale.
Key Responsibilities
- Own the end-to-end architecture across ingestion (Flink + PySpark), storage (Iceberg on GCS with BigLake Metastore), transformation (dbt over Starburst), FHIR serialization (flat FHIR Iceberg → bundles → FHIR-Repository), and consumption (Starburst, FHIR API, data products).
- Maintain the architecture specification document and its companion design docs.
- Lead the design and evolution of the Common Data Model (CDM): dimensional, fact, bridge, and reference tables; SCD2 semantics; hash key conventions; write-authority matrix.
- Define the bidirectional FHIR flow with origin-tag-based loop prevention and the FHIR repository egress interceptor contract.
- Define patient identity resolution architecture using Informatica MDM and associated pipelines.
- Set standards for naming conventions, hash algorithms, Iceberg table properties, DLQ taxonomy, observability, and security (PHI handling, encryption, audit logging).
- Approve significant changes to platform-wide rules and review architectural specs.
- Drive high-volume capacity design, perform capacity planning, validation, and parallel-run cutover.
- Evaluate and recommend technology choices; build option analyses with trade-off matrices.
- Provide architectural reviews, mentor engineers and modelers, and lead review meetings with customer stakeholders.
- Establish disaster recovery, backup, and reprocessing strategy with RPO <5 minutes and RTO <4 hours.
Required Skills and Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field.
- 8+ years of data engineering experience with 3-5 years in a Data Architect or Lead Data Engineer role.
- Deep expertise in architecting and operating large-scale data platforms on Google Cloud Platform (Cloud Storage, Dataproc, Cloud Composer, Secret Manager, IAM, networking).
- Hands‑on experience with Apache Iceberg in production (table properties, partitioning, schema evolution).
- Production experience with Starburst (Galaxy or Enterprise) or Trino for analytical workloads.
- Strong experience with Apache Flink and Apache Kafka for streaming architectures.
- Expert SQL and Python programming skills; familiarity with PySpark, PyFlink, and dbt.
- Comprehensive understanding of healthcare data standards (HL7v2, CCDA, FHIR R4 with US Core 6.1 profiles).
- Hands‑on experience with FHIR runtime platforms (FHIR-Repository or HAPI FHIR).
- Experience with master data management for patient identity resolution (Informatica MDM preferred).
- Experience designing for HIPAA-regulated environments (PHI handling, encryption, audit logging).
- Demonstrated ability to lead complex technical initiatives and influence stakeholders.
- Excellent written and verbal communication skills.
Nice-to-Have Skills
- GCP Professional Data Engineer or Cloud Architect certification.
- Experience with governance platforms such as Atlan, Collibra, Alation, or Unity Catalog.
- Experience with multi-region active-active or active-passive deployments on GCP.
- Production experience operating Confluent Cloud at significant scale.
- Familiarity with spec-driven development workflows and AI-assisted code generation.
- Experience replacing or sunsetting legacy healthcare data platforms.