Get more replies from employers
Send a job-specific resume in minutes.
The role focuses on designing and governing a comprehensive data architecture for a Lakehouse environment, spanning HDFS, Ozone, Spark, Hive, and streaming layers. You will architect policy-driven data governance, implement security, lineage, and quality controls, and drive CI/CD for data pipelines and infrastructure.
You will lead architecture reviews, mentor engineers, and collaborate with security, compliance, and domain SMEs to ensure robust, auditable, and scalable data platforms in a large
The role will also architect a Rules Engine for business validations and transformations and establish a comprehensive Data Governance layer covering security lineage quality compliance and operational readiness
Lead HLD and LLD for the Lakehouse covering HDFS Ozone YARN Hive 3 Impala Spark Kafka NiFi Ranger Atlas Knox KMS and HMS with HA Define storage and compute topology cluster sizing multi-environment architecture and network zones Design perimeter security using Knox and TLS Establish Bronze Silver and Gold standards including naming data contracts partitioning data quality checkpoints schema evolution retention and recovery patterns Align serving models for BI and analytics using Impala and Phoenix Design batch pipelines using Informatica BDM Spark and Hive Design speed and streaming paths using NiFi Kafka Hudi Kudu HBase and Phoenix Define the convergence and serving of batch and streaming outputs into governed data stores Architect a declarative Rules Engine covering validations standardization derivations survivorship consent PII handling and SLA routing Ensure rules are metadata-driven versioned testable and CI CD integrated Implement data governance using Ranger and Atlas including policies masking classifications lineage and glossary management Define data lifecycle retention legal holds PII PCI SOX requirements and audit trails Establish CI CD for infrastructure schemas table evolution rules packs and ETL deployments Build SLA dashboards alerts runbooks and observability processes Architect disaster recovery and business continuity using BDR Replication Manager and replication strategies Partner with Domain SMEs Platform Administrators Security and Compliance teams Lead Architecture Review Boards and mentor developers and data engineering teams
10+ years of experience in data architecture/engineering.Bacheloru2019s or Masteru2019s in Computer Science, Engineering, or related field or equivalent experience.3u20135+ years of experience architecting on Cloudera CDH/CDP in production.Proven experience delivering on-premises Lakehouse environments using Medallion and Lambda architectures.Strong understanding of industry standards and best practices in Data Engineering.Good understanding of Informatica DEI suite for framework/pattern design and pushdown to Spark/Hive.Deep practical exposure to Apache Iceberg, Apache Hudi, Apache Kudu, HBase, Phoenix, Hive 3, and Impala.Strong SQL expertise, including windowing, partitioning, MERGE/UPSERT, and cost-based tuning.Strong Impala/Phoenix performance tuning experience.Experience with Kerberos, TLS, AD/LDAP, Ranger policies, Atlas lineage/glossary, and SDX concepts.Experience designing metadata-driven Rules Engines and data quality frameworks for both batch and streaming environments.Solid Linux fundamentals.Experience with Git and CI/CD.Strong documentation and stakeholder communication skills.Experience with Kafka patterns and CDC from RDBMS is preferred.Familiarity with Ozone, KMS/Key Trustee, and air-gapped deployments is preferred.SRE experience, including SLOs, error budgets, RCA, and disaster recovery exercises, is preferred.Cloudera CDP, Informatica, and Security certifications are a plus.Regulatory experience in PII/PCI/SOX/GDPR environments is preferred.