Position: Iceberg DBA / Lakehouse Operations Engineer
Location : Irving, TX
Visa Preference : USC/ GC's
Job Summary
We are looking for an experienced Iceberg DBA / Lakehouse Operations Engineer to own the reliability, performance, and operational integrity of enterprise-scale Apache Iceberg data platforms.
The role supports a multi-engine Lakehouse environment across Spark, Hive, Impala, and Trino, with a strong focus on table optimization, metadata management, performance tuning, production support, and Hive/Teradata-to-Iceberg modernization.
Key Responsibilities
- Own day-to-day operations and reliability of Apache Iceberg tables at TB/PB scale.
- Perform table maintenance including:
- Compaction and small-file management
- Snapshot expiration and metadata cleanup
- Optimize table performance through partitioning, file sizing, clustering, ordering, and pruning.
- Manage Iceberg metadata, snapshots, retention, archival, and time-travel capabilities.
- Ensure consistent query performance across Spark/CDE and Hive/Impala/CDW.
- Troubleshoot query failures, inefficient execution plans, metadata issues, and performance bottlenecks.
- Support data modeling and Lakehouse architecture, including 3NF and Bronze/Silver/Gold (Medallion) patterns.
- Ensure data accuracy through validation, reconciliation, and source-to-Iceberg checks.
- Provide L2/L3 production support, including P1/P2 incident resolution, RCA, and preventive actions.
- Support security and governance through Ranger policies and RBAC.
- Collaborate with Data Engineering, Platform, and Application teams on Lakehouse modernization and operational improvements.
Required Skills
- Strong hands‑on experience with Apache Iceberg and/or Hive-based data lakes.
- Strong expertise in Iceberg table management, optimization, metadata, and performance tuning.
- Experience managing large‑scale TB/PB datasets.
- Strong knowledge of Spark SQL, Hive, Impala, NiFi, and Trino.
- Strong understanding of partitioning, Parquet/ORC, and distributed query processing.
- Knowledge of data modeling and Medallion Lakehouse architecture.
- Experience with production support, troubleshooting, RCA, and on‑call environments.
- Scripting/automation skills using Python and/or Shell.
Preferred Skills
- Experience with Hive‑to‑Iceberg or Teradata‑to‑Iceberg migrations.
- Experience with Cloudera CDP, CDE, and CDW.
- Cloud experience with AWS and/or Azure.
- Experience with enterprise data modernization initiatives.
- Spark / Hive / Impala / Trino
- Cloudera CDP (CDE/CDW)
- Mission‑critical analytics and reporting workloads