Apache Iceberg Engineer

Smart IT Frame LLC

Sunnyvale (CA)

On-site

USD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An innovative firm is seeking an experienced Apache Iceberg Engineer to design and optimize large-scale data lakehouse solutions. This role involves collaborating with data engineers and scientists to create efficient data architectures using cutting-edge technologies like Apache Spark and cloud storage solutions. The ideal candidate will possess strong expertise in big data frameworks, data governance, and cloud platforms. Join a forward-thinking team that values creativity and technical excellence, and make a significant impact on data management practices in a dynamic environment. If you're passionate about data engineering and eager to tackle complex challenges, this opportunity is perfect for you.

Qualifications

  • 3+ years of experience in Big Data, Data Engineering, or Cloud Data Warehousing.
  • Hands-on experience with Apache Iceberg in a production environment.

Responsibilities

  • Design and optimize Iceberg-based data lake architectures for large-scale datasets.
  • Develop data ingestion and transformation pipelines using Spark and Flink.

Skills

Apache Iceberg
Big Data Processing
Apache Spark
Flink
SQL
Data Governance
Cloud Storage Solutions
Data Lakehouse Architectures
Python
Java

Education

Bachelor's degree in Computer Science
Master's degree in Data Engineering

Tools

AWS S3
Google Cloud Storage
Azure Data Lake Storage
Terraform
Kubernetes
Airflow

Job description

We are looking for an experienced Apache Iceberg Engineer to design, develop, and optimize large-scale data lakehouse solutions leveraging Apache Iceberg. The ideal candidate will have expertise in big data processing frameworks (Apache Spark, Flink, Presto, Trino, Hive) and cloud-based data platforms like AWS S3, Google Cloud Storage, or Azure Data Lake Storage. You will work closely with data engineers, data scientists, and DevOps teams to ensure efficient, scalable, and reliable data architecture.

Key Responsibilities:

  • Design, implement, and optimize Iceberg-based data lake architectures for large-scale datasets.
  • Develop data ingestion, transformation, and query optimization pipelines using Spark, Flink, or Presto/Trino.
  • Ensure ACID compliance, schema evolution, and partition evolution in Iceberg tables.
  • Implement time travel, versioning, and snapshot management for historical data analysis.
  • Optimize metadata management and query performance in Iceberg-based data lakes.
  • Integrate Apache Iceberg with cloud storage solutions (AWS S3, GCS, ADLS) and data warehouses.
  • Implement best practices for data governance, access control, and security within an Iceberg-based environment.
  • Troubleshoot performance issues, metadata inefficiencies, and schema inconsistencies in Iceberg tables.
  • Collaborate with DevOps, ML engineers, and BI teams to enable smooth data workflows.

Required Qualifications:

  • Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field.
  • 3+ years of experience in Big Data, Data Engineering, or Cloud Data Warehousing.
  • Hands-on experience with Apache Iceberg in a production environment.
  • Strong expertise in Apache Spark, Flink, Trino, Presto, or Hive for big data processing.
  • Proficiency in SQL and distributed query engines.
  • Experience working with cloud storage solutions (AWS S3, GCS, ADLS).
  • Knowledge of data lakehouse architectures and modern data management principles.
  • Familiarity with schema evolution, ACID transactions, and partitioning techniques.
  • Experience with Python, Scala, or Java for data processing.

Preferred Qualifications:

  • Experience in real-time data processing using Flink or Kafka.
  • Understanding of data governance, access control, and compliance frameworks.
  • Knowledge of other data lake frameworks like Delta Lake (Databricks) or Apache Hudi.
  • Hands-on experience with Terraform, Kubernetes, or Airflow for data pipeline automation.
Seniority level

Mid-Senior level

Employment type

Full-time

Job function

Other

Industries

Software Development and IT Services and IT Consulting

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DataOps Engineer
DataOps Engineer

Woongjin, Inc • Englewood Cliffs (NJ)

On-site
USD 110,000
DataOps Engineer
DataOps Engineer

SBT Global, Inc. • Englewood Cliffs (NJ)

On-site
USD 140,000 - 220,000
Remote Apache Iceberg Engineer - Data Lake Performance
Remote Apache Iceberg Engineer - Data Lake Performance

BairesDev • Peru (IL)

On-site
USD 120,000 - 180,000
100% remote work (from anywhere)
USD or local currency compensation
Hardware and software setup for work"s
+4
Sr. Staff Software Engineer - Apache Iceberg
Sr. Staff Software Engineer - Apache Iceberg

Cloudera • Washington

Hybrid
USD 184,000 - 230,000
Generous PTO Policy
Flexible WFH Policy
Mental & Physical Wellness programs
+1
Staff Lakehouse Engineer - Iceberg, Spark & Snowflake
Staff Lakehouse Engineer - Iceberg, Spark & Snowflake

Affirm • Seattle (WA)

On-site
USD 204,000 - 290,000
Health insurance
Flexible Spending Wallets
Time off
+1
Senior Data Lake Engineer — Iceberg, Multi-Cloud
Senior Data Lake Engineer — Iceberg, Multi-Cloud

HR Tech Job • Pleasanton (CA)

Hybrid
USD 190,000 - 286,000
Data Engineer
Data Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 90,000 - 120,000
Data Engineer
Data Engineer

Blutic • Dallas (TX)

Hybrid
USD 90,000 - 120,000
Senior Data Platform Architect - Spark, Iceberg & Lakehouse
Senior Data Platform Architect - Spark, Iceberg & Lakehouse

FloQast, Inc. • San Jose (CA)

On-site
USD 188,000 - 282,000
Sr. Staff Software Engineer - Apache Iceberg
Sr. Staff Software Engineer - Apache Iceberg

Cloudera • Washington

On-site
USD 184,000 - 230,000
Generous PTO
Unplugged days
Flexible WFH
+6