Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs

Irvine (CA)

On-site

USD 120,000 - 180,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitus Labs in Irvine, CA, is seeking a mid-level Data Engineer to own and scale an AWS-based data lakehouse. You will write production code in Java and Python daily, build robust data pipelines, and contribute to platform design and data governance.

You will collaborate with data scientists, analytics engineers, and product teams, mentor junior engineers, and help maintain data quality, reliability, and security across the data platform.

Qualifications

  • 3+ years of professional data engineering experience on AWS
  • Production data pipelines with defined freshness/reliability SLAs
  • Experience with medallion lakehouse concepts and Iceberg
  • Strong Java or Python development skills for Spark jobs

Responsibilities

  • Build and scale AWS-based data lakehouse on S3 with Iceberg
  • Develop high-throughput ETL/ELT pipelines using AWS Glue, EMR, and Lambda
  • Implement schema evolution, partitioning, and compaction for Iceberg tables
  • Write production-grade pipeline code in Java and Python
  • Collaborate with data scientists and analytics engineers
  • Mentor junior engineers and promote data governance standards

Skills

Java
Python
Spark
SQL
Data pipelines
Data quality
Mentoring
Team collaboration
Documentation

Education

Bachelor's degree in CS/engineering

Tools

AWS Glue
EMR
Kinesis
Iceberg
Airflow
CodePipeline
GitHub Actions
Apache Kafka

Job description

We are looking for a mid-level Data Engineer to join our Data Platform team and take ownership of building and scaling our AWS-based data lakehouse. You will architect and deliver robust, production-grade data pipelines, work closely with data scientists, analytics engineers, and product teams, and set the technical direction for how data flows across the organization. This is a hands‑on engineering role — you will write production code in Java and Python every day, while also contributing to platform design decisions, mentoring junior engineers, and driving best practices around data quality, reliability, and governance.

Key Responsibilities
Data Lakehouse Development
  • Build and extend medallion-architecture data lakehouse layers (Bronze / Silver / Gold) on AWS S3 using the Apache Iceberg table format.
  • Develop and maintain high-throughput ETL/ELT pipelines using AWS Glue, EMR (Spark), and Lambda.
  • Implement schema evolution, partitioning strategies, and compaction processes for Iceberg tables to optimize storage and query performance.
  • Write production-quality pipeline code in Java and Python, following team conventions for structure, testing, and maintainability.
Real-Time & Batch Streaming
  • Build and operate event-driven data pipelines using Amazon Kinesis Data Streams, Kinesis Firehose, or Apache Kafka (MSK).
  • Implement exactly-once or at-least-once processing semantics for streaming workloads using Apache Flink or Spark Structured Streaming on EMR.
  • Monitor and help tune cost and performance across AWS services including S3, Glue, Athena, Redshift Spectrum, EMR, Lambda, Step Functions, and EventBridge.
  • Follow platform security standards in day-to-day work: IAM least-privilege policies, KMS encryption, and VPC networking.
  • Maintain and extend CI/CD pipelines for data workloads using AWS CodePipeline, GitHub Actions, or equivalent.
Data Quality & Governance
  • Add data quality checks using frameworks such as Great Expectations or Deequ, and integrate validation steps into pipeline orchestration.
  • Help maintain and uphold data contracts between producing and consuming systems.
  • Contribute to data cataloguing and lineage tracking using AWS Glue Data Catalog or Apache Atlas.
Collaboration & Ways of Working
  • Partner with data scientists, ML engineers, and analysts to understand data requirements and deliver performant, well-documented datasets.
  • Take an active part in code reviews, design discussions, and pair programming, and support junior engineers where you can.
  • Document your pipeline designs and decisions, and contribute to the internal engineering knowledge base.
Required Qualifications
Experience
  • 3+ years of professional data engineering experience, with at least 1–2 years on AWS cloud platforms.
  • Experience building and supporting production data pipelines, ideally on large datasets with defined freshness or reliability SLAs.
  • Working knowledge of data lakehouse concepts — medallion pattern and open table formats (Iceberg preferred; Delta Lake or Hudi acceptable).
Programming Languages
  • Java: Working proficiency in Java (8+) for Spark jobs and pipeline components, with familiarity with Maven or Gradle build systems.
  • Python: Experience with pandas, PySpark, boto3, and automation tooling; strong skills in one language and willingness to ramp up on the other.
AWS Core Services
  • AWS Core Services, (Spark/Flink), Lambda, EC2.
  • Streaming: Kinesis Data Streams, Kinesis Firehose, or MSK (Managed Kafka).
  • Orchestration: Step Functions, MWAA (Managed Airflow), or EventBridge Scheduler.
  • Querying: Athena, Redshift, or Redshift Spectrum.
  • Security & Governance: Comfortable working within IAM, KMS, Secrets Manager, and VPC setups.
  • DevOps: Exposure to AWS CDK or CloudFormation, and to CodePipeline or equivalent CI/CD tools.
Data Processing Frameworks
  • Apache Spark (PySpark and/or Spark Java API) — distributed transformations and a working grasp of performance tuning.
  • Apache Iceberg — reading and writing tables, time travel, and basic table maintenance.
  • SQL — strong SQL for data transformation, including window functions, CTEs, and query tuning.
  • Must be Chinese Mandarin fluent.
Preferred Qualifications
  • AWS Certified Data Engineer – Associate or AWS Certified Solutions Architect certification.
  • Experience with dbt for SQL-based transformation layers on top of the lakehouse.
  • Familiarity with ML platform integration: feature stores (SageMaker Feature Store), model serving data needs, or MLflow experiment tracking.
  • Experience with real-time OLAP engines such as Apache Druid or ClickHouse.
  • Experience with Lake Formation fine-grained access control, or with data cataloguing and lineage tooling.
  • Exposure to data mesh or data product thinking — domain ownership and
Tech Stack at a Glance
Languages

Java/ Python 3

AWS (S3, Glue, EMR, Kinesis, Athena, Lambda, Step Functions, Lake Formation, CDK)
Processing

Apache Spark, Apache Flink, Spark Structured Streaming

Table Format
Streaming
Orchestration

Apache Airflow (MWAA), AWS Step Functions

IaC & CI/CD

AWS CDK / Terraform, GitHub Actions / CodePipeline

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Python, SQL, AWS
Data Engineer - Python, SQL, AWS

Compunnel, Inc. • Durham (NC)

On-site
USD 95,000 - 120,000
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
AWS Data Engineer
AWS Data Engineer

Spectraforce Technologies • Newark (NJ)

Hybrid
USD 120,000 - 170,000
Data Engineer – Baltimore City, MD
Data Engineer – Baltimore City, MD

Creative Information Technology India • Falls Church (VA)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

Compunnel, Inc. • Boston (MA)

On-site
USD 110,000 - 140,000
AWS Data Engineer
AWS Data Engineer

MMD Services, Inc • Rosemont (IL)

Hybrid
USD 135,000 - 150,000
401(k)
401(k) matching
Dental insurance
+5
AWS Data Engineer
AWS Data Engineer

Capgemini • Newark (NJ)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

Oscar • Grand Prairie (TX)

On-site
USD 110,000 - 150,000
Medical coverage
Dental coverage
Vision coverage
+3
Senior Data Engineer: Real-Time Pipelines on AWS
Senior Data Engineer: Real-Time Pipelines on AWS

eOne Infotech • Fort Mill (SC)

On-site
USD 90,000 - 120,000
Senior AWS Data Engineer — Streaming & Data Lake Architect
Senior AWS Data Engineer — Streaming & Data Lake Architect

Tech Mirrors • Malvern

On-site