Big Data Engineer

TP-Link

Irvine (CA)

On-site

USD 100,000 - 120,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Free snacks and drinks
Fully paid medical, dental, and vision
401K contribution
Bi-annual reviews and pay raises
Gym membership
Team-building events

Job summary

TP-Link Systems Inc. in Irvine, CA is seeking a data engineer to design and operate scalable ETL pipelines on AWS cloud.

You will build pipelines on EMR/Spark orchestrated by Airflow, write SQL and Python, ensure data quality, and partner with analysts to deliver reliable datasets for dashboards and analytics. Experience with dimensional modeling, star schema, and lakehouse formats recommended; comfortable with Parquet/ORC, S3, Git, and cost-conscious engineering.

Qualifications

  • 2–4 years of hands-on data development in production environments.

Responsibilities

  • Build, run, and own ETL pipelines on EMR (Spark) orchestrated in Airflow.
  • Write SQL and Python for data ingestion, transformation, and datasets for dashboards.
  • Own data quality checks and extend data models (fact & dimension, star schema).
  • Keep jobs efficient in runtime and cost, and automate repetitive tasks.
  • Collaborate with analysts and stakeholders to deliver usable datasets.

Skills

SQL
Python
Airflow
AWS
Data modeling
Git
Pandas
PySpark

Education

Bachelor's degree in Computer Science or related field

Tools

Spark
Databricks
Airflow
Databases (SQL)
boto3

Job description

ABOUT US:

Headquartered in the United States, TP-Link Systems Inc. is a global provider of reliable networking devices and smart home products, consistently ranked as the world’s top provider of Wi‑Fi devices. The company is committed to delivering innovative products that enhance people’s lives through faster, more reliable connectivity. With a commitment to excellence, TP-Link serves customers in over 170 countries and continues to grow its global footprint.

We believe technology changes the world for the better! At TP-Link Systems Inc, we are committed to crafting dependable, high-performance products to connect users worldwide with the wonders of technology.

Embracing professionalism, innovation, excellence, and simplicity, we aim to assist our clients in achieving remarkable global performance and enable consumers to enjoy a seamless, effortless lifestyle.

KEY RESPONSIBILITIES
  • Build, run, and own ETL pipelines on EMR (Spark) orchestrated in Airflow — including their monitoring, alerting, and recovery.
  • Write SQL and Python for data ingestion, transformation, and the datasets that analysts and dashboards depend on.
  • Own data quality for what you build — run the checks before you ship and add new ones where they're missing; confirm the results are correct, not only that the job completed.
  • Build and extend the data models for your area: fact and dimension tables following the team's layering conventions.
  • Keep your jobs efficient — watch runtime and cost, and raise slow or expensive jobs rather than living with them.
  • Use and extend the team's shared patterns and templates, and turn work you find yourself repeating into something automated or reusable.
  • Work with analysts and business stakeholders to turn requests into datasets that actually get used.
REQUIRED QUALIFICATIONS
  • 2–4 years of hands‑on data development in a production environment — pipelines that run on a schedule with real downstream consumers, that you were responsible for including when they broke.
  • Strong SQL: window functions, complex multi‑table joins, incremental loads, and the ability to work out why a query is slow.
  • Python for production ETL and tooling (PySpark, pandas, boto3) — code that runs on a schedule, not only notebooks.
  • Hands‑on experience with Spark on a cloud platform: AWS EMR, Databricks, Glue, or equivalent.
  • Production experience with a scheduler, Airflow preferred: DAG design, dependencies, retries, and reruns that are safe to repeat.
  • Working understanding of dimensional modeling: fact and dimension tables, star schema, and warehouse layering.
  • Comfortable working on AWS (S3 with Parquet / ORC), Linux, and Git — and comfortable picking up new tools as the platform evolves.
  • Effective use of AI to solve data problems — using it to move faster on SQL, debugging, and unfamiliar schemas, with the judgment to catch output that looks right but isn't.
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
PREFERRED QUALIFICATIONS
  • Data ingestion or CDC tooling: DataX, Sqoop, Debezium, Fivetran, or similar.
  • An OLAP / MPP engine: StarRocks, Doris, ClickHouse, Redshift, or similar.
  • Spark performance work: partitioning, shuffle, skew, and memory tuning.
  • AWS cost optimization: EMR instance sizing and Spot strategy, S3 lifecycle policies.
  • Lakehouse formats (Iceberg, Hudi, Delta), streaming (Kafka, Flink), dbt, or a data quality framework.
  • QuickSight or another BI tool: dataset and permission design.
  • Self‑directed learning: a project, an open‑source contribution, or a tool you picked up on your own and put to real use

Base Salary Range: $100 - 120K

  • Free snacks and drinks
  • Fully paid medical, dental, and vision insurance (partial coverage for dependents)
  • Contributions to 401K funds
  • Bi‑annual reviews, and annual pay increases
  • Health and wellness benefits, including free gym membership
  • Quarterly team‑building events
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer
Big Data Engineer

Socket.dev • Irvine (CA)

On-site
USD 120,000 - 180,000
Free snacks & drinks
Full medical/dental/vision insurance
401K contributions
Data Engineer
Data Engineer

Oscar • Grand Prairie (TX)

On-site
USD 110,000 - 150,000
Medical coverage
Dental coverage
Vision coverage
+3
Senior Data Engineer
Senior Data Engineer

Jobtailor • Naperville (IL)

On-site
USD 120,000 - 160,000
Data & Software Engineer
Data & Software Engineer

Avalore.ai • Chantilly (VA)

On-site
USD 130,000 - 180,000
Health care benefits
Retirement plan (401k)
Life Insurance
+4
Senior Big Data Engineer
Senior Big Data Engineer

TP-Link Systems Inc. • Irvine (CA)

On-site
USD 150,000 - 180,000
Free snacks and drinks
Fully paid medical insurance
401k contributions
+3
Data Engineer ID43228
Data Engineer ID43228

AgileEngine • Miami (FL)

On-site
USD 90,000 - 120,000
Professional growth opportunities
Competitive USD-based compensation
Flexible work schedule
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Senior PySpark Data Developer
Senior PySpark Data Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Data & Software Engineer
Data & Software Engineer

Avalore • Chantilly (VA)

On-site
USD 120,000 - 180,000
Health insurance
401k
Life Insurance
+3