Big Data Engineer

Socket.dev

Irvine (CA)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Free snacks & drinks
Full medical/dental/vision insurance
401K contributions

Job summary

TP-Link Systems Inc. seeks a data platform engineer to own end-to-end data pipelines and quality across analytics workloads. You will design ETL on EMR using Spark and Airflow, ingesting from OLTP sources with DataX and CDC, and model data with dimensional schemas for real analytical use cases.

The role requires production experience with Spark, Airflow, and data ingestion tools, plus AWS, Linux, and Git. You’ll collaborate with analysts to turn requirements into reusable datasets and scalable

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 2+ years hands-on data development in production environments.
  • Experience owning a pipeline or component in production and being accountable.
  • Strong SQL with analytics functions and multi-table joins.
  • Production-grade Python with ETL tooling (PySpark / pandas / boto3).
  • Experience with Spark on EMR and/or Hive, partitioning and memory tuning.
  • Airflow DAG design, retries, and SLA alerting.
  • Experience with data ingestion tools (DataX, Sqoop, Debezium).
  • Exposure to StarRocks, Doris, or ClickHouse and data warehouse modeling.
  • Comfort with AWS, Linux, and Git.
  • Ability to learn new engines quickly and use AI to assist data problems.

Responsibilities

  • Own a data platform component end to end including pipelines and monitoring.
  • Design and build ETL on EMR (Spark) and Airflow; ingest from OLTP sources with DataX/CDC.
  • Model data using dimensional schemas for analytical use cases.
  • Ensure data quality and write monitoring checks; resolve issues upstream.
  • Tune runtime and infrastructure costs; choose appropriate engines.
  • Create reusable components and automate repetitive tasks.
  • Collaborate with analysts to translate requirements into usable datasets.

Skills

SQL
Python
Spark EMR
Airflow
ETL
Amazon Web Services
Linux
Git
Pandas
boto3
Data ingestion
CDC

Education

Bachelor's degree in CS/IS or equivalent

Tools

Spark on EMR
DataX
Sqoop
Debezium
StarRocks
Doris
ClickHouse
dbt
Glue Data Catalog
QuickSight

Job description

ABOUT US:

Headquartered in the United States, TP-Link Systems Inc. is a global provider of reliable networking devices and smart home products, consistently ranked as the world's top provider of Wi-Fi devices. The company is committed to delivering innovative products that enhance people’s lives through faster, more reliable connectivity. With a commitment to excellence, TP-Link serves customers in over 170 countries and continues to grow its global footprint.

We believe technology changes the world for the better! At TP-Link Systems Inc, we are committed to crafting dependable, high-performance products to connect users worldwide with the wonders of technology.

Embracing professionalism, innovation, excellence, and simplicity, we aim to assist our clients in achieving remarkable global performance and enable consumers to enjoy a seamless, effortless lifestyle.

KEY RESPONSIBILITIES
  • Own a component of the data platform end to end - its pipelines, tables, quality checks, monitoring, and recovery.
  • Design and build ETL on EMR (Spark) and Airflow, and data ingestion from OLTP sources via DataX / CDC, including incremental sync and upstream schema changes.
  • Model the data for your area: dimensional models and layering (ODS / DWD / DWS / ADS) that serve real analytical use cases.
  • Own data quality for your component - write the checks and freshness monitoring, and resolve issues before they reach downstream users.
  • Tune runtime and infrastructure cost for the jobs you own, and choose the right engine for a workload (StarRocks serving vs. batch on EMR).
  • Build components other engineers can reuse, and automate work that would otherwise repeat.
  • Work directly with analysts and business stakeholders to turn requirements into models and datasets that actually get used.
REQUIRED QUALIFICATIONS
  • Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
  • 2+ years of hands-on data development in a production environment - scheduled pipelines with real downstream consumers.
  • Experience owning a pipeline or component in production: you were the person accountable for it, including when it broke.
  • Strong SQL: window functions, complex multi-table joins, incremental and idempotent writes; able to read an execution plan and diagnose data skew or slow queries.
  • Production-grade Python: maintainable, testable ETL and tooling code (PySpark / pandas / boto3) with proper error handling and logging.
  • Hands-on with Spark on EMR, or equivalent Hadoop / Hive experience, including partitioning, shuffle, and memory tuning.
  • Production experience with Airflow: DAG design, dependencies, retries, idempotent reruns, backfills, and SLA alerting.
  • Experience with a data ingestion or CDC tool in production (DataX, Sqoop, Debezium, or similar): extracting from OLTP sources, incremental sync, and handling upstream schema changes.
  • Experience with an OLAP engine - StarRocks, Doris, or ClickHouse: table models, partitioning and bucketing, materialized views, and query tuning.
  • Solid grasp of data warehouse modeling: dimensional modeling, slowly changing dimensions, and layering conventions.
  • Comfortable on AWS (S3 with Parquet / ORC and sensible partitioning, IAM basics, day-to-day EMR operations), Linux, and Git.
  • Comfortable picking up new tools as the platform evolves - we'd rather hire someone who learns a new engine quickly than someone who has only ever used ours.
  • Effective use of AI to solve data problems - using it to move faster on SQL, debugging, and unfamiliar schemas, with the judgment to catch output that looks right but isn't.
PREFERRED QUALIFICATIONS
  • AWS cost optimization: EMR instance sizing and Spot strategy, S3 lifecycle policies, StarRocks vs. Athena trade-offs.
  • Lakehouse table formats: Iceberg, Hudi, or Delta.
  • Streaming: Kafka with Flink or Spark Structured Streaming.
  • dbt, Glue Data Catalog, data lineage, or a data quality framework.
  • QuickSight dataset, SPICE, and row-level permission design.
Benefits

Salary range: TBD

  • Free snacks and drinks
  • Fully paid medical, dental, and vision insurance (partial coverage for dependents)
  • Contributions to 401K funds
  • Bi-annual reviews, and annual pay increases
  • Health and wellness benefits, including free gym membership
  • Quarterly team-building events

At TP-Link Systems Inc., we are continually searching for ambitious individuals who are passionate about their work. We believe that diversity fuels innovation, collaboration, and drives our entrepreneurial spirit. As a global company, we highly value diverse perspectives and are committed to cultivating an environment where all voices are heard, respected, and valued. We are dedicated to providing equal employment opportunities to all employees and applicants, and we prohibit discrimination and harassment of any kind based on race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. Beyond compliance, we strive to create a supportive and growth-oriented workplace for everyone. If you share our passion and connection to this mission, we welcome you to apply and join us in building a vibrant and inclusive team at TP-Link Systems Inc.

Please, no third-party agency inquiries, and we are unable to offer visa sponsorships at this time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer
Big Data Engineer

TP-Link • Irvine (CA)

On-site
USD 100,000 - 120,000
Free snacks and drinks
Fully paid medical, dental, and vision
401K contribution
+3
Senior Big Data Engineer
Senior Big Data Engineer

TP-Link Systems Inc. • Irvine (CA)

On-site
USD 150,000 - 180,000
Free snacks and drinks
Fully paid medical insurance
401k contributions
+3
Cloud Software Engineer - Database
Cloud Software Engineer - Database

TP-Link Corporation Limited • Irvine (CA)

On-site
USD 90,000 - 130,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
401k contributions
+2
Senior Cloud Software Engineer - Database
Senior Cloud Software Engineer - Database

TP-Link Corporation Limited • Irvine (CA)

On-site
USD 120,000 - 150,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
Contributions to 401k funds
+2
Senior Cloud Software Engineer - Database
Senior Cloud Software Engineer - Database

TP-Link • Irvine (CA)

On-site
USD 120,000 - 150,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
Contributions to 401k funds
+2
Senior Cloud Engineer - Distributed Database & Middleware
Senior Cloud Engineer - Distributed Database & Middleware

TP-Link • Irvine (CA)

On-site
USD 150,000 - 180,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
401k contributions
+3
Cloud Software Engineer, Backend - Enterprise
Cloud Software Engineer, Backend - Enterprise

TP-Link Systems Inc. • Irvine (CA)

On-site
USD 120,000 - 180,000
Free snacks and drinks
Fully paid insurance
401k contributions
+3
Data-Driven HR Operations Associate
Data-Driven HR Operations Associate

VIGI Retail Solution • Irvine (CA)

On-site
USD 73,000 - 83,000
Free snacks and drinks
Fully paid medical, dental, and vision insurance
401k contributions
+2
Cloud Software Engineer, Backend - Enterprise
Cloud Software Engineer, Backend - Enterprise

TP-Link Corporation Limited • Irvine (CA)

On-site
USD 120,000 - 180,000
Free snacks and drinks
Fully paid medical, dental, and vision
401k contributions
+4
Senior Fullstack Software Engineer
Senior Fullstack Software Engineer

Worky • Irvine (CA)

On-site
USD 170,000 - 200,000
Free snacks and drinks
Fully paid medical, dental, and vision
401K contributions
+3