Data Engineer, Bilingual Mandarin

Jobtailor

California (MO)

On-site

USD 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a data engineer to design and maintain scalable data pipelines across online and offline systems. You will ingest data from multiple sources, ensure data quality, and support analytics platforms used by product, operations, and growth teams.

The role emphasizes hands-on coding in Java or Python, strong SQL, and experience with modern data stack tools such as SeaTunnel, Kafka, Flink, and Doris.

Qualifications

  • 1–4 years of experience in data engineering or related roles.
  • Hands-on in data ingestion, pipelines, and data warehouse development.
  • Strong SQL skills and ability to optimize queries.
  • Knowledge of data warehouse layers and modeling concepts.
  • Experience with REST APIs, OAuth, pagination, or similar patterns.

Responsibilities

  • Ingest data from internal and external sources and ensure data quality.
  • Build offline/real-time pipelines using SeaTunnel, Kafka, Flink, or Spark.
  • Support data governance, metadata management, and data catalogs.
  • Collaborate with analysts and engineers across time zones.
  • Contribute to data service APIs and BI data provisioning.

Skills

SQL Proficiency
Data Ingestion
Data Pipeline Development
Data Modeling
Java
Python
Data Quality
Data Governance
Git
English Communication

Tools

SeaTunnel
Kafka
Flink
Spark
Doris
DolphinScheduler
StreamPark
Airflow
dbt
Databricks
PostgreSQL

Job description

  • Ingest data from internal and external business systems, third-party platforms, SaaS products, and external data sources; handle data collection, sync, cleansing, and loading
  • Participate in building offline and real-time data pipelines using SeaTunnel, Kafka, Flink, Spark, or similar technologies to improve ingestion stability and processing efficiency
  • Handle practical challenges in data sync: authentication, pagination, rate limiting, failure retry, incremental sync, backfill, schema changes, and task anomalies
  • Participate in layered data warehouse development across ODS, DWD, DWS, and ADS layers; build and maintain data models
  • Support business domain modeling, metric standardization, shared data model development, and core table maintenance
  • Optimize data organization and query performance on OLAP engines such as Doris to provide stable data support for product, operations, growth, customer success, and management analytics
  • Build and maintain data quality rules for core data pipelines; ensure data accuracy, completeness, consistency, and timeliness
  • Participate in data validation, anomaly detection, alerting, and issue resolution; help improve stability of critical data pipelines
  • Contribute to data governance capabilities including DataHub or similar tools; improve metadata management, data lineage, data asset catalog, and data standards
  • Participate in building data platform capabilities including data development, task scheduling, monitoring, quality management, governance, and service delivery modules
  • Use tools such as DolphinScheduler and StreamPark for task management, scheduling orchestration, and real-time task operations
  • Support the data service layer by delivering standardized APIs, metric services, and data capabilities to internal systems, analytics applications, and business tools
  • Support underlying data for tools like Superset; ensure data availability for BI dashboards, metric boards, and business monitoring
  • Participate in designing and implementing data development automation tools and engineering agents
  • Explore AI agent applications in data development, governance, quality detection, task operations, anomaly diagnosis, and documentation generation
  • Leverage large language models and automation tools to improve data engineering efficiency, task stability, and platform intelligence
Requirements
  • **Must-Have Experience**
  • 1–4 years of experience in data engineering, data platforms, data warehousing, backend development, analytics engineering, or a related role
  • Real project experience in data ingestion, data pipelines, data warehouse development, data modeling, data services, or data platform work
  • Strong learning ability and execution skills; able to independently drive small-to-medium data engineering tasks with clear objectives
  • **SQL Skills**
  • Proficient in SQL for querying, cleansing, aggregation, deduplication, comparison, validation, and metric calculation
  • Familiar with joins, window functions, CTEs, aggregation analysis, incremental logic, and basic performance optimization
  • Understands data warehouse layering concepts: fact tables, dimension tables, subject domains, metric definitions, and shared models
  • **Data Development**
  • Proficient in Java or Python for API integration, data processing, automation scripting, and file handling
  • Understands common engineering patterns: REST APIs, OAuth/API keys, pagination, rate limiting, retry logic, error handling, logging, and task idempotency
  • Good code structure habits; writes clean, maintainable, and reusable code
  • Familiar with Git, code review practices, README documentation, logging, testing, and collaborative engineering workflows
  • **Pipeline & Platform Tools**
  • Familiar with one or more of: SeaTunnel, Kafka, Flink, Spark (data integration, real-time, or offline processing)
  • Familiar with one or more of: Doris, ClickHouse, Snowflake, BigQuery, Redshift, Databricks, PostgreSQL (data warehouse, OLAP, or lakehouse systems)
  • Familiar with one or more of: DolphinScheduler, StreamPark, Airflow, Dagster, Prefect, dbt (scheduling, development, or task management tools)
  • Understands data pipeline operations: scheduling, dependencies, monitoring, failure retry, backfill, version management, and deployment processes
  • Candidates are not expected to master all tools, but must have a solid data engineering foundation and the ability to quickly learn new tech stacks
  • **Data Quality & Governance Mindset**
  • Understands data quality dimensions: accuracy, completeness, consistency, uniqueness, timeliness, and anomaly detection
  • Proactively designs data validation rules and can identify and locate data anomalies
  • Familiar with metadata management, data lineage, data asset catalogs, and data standards; experience with DataHub or similar platforms is a plus
  • **Collaboration & Communication**
  • Able to communicate data requirements with analysts, business stakeholders, backend engineers, and product managers
  • Clearly describes problems, solutions, risks, progress, and deliverables
  • Comfortable with cross-timezone collaboration; strong written and spoken English communication skills
  • Willing to participate in regular fixed collaboration sessions with China-based teams and drive work through documentation and async communication
  • **Nice-to-Have**
  • Experience integrating third-party SaaS data: CRM, ERP, marketing platforms, customer service systems, logistics, e-commerce, payment systems, or ad platforms
  • Experience building data lakehouses, data middle platforms, data platforms, or enterprise-level data warehouses
  • Experience developing data service APIs, metric services, internal data products, or lightweight backend services
  • Experience with data quality frameworks, data lineage, metadata management, data catalogs, observability, or monitoring and alerting
  • AWS, GCP, or Azure cloud platform experience
  • Docker, CI/CD, Terraform, Kubernetes, or basic DevOps experience
  • Experience with LLMs, AI Agents, code generation, automated testing, task inspection, data quality agents, or engineering efficiency tooling
  • Experience with cross-border teams, international business, supply chain, e-commerce, logistics, marketing, or customer success data scenarios
Core Competencies

Proficient in Data Engineering, Data Warehousing, and Data Pipeline Development, with strong SQL and programming skills in Java or Python. Demonstrates a solid understanding of Data Quality, Governance, and collaboration with cross-functional teams.

Highest-signal resume keywords
  • Data Ingestion
  • Data Pipeline Development
  • SQL Proficiency
  • Data Quality Management
  • Data Governance
ATS Optimization Keywords
Hard Skills
  • Data Engineering
  • Data Warehousing
  • Data Modeling
  • SQL
  • Java
  • Python
  • Data Quality Rules
  • Data Validation
  • Data Pipeline Operations
  • Data Governance
Soft Skills
  • Strong Communication
  • Collaboration
  • Problem-Solving
  • Execution Skills
  • Learning Ability
Industry Keywords
  • Data Platforms
  • SaaS Products
  • Data Quality Dimensions
  • Metadata Management
  • Data Lineage
  • Data Asset Catalog
  • Cross-Border Teams
  • Cloud Platforms
  • Data Lakehouses
  • Enterprise-Level Data Warehouses
Tools & Technologies
  • SeaTunnel
  • Kafka
  • Flink
  • Spark
  • Doris
  • DolphinScheduler
  • StreamPark
  • DataHub
  • Git
  • OLAP Engines
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Jobtailor • Chicago (IL)

On-site
USD 130,000 - 180,000
Staff Data Engineer
Staff Data Engineer

Jobtailor • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Staff Engineer – Data Engineering
Staff Engineer – Data Engineering

Jobtailor • Arizona

On-site
USD 140,000 - 190,000
Data Engineering Team Lead
Data Engineering Team Lead

Jobtailor • New York (NY)

On-site
USD 170,000 - 230,000
Full-Stack Software Engineer, Platform Data
Full-Stack Software Engineer, Platform Data

Jobtailor • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New York (NY)

On-site
USD 140,000 - 190,000
Data Engineering Manager
Data Engineering Manager

Jobtailor • Colorado

On-site
USD 140,000 - 200,000
Data Platform Engineer II
Data Platform Engineer II

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Jobless • New Jersey

On-site
USD 130,000 - 170,000
Lead Business Operations Data Engineer
Lead Business Operations Data Engineer

Jobtailor • Colorado

On-site
USD 140,000 - 170,000