Principal Engineer - Data Engineering

Western Digital

Singapore

On-site

SGD 150,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Western Digital is hiring for a data engineering role focused on ML-ready scientific data pipelines. You will build feature stores, version data, and ensure data governance for AI workflows in a scalable cloud environment.

The role emphasizes end-to-end pipeline engineering, data quality, and collaboration with ML engineers to enable robust model training and deployment at scale in Singapore.

Qualifications

  • Strong SQL skills with advanced data transformations.
  • Experience designing batch pipelines with reliability and fault tolerance.
  • Knowledge of data quality checks, schema validation and distribution monitoring.
  • Ability to trace data origin and transformations for reproducibility of training data.
  • Awareness of how data pipelines feed ML model training and evaluation.

Responsibilities

  • Build and maintain reliable, versioned data pipelines for ML-ready features.
  • Design data quality checks across datasets and alert ML team when issues arise.
  • Implement training data versioning and lineage for reproducibility.
  • Establish data contracts and governance with access control and retention policies.
  • Contribute to real-time sensor data ingestion and synthetic data pipeline workstreams.
  • Maintain ML data layer components tightly integrated with AI platform.

Skills

SQL proficiency
Batch pipeline design
Data quality principles
Data versioning & lineage
ML data lifecycle awareness

Education

Bachelor's or Master’s degree in AI/CS/Data Eng

Tools

Airflow
Prefect
AWS Glue Jobs

Job description

Company Description

WD is building the infrastructure behind the AI-driven data economy.

As AI scales, so does data. Every interaction, every model, every system generates data that must be stored, managed, and made accessible over time. That's where we come in.

We combine deep engineering expertise with global-scale manufacturing to deliver the storage systems that make AI possible, powering hyperscale data centers, cloud platforms, and enterprise infrastructure worldwide.

This isn't theoretical work. It's real systems, at real scale, people solving some of the hardest challenges in technology today.

We're looking for peoplewho want to build, solve, and operate at that level.

Join us and let's shape the future of data.

Job Description
About This Role - The Mission

The data you will build pipelines for is not transactional data or clickstream data. It is experimental measurement data from precision product development instruments - each data point costs real time and resources to generate. Getting the data infrastructure right for this kind of scientific data is a genuinely different engineering challenge from standard web-scale or financial data work. You will develop rare expertise in ML-ready scientific data pipelines that very few data engineers in Singapore or globally have built.

Key Responsibilities
  • Feature Engineering Pipelines: Build and maintain reliable, versioned feature engineering pipelines that transform raw engineering, sensor, and operational data into structured ML-ready feature sets - delivered to the specification defined
  • Data Quality Frameworks: Design and operate data quality checks covering completeness, schema consistency, statistical distribution stability, and label accuracy across all AI training datasets. Alert the ML team when data quality degrades before it impacts model training. Collaborate with team who performs final downstream validation.
  • Data Versioning, Lineage & Drift Detection: Build and maintain training data versioning and lineage tracking - ensuring full reproducibility of all model training runs and early alerting when deployment data diverges from training distributions.
  • Data Contracts & Governance - Guided Implementation: Implement and maintain agreed data contracts between upstream data producers and downstream ML consumers, following governance standards established with guidance from ML Engineer. Establish access control and retention practices for all AI data assets.
  • Real-Time Streaming - Sensor Data Ingestion: Contribute to real-time sensor data ingestion pipelines under technical direction. Develops operational ownership progressively over 6-12 months. Not a solo day-1 requirement.
  • Synthetic Data Pipeline Support: Build pipeline infrastructure to operationalize synthetic data generation workstreams. With generative model methodology provided, builds ingestion, storage, and versioning infrastructure.
  • MLOps Data Layer:Build and maintain the training dataset registry, feature store, and model input validation - tightly integrated with the AI platform (AWS Kubernetes, PortKey, Agent Gateway, LangFuse, AWS Guardrails, Elastic Search etc.).
Qualifications
Requirements:
Education
  • Bachelor's orMaster's degreein AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or related field. AI major preferred; strong data engineering fundamentals required.
Experience
  • Fresh to 1 year.Demonstrated project experience building end-to-end data pipelines - academic, personal, or internship contexts - is the primary evaluation criterion. Python, SQL, and pipeline design fundamentals must be solid and demonstrable through project evidence.
Must Have Skills:
  • SQL:Strong proficiency - complex queries, window functions, data transformation logic
  • Scalable Pipeline Design:Batch pipeline architecture; reliability, schema management, fault tolerance; Pipeline orchestration (Airflow, Prefect, AWS Glue Jobs)
  • Data Quality Principles:Completeness checks, schema validation, distribution stability monitoring
  • Data Versioning & Lineage:Reproducibility of training data; ability to trace data origin and transformations
  • ML Data Lifecycle Awareness:Basic understanding of how data pipelines connect to ML model training. Awareness that data quality and pipeline design affect model performance downstream - specifically, awareness of risks like train/test data leakage and label quality impact on model accuracy. Does not require prior ML work experience; requires curiosity and conceptual understanding.
Good to have Skills:
Additional Information

#LI-FN1

WD thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

WD is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at jobs.accommodations@wdc.com to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

WD thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

WD is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at jobs.accommodations@wdc.com to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Engineer - Data Engineering
Principal Engineer - Data Engineering

Western Digital Corporation • Singapore

Hybrid
SGD 60,000 - 100,000
Principal Engineer - Data Engineering
Principal Engineer - Data Engineering

WD • Singapore

On-site
SGD 55,000 - 85,000
Principal Engineer - Data Engineering
Principal Engineer - Data Engineering

WD Media (Singapore) Pte Ltd • Singapore

On-site
SGD 48,000 - 80,000
Staff Engineer - Machine Learning
Staff Engineer - Machine Learning

Western Digital Corporation • Singapore

Hybrid
SGD 48,000 - 78,000
Staff Engineer - Machine Learning
Staff Engineer - Machine Learning

Western Digital • Singapore

On-site
SGD 60,000 - 90,000
Principal Engineer - Machine Learning
Principal Engineer - Machine Learning

WD Media (Singapore) Pte Ltd • Singapore

On-site
SGD 90,000 - 120,000
Staff Engineer - Machine Learning
Staff Engineer - Machine Learning

WD • Singapore

On-site
SGD 60,000 - 90,000
Principal Engineer - Machine Learning
Principal Engineer - Machine Learning

WD • Singapore

On-site
SGD 120,000 - 180,000
Principal Engineer - Machine Learning
Principal Engineer - Machine Learning

Western Digital • Singapore

On-site
SGD 90,000 - 130,000
Senior Technologist - Machine Learning
Senior Technologist - Machine Learning

Western Digital Corporation • Singapore

Hybrid
SGD 180,000 - 240,000