Senior Data Engineer AI

Anchor Search Group Pte Ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anchor Search Group Pte Ltd is seeking a Senior Data Engineer to design and operate large-scale data pipelines for AI-ready datasets across FDE engagements and production environments.

You will own ingestion, cleaning, transformation, and governance while building batch and streaming pipelines that support RAG systems and vector stores at scale. Collaboration with AI Engineers is essential to define data readiness and metadata standards.

Qualifications

  • 10+ years in data engineering with production-scale pipeline design.

Responsibilities

  • Design and build ingestion, cleaning, and transformation pipelines to AI-ready datasets.
  • Build batch and streaming pipelines (Airflow/Prefect/Kafka) for reliable data flow.
  • Own data quality and governance for upstream model/RAG pipelines.
  • Flag data gaps that could degrade model performance.
  • Collaborate with AI Engineers to align data architecture with model needs.
  • Mentor junior data engineers and contribute reusable ingestion/indexing patterns.

Skills

Strong SQL
Python/Scala/Java
Airflow/Spark/Kafka
AI/ML data pipelines
Data governance & PII handling
Production-scale pipelines
Data quality & validation

Tools

pgvector
Pinecone
Weaviate
OpenSearch
Presidio

Job description

You will operate across both fast-moving Forward Deployed Engineering (FDE) engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work — bringing the same rigor and a reusable, asset-fed approach to both.

Responsibilities
Data Pipeline Engineering & AI-Readiness
  • Design and build ingestion, cleaning, and transformation pipelines that turn messy, real-world client data into AI-ready datasets.
  • Build batch and streaming pipelines (Airflow/Prefect/Kafka) that keep data flowing reliably into AI systems without manual intervention.
  • Own data quality — deduplication, schema validation, completeness checks — upstream of any model or RAG pipeline.
  • Proactively flag data gaps or quality issues that would degrade model/RAG performance downstream, before they surface as an AI Engineer's problem in testing.
RAG & Vector Store Architecture
  • Architect document/data ingestion and indexing pipelines for Retrieval-Augmented Generation (RAG) systems — chunking strategy, embeddings, hybrid/vector search.
  • Design and operate vector database and search infrastructure (pgvector/Pinecone/OpenSearch) at production scale and query volume.
Data Governance &Compliance
  • Implement PII redaction, data residency, and access-control patterns aligned to PDPA and sector-specific requirements (Healthcare, Government, Transport).
  • Maintain clear data lineage and metadata governance so engagement teams and auditors can trace how client data flows into AI outputs.
FDE &Development/Maintenance Coverage
  • During FDE engagements: rapidly assess and prepare a client's data landscape during Discover/POC, identifying data-readiness gaps early.
  • During system development & maintenance engagements: build and operate production-scale data pipelines handling the full volume and complexity of live client systems (e.g., Healthcare or Transport data at scale).
  • Contribute reusable ingestion/indexing patterns back into the shared internal asset library to accelerate future engagements.
Collaboration &Leadership
  • Partner closely and continuously with AI Engineers and AI Architects — understanding what a given model, RAG pipeline, or agent actually needs from the data layer and translating that into concrete pipeline and schema design decisions.
  • Own the definition of "AI-ready" data for each engagement jointly with AI Engineers — agreeing on chunking strategy, metadata, freshness, and quality thresholds before pipelines are built, not after retrieval quality suffers.
  • Sit in solution design conversations alongside AI Engineers and AI Architects, so data architecture and model/RAG architecture are designed together rather than data being treated as a downstream dependency.
  • Mentor junior data engineers and set data engineering standards across engagements.
Requirements
  • 10+years in data engineering, including production-scale pipeline design (not just analytics/reporting pipelines).
  • Strong SQL and at least one systems language (Python/Scala/Java); hands-on with batch and streaming frameworks (Airflow, Spark, Kafka).
  • Experience building data pipelines for AI/ML or RAG use cases — embeddings, vector indexing, hybrid search.
  • Solid understanding of data governance, PII handling, and access-control patterns in regulated environments.
  • Comfortable moving between fast, exploratory data assessment (FDE/POC) and disciplined, high-volume production pipeline engineering (system development & maintenance).
  • Working understanding of core AI/LLM concepts — tokenization, embeddings, chunking strategy, context windows, RAG, and agentic workflows — sufficient to hold areal technical conversation with AI Engineers and AI Architects about what "AI-ready" data means for a given use case, not just how to move and clean it.
Preferred Qualifications
  • Experience with vector databases (pgvector, Pinecone, Weaviate) and search platforms(OpenSearch/Azure AI Search).
  • Exposure to Singapore Government data environments (GCC/HCC) and compliance regimes(IM8, PDPA).
  • Experience with sector-specific data complexity — Healthcare (clinical data governance) or Transport/Aviation systems.
  • Familiarity with data cataloguing and lineage tooling.
  • Prior experience embedded within an AI/ML delivery team (not just a data platform team) — i.e., has sat alongside AI Engineers day-to-day and adjusted pipeline/schema design based on model or RAG performance feedback.
Tech Stack(Illustrative)
  • Languages: Python, SQL (Scala/Java a plus)
  • Pipelines: Airflow/Prefect, Spark, Kafka/Debezium
  • Storage/Search: Postgres, S3/Blob, pgvector/Pinecone/Weaviate, OpenSearch/Azure AI Search
  • Governance: Presidio (PII redaction), data catalogue/lineage tooling
  • Cloud: AWS/Azure/GCP; GCC/HCC exposure a plus
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer/Data Engineer
Senior Data Engineer/Data Engineer

apba tg human resource pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
#EG Senior Data Engineer
#EG Senior Data Engineer

NCS • Singapore

On-site
SGD 120,000 - 180,000
AI Data Engineer
AI Data Engineer

UOB KAY HIAN PRIVATE LIMITED • Singapore

Hybrid
SGD 120,000 - 190,000
AI Data Engineer
AI Data Engineer

UOB Kay Hian Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Data Engineer
Data Engineer

AVENSYS CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Competitive base salary
Career progression
Fun collaborative environment
Senior AI Engineer
Senior AI Engineer

NCS PTE. LTD. • Singapore

On-site
SGD 140,000 - 210,000
AI Engineer - Agentic & GenAI Systems
AI Engineer - Agentic & GenAI Systems

JOY CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Data Scientist
Senior Data Scientist

digital biz solutions pte. ltd. • Singapore

On-site
SGD 180,000 - 280,000
Data Scientist
Data Scientist

Riskdata Consulting • Singapore

On-site
SGD 90,000 - 170,000
Senior AI Data Engineer
Senior AI Data Engineer

jobline resources pte. ltd. • Singapore

On-site
SGD 90,000 - 180,000