Senior Data Engineer

NCS Group

Singapore

On-site

SGD 120,000 - 160,000

Full time

9 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NCS Group in Singapore is seeking a Senior Data Engineer to design AI-ready data pipelines for production-scale AI, ML and RAG systems. You will own ingestion, cleaning, transformation, data quality checks and governance to enable reliable AI outputs.

Collaborate with AI Engineers and Architects to align data architecture with model needs, implement vector store strategies (pgvector, Pinecone, OpenSearch) and ensure compliance across regulated domains.

Qualifications

  • 5+ years in data engineering, delivering production-scale pipelines.
  • Strong SQL and at least one systems language (Python/Scala/Java).
  • Hands-on with batch and streaming frameworks (Airflow, Spark, Kafka).
  • Experience building data pipelines for AI/ML or RAG use cases (embeddings, vector indexing).
  • Solid data governance, PDPA-like patterns, and data residency awareness.

Responsibilities

  • Design/build AI-ready ingestion, cleaning and transformation pipelines.
  • Create batch/streaming pipelines (Airflow/Prefect/Kafka) ensuring reliable data flow.
  • Own data quality, schema validation and completeness checks upstream of models/RAG.
  • Architect RAG data ingestion/indexing pipelines and vector search infrastructure.

Skills

5+ years experience
SQL
Python/Scala/Java
Batch & streaming
AI/ML & RAG
Data governance
FDE/POC
Data architecture

Tools

Airflow
Spark
Kafka
pgvector
Pinecone
Weaviate
OpenSearch
Azure AI Search

Job description

NCS is a leading AI Tech Services company. With a 15,000-strong team across the Asia Pacific, NCS scales its platforms and capabilities to provide clients with greater agility and AI expertise across a range of Industries. Embracing a strong ecosystem of global partners, NCS transforms technology services delivery combining AI with digital resilience to drive real business impact. NCS is a subsidiary of the Singtel Group.

This is a Senior Data Engineer role within NCS AI Central’s Forward Deployed Engineering model, focused on preparing client data for production-grade AI, ML, and RAG systems.

The role owns the design and delivery of AI-ready data pipelines, including ingestion, cleaning, transformation, data quality checks, batch and streaming workflows, and production-scale operations. A key focus is building reliable data foundations for RAG and vector search, covering document ingestion, chunking, embeddings, hybrid search, and vector databases such as pgvector, Pinecone, Weaviate, OpenSearch, or Azure AI Search.

You will also handle data governance and compliance, including PII redaction, access controls, data residency, metadata, and lineage, especially for regulated sectors such as Healthcare, Government, and Transport. They will work closely with AI Engineers and AI Architects to define what “AI-ready” data means for each engagement, from fast POC/POV work through to hardened production systems.

What will you do?
1. Data Pipeline Engineering & AI-Readiness
  • Design and build ingestion, cleaning, and transformation pipelines that turn messy, real-world client data into AI-ready datasets.
  • Build batch and streaming pipelines (Airflow/Prefect/Kafka) that keep data flowing reliably into AI systems without manual intervention.
  • Own data quality — deduplication, schema validation, completeness checks — upstream of any model or RAG pipeline.
  • Proactively flag data gaps or quality issues that would degrade model/RAG performance downstream, before they surface as an AI Engineer's problem in testing.
2. RAG & Vector Store Architecture
  • Architect document/data ingestion and indexing pipelines for Retrieval-Augmented Generation (RAG) systems — chunking strategy, embeddings, hybrid/vector search.
  • Design and operate vector database and search infrastructure (pgvector/Pinecone/OpenSearch) at production scale and query volume.
3. Data Governance & Compliance
  • Implement PII redaction, data residency, and access-control patterns aligned to PDPA and sector-specific requirements (Healthcare, Government, Transport).
  • Maintain clear data lineage and metadata governance so engagement teams and auditors can trace how client data flows into AI outputs.
4. FDE & Development/Maintenance Coverage
  • During FDE engagements: rapidly assess and prepare a client's data landscape during Discover/POC, identifying data-readiness gaps early.
  • During system development & maintenance engagements: build and operate production-scale data pipelines handling the full volume and complexity of live client systems (e.g., Healthcare or Transport data at scale).
  • Contribute reusable ingestion/indexing patterns back into the shared internal asset library to accelerate future engagements.
  • Partner closely and continuously with AI Engineers and AI Architects — understanding what a given model, RAG pipeline, or agent actually needs from the data layer, and translating that into concrete pipeline and schema design decisions.
  • Own the definition of "AI-ready" data for each engagement jointly with AI Engineers — agreeing on chunking strategy, metadata, freshness, and quality thresholds before pipelines are built, not after retrieval quality suffers.
  • Sit in solution design conversations alongside AI Engineers and AI Architects, so data architecture and model/RAG architecture are designed together rather than data being treated as a downstream dependency.
  • Mentor junior data engineers and set data engineering standards across engagements.
The ideal candidate should possess:
  • 5+ years in data engineering, including production-scale pipeline design (not just analytics/reporting pipelines).
  • Strong SQL and at least one systems language (Python/Scala/Java); hands-on with batch and streaming frameworks (Airflow, Spark, Kafka).
  • Experience building data pipelines for AI/ML or RAG use cases — embeddings, vector indexing, hybrid search.
  • Solid understanding of data governance, PII handling, and access-control patterns in regulated environments.
  • Comfortable moving between fast, exploratory data assessment (FDE/POC) and disciplined, high-volume production pipeline engineering (system development & maintenance).
  • Working understanding of core AI/LLM concepts — tokenization, embeddings, chunking strategy, context windows, RAG, and agentic workflows — sufficient to hold a real technical conversation with AI Engineers and AI Architects about what "AI-ready" data means for a given use case, not just how to move and clean it.
Preferred Qualifications
  • Experience with vector databases (pgvector, Pinecone, Weaviate) and search platforms (OpenSearch/Azure AI Search).
  • Exposure to Singapore Government data environments (GCC/HCC) and compliance regimes (IM8, PDPA).
  • Experience with sector-specific data complexity — Healthcare (clinical data governance) or Transport/Aviation systems.
  • Familiarity with data cataloguing and lineage tooling.
  • Prior experience embedded within an AI/ML delivery team (not just a data platform team) — i.e., has sat alongside AI Engineers day-to-day and adjusted pipeline/schema design based on model or RAG performance feedback.
Tech Stack (Illustrative)
Why Join NCS?
Grow with Us
  • Work on cutting-edge AI products that shape the future of technology
  • Collaborate with talented, passionate teams across research, engineering, and design
  • Access continuous learning opportunities and career development pathways
Make an Impact
  • Transform AI research into products that solve real problems for clients and users
  • Drive innovation in a leading Technology Services Firm with regional presence
  • Contribute to building a better future through responsible, human-centred AI
Thrive in Our Culture
  • Experience a human-to-human approach where relationships and collaboration matter
  • Be part of Team NCS, where bold ideas meet practical execution
  • Enjoy a supportive environment that values diversity, inclusion, and respect

We are driven by our AEIOU beliefs—Adventure, Excellence, Integrity, Ownership, and Unity and we seek individuals who embody these values in both their professional and personal lives. We are committed to our Impact: Valuing our clients, Growing our people, and Creating our future.

Together, we make the extraordinary happen.

Learn more about us at ncs.co and visit our LinkedIn career site.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

#EG Senior Data Engineer
#EG Senior Data Engineer

NCS Group • Singapore

On-site
SGD 120,000 - 180,000
#EG Data & AI Architect
#EG Data & AI Architect

NCS Group • Singapore

On-site
SGD 180,000 - 250,000
Google Professional Data Engineer / ML
AWS Machine Learning / Data Analytics
Microsoft Azure AI Engineer Associate
+1
#EG Data Scientist
#EG Data Scientist

NCS Group • Singapore

On-site
SGD 90,000 - 140,000
#EG AI Engineer
#EG AI Engineer

NCS Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

NCS Group • Singapore

On-site
SGD 150,000 - 230,000
Senior Data Scientist
Senior Data Scientist

Ncs • Singapore

On-site
SGD 120,000 - 180,000
EG Senior AI Native Builder
EG Senior AI Native Builder

Ncs • Singapore

On-site
SGD 180,000 - 260,000
Cutting-edge AI products
Continuous learning
Career development
Senior AI Native Builder
Senior AI Native Builder

NCS PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
Data Engineer
Data Engineer

NCS Group • Singapore

On-site
SGD 60,000 - 90,000
#EG Data Scientist
#EG Data Scientist

NCS Pte Ltd • Singapore

On-site
SGD 90,000 - 140,000