Senior Data Engineer

SINGAPORE TELECOMMUNICATIONS LIMITED

Singapore

On-site

SGD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Singtel is building a new AI & Data Analytics unit to empower scalable content extraction and intelligent data processing across the enterprise. The role involves designing end-to-end document parsing pipelines, OCR/VLM systems, embedding models, and RAG-enabled knowledge bases on cloud platforms, with emphasis on reliability and governance.

You will collaborate with senior engineers, implement CI/CD, and integrate tools like Kafka, LangChain, PyMuPDF, and Databricks to drive AI-led

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • 1–3 years of experience in data engineering or data platform development.
  • Ability to independently build basic batch or streaming document processing pipelines.
  • Hands-on experience with Python, PyTorch, PEFT, LoRA, OCR, VLM, LLM, prompting, Text-to-Speech models and SQL for data transformation.
  • Familiarity with Apache Spark (PySpark) and large-scale data processing concepts.
  • Experience with knowledge base and RAG solutions for agentic AI use cases.
  • Self-starter with strong problem-solving skills and collaboration.
  • Strong documentation and communication skills.

Responsibilities

  • Design, build, productionize, and support scalable document parsing and content extraction pipelines on a modern hybrid cloud platform, handling structured and unstructured data including PDFs, PPTX, Excel, Word, images, scanned documents, videos, audios, charts, tables.
  • Design and implement OCR, VLM, and multimodal AI pipelines leveraging state-of-the-art models for accurate content extraction, layout understanding, and table/chart interpretation.
  • Develop and optimize embedding models, including fine-tuning and domain adaptation to support organization-specific terminology, keywords, and semantic search use cases.
  • Implement knowledge ingestion, indexing, and retrieval pipelines for RAG and agentic GenAI systems, ensuring high recall, precision, and low-latency retrieval.
  • Build and maintain data ingestion pipelines for batch and streaming data sources using tools like Kafka, LlamaIndex, LangChain, PyMuPDF, MSAL.
  • Implement layout-aware parsing to preserve document structure, hierarchy, and reading order (using OCR, VLM, and multimodal AI solutions).
  • Extract accurate numeric and semantic data from complex charts and irregular tables while preventing hallucinations.
  • Assist in integrating data from diverse source systems (files, APIs, databases, streaming).
  • Implement, fine-tune, and optimize embedding models to support organization-specific terminology, improving semantic search and GenAI performance.
  • Help maintain metadata and pipeline documentation for transparency and traceability.
  • Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components.
  • Build and maintain knowledge base and RAG solution on cloud (Azure).
  • Implement and operate knowledge base storage, lifecycle management and embedding/vectorization.
  • Contribute to automation efforts using version control and CI/CD workflows.
  • Apply basic data governance and access control policies during implementation.

Skills

Python
PyTorch
PEFT
LoRA
OCR
VLM
LLM
Prompting
SQL
Spark/PySpark
Knowledge base
RAG
Data governance
Documentation

Education

Bachelor’s degree in Computer Science, Engineering, or a related field

Tools

Kafka
LlamaIndex
LangChain
PyMuPDF
MSAL
Microsoft Fabric
Databricks
Delta Lake
Azure

Job description

Powering the Future with AIDA

To lead the next phase of our AI evolution, we’ve launched a new business unitAIDAArtificial Intelligence & Data Analytics– a strategic engine driving our transformation designed to scale our AI ambitions with precision and purpose.This marks apivotal shiftin how we operate, innovate, and serve to embed intelligence into every layer of our business.

AtSingtel, this is more than a technology upgrade. It’s astrategic transformationthat redefines how value is created across the enterprise core—augmenting human capabilitiesand unlocking entirely new potential. It is a transformation journey by aligningpeople, platforms, and processesunder one cohesive strategy. Our mission is to buildAI literacy, and foster a culture whereintelligence empowers people.

We welcome you to join uson a transformational journey that’s reshaping the telecommunications industry — and redefining what’s possible with AI at its core.Grow with usin a workplace that championsinnovation, embracesagility, and putshuman potentialat the heart of everything we do.

Be a Part Of Something BIG!
  • Design, build, productionize, and support scalable document parsing and content extraction pipelines on a modern hybrid cloud platform, handling structured and unstructured data including PDFs, PPTX, Excel, Word, images, scanned documents, videos, audios, charts, tables etc
  • Design and implement OCR, VLM, and multimodal AI pipelines leveraging state-of-the-art models (e.g., GPT- [4/5]o, vision-language models, Text-to-Speech) for accurate content extraction, layout understanding, and table/chart interpretation
  • Develop and optimize embedding models, including fine-tuning and domain adaptation to support organization-specific terminology, keywords, and semantic search use cases
  • Implement knowledge ingestion, indexing, and retrieval pipelines for RAG (Retrieval-Augmented Generation) and agentic GenAI systems, ensuring high recall, precision, and low-latency retrieval
Make an Impact by:
  • Build and maintain data ingestion pipelines for batch and streaming data sources using tools/library like Kafka, LlamaIndex, LangChain, PyMuPDF, MSAL etc.
  • Implement layout-aware parsing to accurately preserve document structure, hierarchy, and reading order (using OCR, VLM, and multimodal AI solutions).
  • Extract accurate numeric and semantic data from complex charts and irregular tables while preventing hallucinations.
  • Assist in integrating data from diverse source systems (files, APIs, databases, streaming)
  • Implement, fine-tune, and optimize embedding models to support organization-specific terminology, improving semantic search, retrieval accuracy, and downstream GenAI performance.
  • Help maintain metadata and pipeline documentation for transparency and traceability
  • Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components
  • Build and maintain knowledge base and RAG solution on cloud (Azure)
  • Implement and operate knowledge base storage, lifecycle management and embedding/vectorization
  • Contribute to automation efforts using version control and CI/CD workflows
  • Apply basic data governance and access control policies during implementation.
Skills for Success:
  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 1–3 years of experience in data engineering or data platform development. Fresh graduates encouraged to apply as well.
  • Proven ability to independently build basic batch or streaming document processing pipelines
  • Hands-on experience with Python, PyTorch, PEFT, LoRA, OCR, VLM, LLM, prompting, Text-to-Speech models and SQL for data transformation and validation
  • Familiarity with Apache Spark (especially PySpark) and large-scale data processing concepts
  • Experience with implementing knowledge base and RAG solutions for agentic AI use cases
  • Self-starter with strong problem-solving skills and a keen attention to detail
  • Able to work independently while collaborating effectively with senior engineers and other stakeholders
  • Strong documentation and communication skills

Are you ready to say hello to BIG Possibilities?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Data Engineer
Senior AI Data Engineer

Singtel Group • Singapore

On-site
SGD 110,000 - 170,000
Senior AI Engineer: Enterprise AI & Production Systems
Senior AI Engineer: Enterprise AI & Production Systems

Singtel • Singapore

On-site
Confidential
Senior AI Engineer- #AIDA
Senior AI Engineer- #AIDA

Singapore Telecommunications Limited • Singapore

On-site
SGD 120,000 - 180,000
AI Data Engineer (Document Intelligence & GenAI)
AI Data Engineer (Document Intelligence & GenAI)

Krisvconsulting Services Pte Ltd • Singapore

On-site
SGD 60,000 - 100,000
Lead/Senior AI Engineer- #AIDA
Lead/Senior AI Engineer- #AIDA

SINGAPORE TELECOMMUNICATIONS LIMITED • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Architect: Enterprise AI Solutions
Senior AI Architect: Enterprise AI Solutions

Singtel • Singapore

Hybrid
Confidential
Senior AI Engineer- #AIDA
Senior AI Engineer- #AIDA

Singtel Group • Singapore

On-site
SGD 70,000 - 100,000
Senior AI Engineer
Senior AI Engineer

Singtel Group • Singapore

On-site
SGD 70,000 - 100,000
Senior AI Engineer- #AIDA (Singapore, Singapore)
Senior AI Engineer- #AIDA (Singapore, Singapore)

wimatec MATTES GmbH • Singapore

On-site
SGD 70,000 - 90,000
Data & AI Operations Engineer #AIDA
Data & AI Operations Engineer #AIDA

Singtel Group • Singapore

On-site
SGD 60,000 - 90,000