Data Scientist

REGTECH INSIGHT PTE. LTD.

Singapore

On-site

SGD 120,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

REGTECH INSIGHT PTE. LTD. in Singapore seeks a Data Engineer to design, develop and maintain scalable ETL/ELT pipelines for large datasets.

You will use Python, PySpark, Spark SQL and Scala to process structured and unstructured data, and build batch and real-time pipelines with Kafka, Kinesis, Spark Streaming and Airflow. You will design data lakes and data warehouses on cloud platforms, collaborating with data scientists and product teams to deliver enterprise data solutions.

Qualifications

  • Bachelor's or master's degree in computer science or related discipline.
  • Strong professional experience in Big Data, Cloud Computing or related technology domains.
  • Experience working with large-scale enterprise data platforms and production data pipelines.
  • Strong programming and SQL skills.
  • Experience with cloud-based data engineering and modern data processing frameworks.
  • Experience with AI/ML, NLP or Generative AI will be highly advantageous.

Responsibilities

  • Design, develop and maintain scalable ETL/ELT data pipelines for large-volume datasets.
  • Develop high-performance data processing solutions using Python, PySpark, Spark SQL and Scala.
  • Build batch and real-time data pipelines using Kafka, Kinesis, Spark Streaming, AWS Glue and Airflow.
  • Develop data ingestion frameworks integrating RDBMS, APIs, files, cloud storage and streaming platforms.
  • Design and implement data lakes, data warehouses and cloud-based data processing platforms.
  • Work with Databricks, Hadoop, Hive, Cloudera, Presto and Snowflake for large-scale data processing and analytics.
  • Perform data modelling, data transformation, data quality, query optimisation and performance tuning.
  • Develop and optimise SQL solutions across Oracle, SQL Server, PostgreSQL, Teradata, MongoDB and cloud databases.
  • Design and implement data migration solutions involving large-scale enterprise datasets.
  • Develop and support real-time and batch processing architectures for enterprise applications.
  • Integrate data platforms with REST APIs, GraphQL and enterprise applications.
  • Implement CI/CD and DevOps practices using Jenkins, Git, Docker, Kubernetes and OpenShift.
  • Develop cloud-native data solutions using AWS and Azure, including S3, Glue, EMR, Redshift, Kinesis, Lambda, RDS and DynamoDB.
  • Develop AI/GenAI-enabled data solutions involving LLMs, NLP, RAG, Agentic AI and vector databases.
  • Integrate LLM services and AI platforms such as Azure OpenAI, OpenAI APIs, Hugging Face and Google Gemini/ADK.
  • Develop NLP pipelines for text processing, embeddings, summarisation, sentiment analysis, voice-to-text and speaker diarisation.
  • Design and implement vector search and retrieval solutions using Redis, ChromaDB and FAISS.
  • Develop AI-powered APIs and applications using FastAPI, Gradio and Python.
  • Collaborate with architects, data scientists, software engineers, business analysts and product teams to deliver enterprise data solutions.
  • Participate in Agile SDLC activities including requirements analysis, architecture, development, testing, deployment and production support.
  • Troubleshoot complex data, application and platform issues and provide scalable technical solutions.

Skills

Python
PySpark
Apache Spark
Spark SQL
Scala
Hadoop
Hive
Kafka
Presto
Databricks
Cloudera
Snowflake
Airflow
Jenkins
Docker
Kubernetes
Git
Terraform
OpenShift
OpenSearch
AWS
Azure
REST APIs
GraphQL
Power BI
SQL
Data Modelling

Education

Bachelor's or Master's in Computer Science or related

Tools

Databricks
Cloudera
Snowflake
Airflow
Jenkins
Docker
Kubernetes
Git
Terraform
OpenShift
OpenSearch

Job description

Key Responsibilities


  • Design, develop and maintain scalable ETL/ELT data pipelines for large-volume structured and unstructured datasets.

  • Develop high-performance data processing solutions using Python, PySpark, Apache Spark, Spark SQL and Scala.

  • Build batch and real-time data pipelines using Kafka, Kinesis, Spark Streaming, AWS Glue and Airflow.

  • Develop data ingestion frameworks integrating RDBMS, APIs, files, cloud storage and streaming platforms.

  • Design and implement data lakes, data warehouses and cloud-based data processing platforms.

  • Work with Databricks, Hadoop, Hive, Cloudera, Presto and Snowflake for large-scale data processing and analytics.

  • Perform data modelling, data transformation, data quality, query optimisation and performance tuning.

  • Develop and optimise SQL solutions across Oracle, SQL Server, PostgreSQL, Teradata, MongoDB and cloud databases.

  • Design and implement data migration solutions involving large-scale enterprise datasets.

  • Develop and support real-time and batch processing architectures for enterprise applications.

  • Integrate data platforms with REST APIs, GraphQL and enterprise applications.

  • Implement CI/CD and DevOps practices using Jenkins, Git, Docker, Kubernetes and OpenShift.

  • Develop cloud-native data solutions using AWS and Azure, including S3, Glue, EMR, Redshift, Kinesis, Lambda, RDS and DynamoDB.

  • Develop AI/GenAI-enabled data solutions involving LLMs, NLP, RAG, Agentic AI and vector databases.

  • Integrate LLM services and AI platforms such as Azure OpenAI, OpenAI APIs, Hugging Face and Google Gemini/ADK.

  • Develop NLP pipelines for text processing, embeddings, summarisation, sentiment analysis, voice-to-text and speaker diarisation.

  • Design and implement vector search and retrieval solutions using Redis, ChromaDB and FAISS.

  • Develop AI-powered APIs and applications using FastAPI, Gradio and Python.

  • Collaborate with architects, data scientists, software engineers, business analysts and product teams to deliver enterprise data solutions.

  • Participate in Agile SDLC activities including requirements analysis, architecture, development, testing, deployment and production support.

  • Troubleshoot complex data, application and platform issues and provide scalable technical solutions.


Required Technical Skills

Data Engineering:
Python, PySpark, Apache Spark, Spark SQL, Scala, Hadoop, Hive, Kafka, Presto, Databricks, Cloudera, Snowflake


Cloud Technologies:
AWS, Azure, S3, Glue, EMR, Redshift, Kinesis, Lambda, RDS, DynamoDB, OpenSearch


Databases:
SQL Server, Oracle, PostgreSQL, Teradata, MongoDB, Redis


Programming:
Python, Java, Scala, SQL, Shell Scripting, Node.js


AI / GenAI / NLP:
Generative AI, LLM, NLP, RAG, Agentic RAG, LangChain, LangGraph, LlamaIndex, Hugging Face Transformers, Azure OpenAI, OpenAI API, Google Gemini/ADK, PyTorch


Vector & AI Search:
Redis Vector Database, ChromaDB, FAISS, Embeddings, Hybrid Search, Semantic Search


DevOps & Deployment:
Docker, Kubernetes, OpenShift, Jenkins, Git, Terraform, CI/CD


Data Integration & APIs:
REST APIs, GraphQL, FastAPI, API Gateway, AWS Lambda, CDC, Debezium


Data Visualisation:
Power BI, Data Modelling, Reporting and Analytics


Qualifications


  • Bachelor's or master's degree in computer science, Information Technology, Programming & Systems Analysis, Computer Studies or a related discipline.

  • Strong professional experience in Big Data, Cloud Computing or related technology domains.

  • Experience working with large-scale enterprise data platforms and production data pipelines.

  • Strong programming and SQL skills.

  • Experience with cloud-based data engineering and modern data processing frameworks.

  • Experience with AI/ML, NLP or Generative AI will be highly advantageous.


Preferred Experience


  • Enterprise Banking / Financial Services experience.

  • Experience working with large-scale customer, transaction or financial datasets.

  • Experience with data migration and legacy ETL modernization.

  • Experience implementing AI/GenAI solutions within enterprise data platforms.

  • Experience with production deployments and CI/CD environments.

  • Strong understanding of data governance, security, data quality and performance optimization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Riskdata Consulting • Singapore

On-site
SGD 90,000 - 170,000
Data Scientist
Data Scientist

RISKDATA CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Data Scientist
Data Scientist

UNISYNC SYSTEMS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Data Scientist
Data Scientist

JEET ANALYTICS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Data Scientist
Data Scientist

UARROW PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Data Scientist
Data Scientist

UNISONEDGE CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Data Engineering Consultant
Data Engineering Consultant

UARROW PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Data Engineer – PySpark, Databricks & Data Lakehouse
Senior Data Engineer – PySpark, Databricks & Data Lakehouse

D L Resources Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Data Engineering Consultant
Data Engineering Consultant

JEET ANALYTICS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Data & AI Solutions Specialist
Data & AI Solutions Specialist

UNISONEDGE CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000