A complete application in a minute — tailored resume and cover letter, ready to send.
Riskdata Consulting in Singapore seeks a Senior Data Engineer to design, build and optimize scalable data pipelines and AI-enabled data solutions. You will work with Spark/PySpark, SQL, Python and cloud services to deliver batch and real-time processing at scale.
You will collaborate with architects, data scientists and engineers to implement data warehouses and lakes, while enabling CI/CD and containerized deployments with Docker/Kubernetes/OpenShift.
We are looking for a highly skilled and experienced Senior Data Engineer to design, develop and optimize scalable data engineering solutions using Big Data, cloud and modern AI technologies.
The successful candidate will be responsible for developing high-performance data pipelines, data ingestion frameworks and data processing solutions, while working closely with architects, business stakeholders, data scientists and engineering teams.
Design, develop and maintain scalable batch and real-time data pipelines using Apache Spark/PySpark, SQL and Python.
Develop and optimize ETL/ELT pipelines for large-scale data processing.
Build data ingestion solutions using databases, APIs, files and streaming platforms.
Work with AWS services including S3, Glue, EMR, Redshift, Kinesis, Lambda and DynamoDB.
Develop and optimize data processing solutions using Hadoop, Hive, Spark, Kafka, Cloudera and Databricks.
Perform SQL and Spark performance tuning and optimize large-scale data processing workloads.
Design and implement data models, data warehouses and data lake solutions.
Develop CI/CD pipelines and support containerized deployments using Docker, Kubernetes and OpenShift.
Integrate Generative AI and NLP capabilities into enterprise data applications where required.
Develop solutions using LLM frameworks, RAG, vector databases and AI APIs.
Collaborate with solution architects, data scientists, software engineers and business stakeholders to deliver production-ready solutions.
Participate in system design, development, testing, deployment and production support.
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related field.
Minimum 8 years of relevant experience in Data Engineering / Big Data, with strong experience in large-scale data processing.
Strong programming experience in Python, Java and/or Scala.
Strong hands-on experience with Apache Spark/PySpark, SQL, Hadoop and Hive.
Experience with cloud platforms, particularly AWS.
Experience with data warehouses, relational databases and data lake technologies.
Experience with Kafka, Airflow, Jenkins, Git and CI/CD practices.
Experience with Docker, Kubernetes or OpenShift is advantageous.
Knowledge of Databricks, Snowflake and Cloudera is advantageous.
Exposure to Generative AI, LLMs, RAG, LangChain/LangGraph and vector databases is an advantage.
Strong analytical, problem-solving and communication skills.
Python, Java, Scala, PySpark, Apache Spark, SQL, Hadoop, Hive, Kafka, Databricks, Cloudera, AWS, S3, Glue, EMR, Redshift, Kinesis, Lambda, Airflow, Jenkins, Docker, Kubernetes, OpenShift, Snowflake, MongoDB, Oracle, PostgreSQL and Generative AI technologies.