Role Overview
We are seeking an experienced Senior Data Engineer to design, build, and optimize scalable data pipelines and platforms that support advanced analytics, reporting, and machine learning initiatives. The ideal candidate will have strong hands‑on expertise in Scala, PySpark, and distributed computing frameworks, combined with solid data engineering fundamentals and cloud exposure. You will work closely with data scientists, analysts, and product teams to deliver reliable, high‑performance data solutions in a fast‑paced, agile environment.
Key Responsibilities
- Design, develop, and maintain robust, scalable data pipelines and ETL processes.
- Build and optimize distributed data processing systems using Apache Spark, Hadoop, and related technologies.
- Develop high-performance data applications using Scala and PySpark.
- Implement and maintain data models, schemas, and data warehousing solutions.
- Optimize SQL and NoSQL queries to ensure high performance, scalability, and reliability.
- Ensure data quality, integrity, and security across data platforms.
- Build and maintain CI/CD pipelines and follow DevOps best practices using Git and automated deployment tools.
- Collaborate in Agile/Scrum teams, participating in sprint planning, reviews, and retrospectives.
- Troubleshoot production issues, perform root cause analysis, and implement long-term fixes.
- Document technical designs, workflows, and best practices.
Required Skills & Qualifications
- Strong hands‑on experience in data engineering, big data, or related roles.
- Strong proficiency in Scala and PySpark for large‑scale data processing.
- Solid understanding of distributed computing frameworks such as Apache Spark and Hadoop.
- Strong experience with SQL and NoSQL databases, data modeling, and performance tuning.
- Experience working with CI/CD pipelines, Git, and Agile development methodologies.
- Exposure to at least one cloud platform (AWS, Azure, or GCP).
- Strong analytical, problem‑solving, and debugging skills.
- Excellent verbal and written communication skills.
Preferred / Good-to-Have Skills
- Experience with streaming platforms such as Kafka, Flink, or Spark Streaming.
- Knowledge of containerization and orchestration tools (Docker, Kubernetes).
- Experience with data warehousing solutions (Snowflake, Redshift, BigQuery, Azure Synapse).
- Familiarity with workflow orchestration tools (Airflow, Argo, Luigi).
- Exposure to data governance, lineage, and security frameworks.