An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Teksystems in Gurugram is seeking a data engineer with Apache Spark and Scala expertise to build large‑scale data pipelines and analytics platforms.
You will work on data lake architectures, Iceberg/Parquet formats, metadata and CI/CD for data projects, collaborating with product, data science, and engineering teams.
Strong Python/SQL skills and cloud experience are a plus.
Role & responsibilities
JD:
Apache Spark strong hands‑on expertise in large‑scale distributed data processing; Scala – advanced programming proficiency for big data engineering
Big Data Ecosystem – strong understanding of distributed computing and data platform components; Data Lake Architecture – experience designing scalable modern data lake solutions
Apache Iceberg – working knowledge of open table formats for large‑scale analytics; Parquet – experience with columnar storage formats
Hive Metastore – metadata management experience; ETL/ELT Pipeline Development – ability to build reliable and scalable data pipelines
Data Modelling & Data Warehousing – strong foundation in analytical data structures; Performance Optimisation – tuning data processing jobs and storage usage
Data Quality Frameworks – ensuring accuracy, reliability, and governance of data; CI/CD Practices – deployment automation and engineering best practices
Python & SQL Optimisation – good‑to‑have skills for scripting and query performance; Cloud Data Platforms – good‑to‑have experience with scalable cloud‑based data platforms
Stakeholder Collaboration – working with Product Managers, business teams, Data Scientists, and Engineering teams Communication & Problem Solving – ability to translate business needs into data solutions