Overview
Design and develop robust ingestion, transformation, and enrichment pipelines using Python, PySpark, and SQL. Collaborate with CEFS data architects, data scientists, and business analysts to translate functional requirements into technical specifications.
Responsibilities
- Design & develop robust ingestion, transformation, and enrichment pipelines with Python, PySpark, and SQL
- Write and optimize complex SQL queries, analytical UDFs, and window functions for data aggregation and reporting
- Collaborate with CEFS data architects, data scientists, and business analysts to translate functional requirements into technical specifications
- Unit-test, integrate-test, and review code
- Maintain CI/CD pipelines (Git, Jenkins, Docker) for automated build, test, and deployment of jobs
- Monitor production workloads and troubleshoot performance bottlenecks, memory issues, and job failures
- Document data lineage, pipeline design, and operational run-books in Confluence/SharePoint
- Keep up to date with latest technologies and trends and provide input, expertise and recommendations
Contributing Responsibilities
- Contribute towards innovation (e.g. AI/ML); suggest new technical practices for efficiency improvement
- Participate in Agile ceremonies (sprint planning, daily standups, retrospectives) and help groom the backlogs
- Mentor junior engineers and champion best practices in Python coding, Spark optimization, and data-engineering patterns
- Evaluate emerging technologies and deliver proof-of-concepts for CEFS
Technical & Behavioral Competencies
- Resourceful to quickly understand complexities involved and provide the way forward
- Experience in technical analysis of n-tier applications with multiple integrations using object oriented, APIs & Microservices approaches
- Strong knowledge about design patterns and development principles
- Inclination and prior experience of working across SQL, Python and ETL
- Hands-on experience with Python (NumPy, pandas, Python Frameworks, Restful APIs, MS-SQL or Oracle)
- PySpark - DataFrames, Spark SQL, Structured Streaming, performance tuning (partitioning, caching, broadcast joins)
- Advanced SQL - complex queries, stored procedures, query optimization
- Knowledge of Python packages such as Pandas, NumPy for data cleaning, wrangling, analysis and visualization
- Experience in development and maintenance of code/scripts across applications, debugging, and production support
- Knowledge of Linux/Unix environment (basic commands, shell scripting), testing, documentation and new framework
- Experience with build tools like Maven and DevOps tools like Bitbucket, Jenkins
- Knowledge of Agile, Scrum, DevOps
- Development experience in a Data Engineering environment
- Willingness to learn and work on diverse technologies
- Self-motivated with good interpersonal skills and a drive to upgrade technologies
- Good communication and coordination skills
Specific Qualifications
- Good to have knowledge of front-end technologies, preferably Flask
Skills Referential
- Technical Skills:
- Behavioral Skills:
- Ability to synthesize / simplify
- Ability to collaborate / teamwork
- Attention to detail / rigor
- Ability to deliver / results driven
Education Level: Bachelor\'s degree or equivalent
Location: Chennai