Role:- Data Engineer (4-7 years)
Location:- Mumbai
We are looking for an experienced Data Engineer to design, develop, and maintain scalable data pipelines and data platforms supporting Analytics, AI/ML, and Business Intelligence use cases.The candidate will work closely with Data Scientists, AI Engineers, Analytics teams, SAP teams, and business stakeholders to build reliable, secure, and high-performance data solutions using AWS Cloud and SAP S/4HANA.
Qualification
- Bachelor's/Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- 4+ years of relevant Data Engineering experience preferred.
- Hands-on experience with AWS Cloud and SAP S/4HANA integration is strongly preferred.
Success Measures
The Data Engineer will be evaluated on:
- Reliability and scalability of data pipelines.
- Data quality and completeness.
- Successful integration of SAP S/4HANA with AWS.
- Pipeline performance and optimization.
- Reduction in processing time and cloud infrastructure cost.
- Timely resolution of production issues.
- Reusability and maintainability of data engineering components.
- Strong data governance and security.
- Enablement of Analytics, AI/ML, and Business Intelligence use cases.
Key Responsibilities
- Design and develop scalable ETL/ELT pipelines for structured and unstructured data.
- Build data ingestion and transformation pipelines using Python, SQL, and AWS Cloud services.
- Develop data integration solutions between SAP S/4HANA and AWS Cloud.
- Work with AWS services such as Amazon S3, AWS Glue, AWS Lambda, Amazon Athena, Amazon CloudWatch, and IAM.
- Implement scalable data storage and processing architectures on AWS.
- Perform data cleansing, transformation, validation, reconciliation, and data-quality checks.
- Design and implement incremental data loading, CDC, watermarking, and SCD Type 1/Type 2 processes.
- Optimize data pipelines for performance, scalability, reliability, and cloud cost efficiency.
- Implement data governance, metadata management, data lineage, security, and access controls.
- Develop monitoring and alerting mechanisms for production data pipelines.
- Troubleshoot pipeline failures, data-quality issues, performance bottlenecks, and production incidents.
- Implement CI/CD practices for data engineering workloads using Git, GitHub, and DevOps tools.
- Collaborate with Data Scientists and AI/ML Engineers to prepare high-quality datasets and feature pipelines.
- Support integration of GenAI, RAG, and Agentic AI solutions with enterprise data where required.
- Document technical architecture, data flows, pipeline logic, data mappings, and operational procedures. Required Technical Skills
Programming
- Strong Python
- Strong SQL
- Shell/Bash scripting Data Engineering
- ETL/ELT
- Data ingestion and integration
- Data pipelines
- Data warehousing
- Data modelling
- Data quality and validation
- Incremental loading / CDC
- Watermark-based processing
- SCD Type 1 & Type 2
- Batch and near-real-time data processing
AWS Cloud
Strong hands-on experience with:
- Amazon S3
- AWS Glue
- AWS Lambda
- Amazon Athena
- Amazon CloudWatch
- AWS IAM
- AWS Step Functions / workflow orchestration
- AWS Event Bridge
- AWS Secrets Manager
- AWS CI/CD services
SAP S/4HANA
- Strong understanding of SAP S/4HANA data structures and business processes.
- Experience extracting/integrating data from SAP S/4HANA into AWS.
- Experience with OData, APIs, CDS Views, IDocs, or other relevant SAP interfaces.
- Understanding of SAP master and transactional data.
- Ability to work with SAP functional and technical teams for data integration requirements.
- Experience handling SAP data quality, reconciliation, and incremental extraction. AI/ML Data Engineering Exposure
- The candidate should be able to support modern AI/ML workloads by:Preparing high quality datasets for ML models.
- Building data pipelines for AI/ML applications.
- Processing structured and unstructured enterprise data.
- Supporting RAG pipelines and embedding-generation workflows.
- Providing governed enterprise data to AI agents and LLM applications. Implementing data validation and monitoring for AI applications.
- Supporting integration between AWS data platforms and AI/ML solutions.
Preferred Skills
- AWS Step Functions / workflow orchestration
- REST API integration
- SAP APIs / OData / CDS Views
- SAP S/4HANA integration
- Data Catalog and metadata management
- Data governance and security
- Docker, CI/CD
- Git / GitHub / DevOps, Infrastructure as Code
- Experience supporting AI/ML data pipelines
- Experience with RAG, embeddings, vector databases, or Agentic AI data workflows
- Experience in enterprise data migration and integration projects
Key Competencies
- Strong problem-solving and analytical skills.
- Strong understanding of AWS data engineering architecture.
- Good understanding of SAP S/4HANA data and integration.
- Ability to troubleshoot production issues independently.
- Strong focus on data quality, reliability, and security.
- Understanding of data governance and access controls.
- Good communication and stakeholder-management skills.
- Ability to work in an Agile/Scrum environment.
- Ability to convert business requirements into scalable technical solutions.