AWS Data Engineer - G Noida

Coforge

Hyderabad, Pune District, Delhi, Gurugram District

Hybrid

INR 1,200,000 - 1,900,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Coforge is seeking a skilled Data Engineer to design, develop and maintain robust ETL/ELT pipelines on Databricks using Python, PySpark and Spark SQL. You will translate business requirements into technical designs and drive delivery with minimal handholding.

You will work with AWS services, implement data quality checks, governance with Unity Catalog and contribute to CI/CD. Experience with SAP S/4HANA and Salesforce integration is a plus.

Qualifications

  • Bachelor's degree in Computer Science, Engineering or a related field (Master's preferred).
  • Databricks Certified Data Engineer (Associate/Professional) and AWS Certified Data Engineer / Solutions Architect are advantageous.

Responsibilities

  • Design, develop and maintain ETL/ELT pipelines on Databricks with Python, PySpark and Spark SQL.
  • Analyze requirements and translate into technical designs with minimal handholding.
  • Design and implement conceptual, logical and physical data models for analytics and master data use cases.
  • Ingest data from SAP S/4HANA, Salesforce, flat files and APIs into the Lakehouse.
  • Develop and operate workloads on AWS (S3, Lambda, IAM, networking) integrated with Databricks.
  • Implement data quality checks, validation rules and error-handling frameworks; prepare test scenarios and validate deliverables.
  • Apply Unity Catalog governance: access controls, lineage, cataloguing and data handling.
  • Optimize Spark jobs and Delta tables for performance and cost.
  • Contribute to CI/CD and produce technical documentation.
  • Support BI and MDM workstreams with curated datasets for Power BI and Informatica IDMC.
  • Provide L2/L3 support and root-cause analysis for pipelines.

Skills

Python
SQL
Apache Spark / PySpark
Data Modelling
Databricks
AWS
Orchestration tools
CI/CD
SAP S/4HANA integration
Power BI datasets
Streaming (Kafka/Kinesis)
Git
Azure DevOps
Delta Lake
Delta Live Tables
Unity Catalog
APIs

Education

Bachelor's degree in Computer Science, Engineering or related field
Master's degree preferred

Tools

Databricks
Delta Lake
Delta Live Tables/Jobs
Unity Catalog
Spark SQL
Git
Azure DevOps
AWS Glue
Power BI

Job description

2. Key Responsibilities

  • Design, develop and maintain robust, scalable ETL/ELT pipelines on Databricks using Python, PySpark and Spark SQL, following medallion (Bronze/Silver/Gold) architecture principles.
  • Independently analyze business requirements, translate them into technical designs and solution options, and drive them to completion with minimal handholding.
  • Design and implement conceptual, logical and physical data models (dimensional and 3NF) for analytics, reporting and master data use cases.
  • Build and manage data ingestion from enterprise sources (SAP S/4HANA, Salesforce, flat files, APIs, streaming) into the Lakehouse.
  • Develop and operate workloads on AWS (S3, Lambda, IAM, networking fundamentals) integrated with the Databricks platform.
  • Implement data quality checks, validation rules, reconciliation and error-handling frameworks; prepare test scenarios and perform thorough unit and business-level validation of own deliverables before handover.
  • Apply Unity Catalog-based governance: access controls, lineage, cataloguing and aligned handling of personal data.
  • Optimize Spark jobs and Delta tables for performance and cost (partitioning, Z-ordering, cluster sizing, job orchestration).
  • Contribute to CI/CD practices for data pipelines (Git-based version control, code reviews, automated deployment) and produce clear technical documentation.
  • Support BI and MDM workstreams by delivering curated, well-modelled datasets for Power BI and Informatica IDMC consumption.
  • Provide L2/L3 support for deployed pipelines, troubleshoot production issues and drive root-cause resolution.

3. Primary (Must-Have) Skills

  • Python: Strong, hands-on programming for data engineering: clean, modular, well-tested code.
  • SQL: Advanced SQL for complex transformations, analysis and performance optimization.
  • Apache Spark / PySpark: Deep understanding of Spark architecture, Data frame API, Spark SQL, performance tuning and debugging of distributed jobs.
  • Data Modelling: Solid experience in dimensional modelling (star/snowflake), normalized modelling, slowly changing dimensions and canonical/master data models.
  • Databricks: Proven project experience with Databricks workspaces, Delta Lake, Delta Live Tables/Jobs, Workflows and Unity Catalog.
  • AWS: Working proficiency with core AWS data services (S3, Glue, Lambda, IAM, CloudWatch) and integration with Databricks.

4. Other Required Skills

  • Experience with orchestration tools (Databricks Workflows or equivalent).
  • Git-based development workflow, code review discipline and CI/CD for data pipelines (Azure DevOps or similar).
  • Exposure to integrating with enterprise systems such as SAP S/4HANA and Salesforce (APIs, CDC, extractors) is highly desirable.
  • Familiarity with Power BI datasets/semantic models and with MDM concepts (e.g., Informatica IDMC) is an advantage.
  • Exposure to streaming technologies (Structured Streaming, Kafka/Kinesis) is a plus.

5. Ways of Working & Behavioral Expectations

  • Independent delivery: Works independently end-to-end: clarifies requirements early, proposes designs, and delivers complete, tested solutions without repeated review cycles.
  • Ownership: Takes accountability for quality and timelines; proactively prepares test scenarios and validates business logic before submitting work.
  • Speed of comprehension: Grasps new requirements and domain logic quickly and converts them into working solutions at the pace the project demands.
  • Communication: Communicates progress, risks and blockers clearly and early; documents solutions to a standard others can maintain.
  • Teamwork: Collaborates effectively with BI developers, MDM consultants, business SMEs and vendor teams in a multi-vendor environment.

6. Qualifications

  • Bachelor's degree in Computer Science, Engineering or a related field (Master's preferred).
  • Relevant certifications are an advantage: Databricks Certified Data Engineer (Associate/Professional), AWS Certified Data Engineer / Solutions Architect.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Aws Data Engineer
Aws Data Engineer

Tata Consultancy Services • Gandhinagar, Pune District, Chennai District

On-site
INR 1,200,000 - 1,800,000
Data Engineer (AWS + Databricks)
Data Engineer (AWS + Databricks)

Alike Thoughts • India

Hybrid
INR 1,200,000 - 2,400,000
Technical Data Engineer
Technical Data Engineer

Moptra Infotech • Dadri, Ghaziabad District, Delhi

Hybrid
INR 2,500,000 - 4,500,000
Data Engineer- Databricks
Data Engineer- Databricks

r3 Consultant • Bengaluru

On-site
INR 1,000,000 - 2,000,000
AWS Data Lead Engineer
AWS Data Lead Engineer

Sonata Software • Hyderabad, Chennai District, Bengaluru

Hybrid
INR 4,000,000 - 9,000,000
AWS certifications encouraged
GenAI data foundations exposure
AWS DATA ENGINEER
AWS DATA ENGINEER

NITYO • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Databricks Data Engineering Architect
Databricks Data Engineering Architect

Blue Cloud Softech Solutions Limited • India

On-site
INR 3,500,000 - 5,500,000
VS01472 - Data Migration Engineer
VS01472 - Data Migration Engineer

E4 Software Services Pvt Ltd. • Mumbai, Pune District, Bengaluru

On-site
INR 1,000,000 - 1,500,000
Aws Data Engineer
Aws Data Engineer

Tekskills • Gurugram District

On-site
INR 1,800,000 - 3,000,000
VS01472 - Data Migration Engineer
VS01472 - Data Migration Engineer

E4 Software Services Pvt Ltd. • Mumbai, Pune District, Bengaluru

On-site
INR 1,000,000 - 1,500,000