Data Engineer

GMG

Gurugram District

On-site

INR 2,000,000 - 3,800,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMG is seeking a highly skilled Data Engineer with a strong focus on AWS and Databricks to design, build, and maintain scalable data pipelines. The role involves ingestion of data from multiple sources, including Google Analytics, and optimizing performance and cost across AWS services and Databricks environments.

The candidate should have hands-on experience with Glue, PySpark, SQL, Athena, Lambda, SNS, and S3, along with CI/CD practices and governance standards.

Qualifications

  • Minimum 6 years in data engineering (core development/design).
  • 3+ years on AWS with hands-on AWS Glue, PySpark, SQL, Athena, Lambda, SNS, S3.
  • 2+ years on Databricks.
  • Experience with testing, CI/CD and automation.

Responsibilities

  • Develop and manage scalable data pipelines for structured, semi-structured, and unstructured data using AWS Glue, PySpark, and SQL.
  • Handle real-time event stream data ingestion and processing from multiple source systems.
  • Ensure efficient data integration into Databricks for advanced processing and analytics.
  • Build and optimize backend systems leveraging AWS services (Glue, Athena, Lambda, SNS, S3).
  • Implement, configure, and manage Databricks environments, including clusters, notebooks, and libraries for performance optimization.
  • Ensure optimal resource utilization for AWS and Databricks clusters to reduce costs.
  • Collaborate across cross-functional teams to achieve organizational goals.

Skills

AWS
Databricks
PySpark
SQL
AWS Glue
Lambda
Athena
Redshift
S3
SNS
CI/CD
Testing
Security
Governance
Performance tuning
Cost optimization
Data pipelines
Google Analytics

Education

Bachelor’s degree in a related field

Tools

Databricks
AWS Glue
PySpark
SQL

Job description

We are seeking a highly skilled Data Engineer specializing in AWS and Databricks. The ideal candidate will design, build, and maintain scalable data pipelines, ensuring efficient data ingestion, processing, and integration from multiple sources—including Google Analytics event data. This role requires deep expertise in AWS Glue, Lambda, Athena, Redshift, Databricks, PySpark, and SQL, alongside strong performance tuning, data security, and cost optimization skills. Candidates with prior experience working in the retail domain will be strongly preferred. Cost optimization skills are essential.

  • Develop and manage ETL pipelines for structured, semi-structured, and unstructured data using AWS Glue, PySpark, and SQL.
  • Handle real-time event stream data ingestion and processing from multiple source systems.
  • Ensure efficient data integration into Databricks for advanced processing and analytics.
  • Build and optimize backend systems leveraging AWS services (Glue, Athena, Lambda, SNS, S3).
  • Implement, configure, and manage Databricks environments, including clusters, notebooks, and libraries for performance optimization.
  • Ensure optimal resource utilization for AWS and Databricks clusters to improve efficiency and reduce costs.
  • Integrate Databricks with various cloud services while following governance and security best practices.
Testing & CI/CD Best Practices
  • Write unit test cases and integration tests to ensure data pipeline reliability.
  • Establish best practices for Databricks CI/CD and implement automation for deployment.
Optimization & Security
  • Apply performance tuning techniques to optimize queries, storage, and processing times.
  • Ensure compliance with security, governance, and industry best practices across AWS and Databricks environments.
  • Monitor system performance and proactively address issues to maintain high availability and reliability.
People Management:

The incumbent holds no direct supervisory responsibilities but is expected to engage effectively within their function and collaboratively across cross-functional teams.

This role contributes meaningfully, whether through operational contributions and/or by offering specialized expertise, guidance, and support, to ensure alignment with either functional and/or strategic organizational goals and objectives.

Knowledge of Glue, PySpark, SQL, Athena, Lambda, SNS, S3

Knowledge of Databricks: Cluster setup, Notebooks, Libraries, CI/CD, Optimization

Data Processing: Event stream ingestion and batch processing

Testing: Writing unit test cases and integration tests

Security & Governance: AWS/Databricks governance standards and best practices

Performance Optimization: Query tuning, cluster performance improvements, cost reduction

Strong problem-solving and analytical skills

Ability to work in a fast-paced, cloud-based data environment

Excellent collaboration and communication skills

Strong attention to detail and commitment to best practices

Experience:

Minimum 6 experience in Data engineering (Core development/design), in which 3+ years on AWS with strong hands on (AWS glue, pyspark, SQL, Athena, lambda, SNS, S3) and 2+ year on Databricks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (AWS)
Data Engineer (AWS)

Weekday (YC W21) • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Data Engineer (AWS, Databricks, PySpark)
Data Engineer (AWS, Databricks, PySpark)

Tata Consultancy Services • Hyderabad, Bengaluru

On-site
INR 4,000,000 - 6,000,000
Aws Data Engineer
Aws Data Engineer

Tekskills • Gurugram District

On-site
INR 1,800,000 - 3,000,000
Data Engineer
Data Engineer

Minfy • India

On-site
INR 1,500,000 - 2,300,000
Data Engineer
Data Engineer

AppSquadz • Dadri

On-site
INR 1,200,000 - 2,500,000
AWS DATA ENGINEER
AWS DATA ENGINEER

NITYO • Bengaluru

On-site
INR 1,000,000 - 1,500,000
AWS Glue Developer
AWS Glue Developer

Elabs Infotech • Pune District

On-site
INR 3,000,000 - 6,000,000
Databricks - Data Engineer
Databricks - Data Engineer

Tredence Inc. • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Databricks / Spark Data Engineer
Databricks / Spark Data Engineer

Saicon • India

On-site
INR 1,200,000 - 1,800,000
Data Engineer ( AWS & Databricks)
Data Engineer ( AWS & Databricks)

Tekskills • Hyderabad

Hybrid
INR 2,500,000 - 4,200,000