Lead Data Engineer

McCormick & Company

Gurugram District

On-site

INR 1,800,000 - 2,400,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

McCormick & Company in India seeks an experienced Data Engineer to build and maintain scalable data pipelines from diverse sources, ensuring availability, reliability, and performance of data products.

You will design ETL/ELT, work with Azure Synapse, PySpark, APIs, and SQL, implement CI/CD for pipelines, and enforce data quality and security across 5-10+ source systems.

Qualifications

  • Bachelors in Mathematics, Statistics, CS, or related field.
  • Microsoft Certified: Fabric Data Engineer Associate or related cloud certifications a plus.
  • 8+ years of data engineering experience.
  • Demonstrated ability coding in one or more languages (PySpark preferred).
  • Experience with building data pipelines from multiple sources.
  • Experience with knowledge graphs a plus.
  • Proven ability to manage multiple priorities and production-grade pipelines.

Responsibilities

  • Build and maintain scalable data pipelines from 5-10+ source systems.
  • Design ETL/ELT solutions with data quality and security.
  • Implement CI/CD to streamline pipeline deployment.
  • Optimize performance for large datasets and workflows.
  • Collaborate with data science and product teams on improvements.
  • Proactively monitor and troubleshoot data pipelines.

Skills

PySpark
SQL
ETL development
Data modeling
Data quality
Data security
CI/CD
Partitioning

Education

Bachelor’s degree in Mathematics, Statistics, Computer Science, Data Analytics/Science, or related field

Tools

Azure Synapse
Databricks
Fabric
APIs

Job description

This role will be accountable for building and maintaining scalable data pipelines from source systems. The Data Engineer will ensure the availability, reliability, and performance of data products by integrating raw data from various sources. Key responsibilities include data modeling, ETL (Extract, Transform, Load) development, and ensuring data quality and security. This role will be accountable for data coming in from 5-10+ source systems.

Design and Execute
  • Partner with data product managers to gather and deliver data pipelines.
  • Design ETL solutions including data quality, data security, and data pipeline resiliency.
  • Execute ETL solutions including data security, data quality and performance requirements.
Data Extraction, Load and Transformation
  • Design, build and implement ELT pipelines to efficiently ingest and transform data from a wide variety of data sources and deliver datasets that meet business requirements.
  • Optimize performance for large datasets and data workflows for performance, scalability, and reliability to support business needs.
  • Develop and maintain scalable data pipelines leveraging Azure Synapse, PySpark, APIs, and SQL & performing advanced data cleaning, transformation, and manipulation to ensure high-quality, and reliable data flows.
  • Implement CI/CD processes to streamline and automate data pipelines deployment
  • Apply data validation frameworks (Great Expectations, Fabric-native tools) to maintain accuracy
  • Utilize partitioning, indexing, clustering strategies to enhance query performance
Process Improvement, Performance and Cost optimization tuning
  • Collaborate with Data Science, AI, and Data product teams to optimize performance and cost effectiveness of their solutions.
  • Identify and support the design of internal process improvements, including automating manual processes, optimizing data product delivery, and redesigning solutions for enhanced scalability.
  • Implement solution adjustments to improve performance and cost-effectiveness of data products.
Issue Resolution and Support
  • Monitor and troubleshoot the data pipelines proactively, which includes leading the support of data-related product pipeline issues to resolve data errors.
  • Provide expert-level support and guidance to data teams across the Enterprise.
Desired Candidate Profile:
  • Bachelor’s degree in Mathematics, Statistics, Computer Science, Data Analytics/Science, or related field
  • or Microsoft Certified: Fabric Data Engineer Associate or related cloud technologies, Fabric IQ/Databricks certifications a plus
  • 8+ years of data engineering experience.
  • Demonstrated ability coding in one or more languages (PySpark preferred).
  • Experience with building data pipelines.
  • Experience with knowledge graphs a plus.
  • Demonstrated ability to manage multiple priorities simultaneously.
  • Demonstrated ownership of production-grade pipelines.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

SOTI Inc • Ernakulam

On-site
INR 1,500,000 - 2,000,000
Lead Data Engineer
Lead Data Engineer

Flexi Careers India • Chennai District

On-site
INR 2,800,000 - 5,200,000
Data Engineer
Data Engineer

Rail Infotech India • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

Version 1 • Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Lead Data Engineer
Lead Data Engineer

Kanerika Inc • Hyderabad

On-site
INR 3,000,000 - 5,200,000
Lead Data Engineer
Lead Data Engineer

Kanerika Inc • Indore District

On-site
INR 3,500,000 - 7,000,000
Data Engineer
Data Engineer

DataPhi • Pune District

On-site
INR 1,200,000 - 1,800,000
Lead Data Engineer
Lead Data Engineer

Apexon • Bengaluru

On-site
INR 1,400,000 - 2,000,000
Lead Data Engineer - Fabric
Lead Data Engineer - Fabric

iLink Digital • Pune District

On-site
INR 1,500,000 - 2,500,000
Microsoft Fabric Data Engineer
Microsoft Fabric Data Engineer

Acme Services Private Limited • Hyderabad

On-site
INR 1,200,000 - 1,800,000