Principal Data Engineer

Dun & Bradstreet India

Hyderabad

On-site

INR 4,000,000 - 6,400,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dun & Bradstreet India is seeking an experienced Data Engineer to build and optimize scalable data pipelines for large datasets, spanning ingestion, normalization, matching, and deduplication. The role emphasizes reliability, performance, and cost efficiency.

You will design foundational data architecture, collaborate with cross-functional teams, and translate business requirements into robust data solutions using Python and SQL.

Qualifications

  • 8-12+ years of hands‑on data engineering or large‑scale data processing experience.
  • Proven track record building production-grade data pipelines and distributed systems.
  • Strong expertise in SQL and Python for data processing and analysis.

Responsibilities

  • Build and optimize scalable data pipelines for large datasets.
  • Design foundational data architecture for identity resolution and ID graphs.
  • Develop systems for data ingestion, normalization, matching, and deduplication.
  • Write efficient Python and SQL code for large-scale processing.
  • Optimize data models, queries, and storage for performance and cost.
  • Build data validation, monitoring, and alerting to ensure data quality.

Skills

SQL
Python
Relational databases
Data modeling
Distributed systems
Analytical thinking

Tools

Docker
Kubernetes
GitHub Actions
Git
Google Cloud
AWS

Job description

  • Build, and optimize scalable data pipelines and ETL/ELT workflows for large, complex datasets
  • Design and implement foundational data architecture supporting identity resolution and ID graph systems
  • Develop and enhance systems supporting identity resolution and ID graph construction (data ingestion, normalization, matching, and deduplication)
  • Write efficient, testable, and maintainable code using Python and SQL for large-scale data processing
  • Optimize data models, queries, and storage strategies for performance, scalability, and cost efficiency
  • Build and maintain data validation, monitoring, and alerting systems to ensure data quality and reliability
  • Troubleshoot, debug, and improve existing data pipelines and infrastructure
  • Own and drive complex data problems end-to-end, from initial design through production deployment
  • Make and influence key technical decisions related to data architecture, scalability, and system design
  • Collaborate with data, platform, DevOps, and product teams to deliver scalable, production ready solutions
  • Translate business and product requirements into practical, performant data solutions
  • Document data pipelines, systems, and workflows clearly
  • Continuously improve system performance, data quality, and pipeline resilience.
  • Contribute to building new capabilities that improve how customers understand and leverage data insights.
Key Requirements:
  • 8-12+ years of hands-on experience in data engineering or large-scale data processing
  • Proven experience building and maintaining production-grade data pipelines and distributed systems
  • Demonstrated experience architecting and delivering large-scale data platforms or mission critical data systems
  • Strong expertise in: o SQL and relational databases (Postgres, BigQuery, Redshift, etc.) Python for data processing and analysis
  • Experience with Google Cloud Platform (BigQuery, Dataflow, Pub/Sub, Cloud Storage, Cloud Functions) and/or AWS (S3, Redshift, EMR, RDS)
  • Experience working with large-scale datasets (hundreds of millions to billions of records)
  • Strong understanding of data modeling, partitioning, indexing, and query optimization
  • Experience with distributed data processing and parallelization techniques
  • Experience moving large volumes of data across systems and architectures
  • Familiarity with CI/CD, containerization, and orchestration tools (Docker, Kubernetes, GitHub Actions, etc.)
  • Strong debugging and troubleshooting skills in complex data environments
  • Experience with version control (Git) and Agile tools (Jira, Confluence, etc.)
  • Highly analytical with strong attention to detail and a data-driven mindset
  • Ability to hit the ground running, quickly understand systems, and deliver independently
  • Comfortable working in a remote, fast-paced, and collaborative environment
  • Proven ability to drive system design and implementation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

GlobalNodes • Gurgaon

On-site
INR 1,500,000 - 2,100,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Proclink • Gandhamguda

On-site
INR 800,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Publicis Production • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Data Engineer
Data Engineer

Techversantinfotech • Ernakulam

On-site
INR 1,000,000 - 1,500,000
Sr. Data Engineer
Sr. Data Engineer

BigThinkCode • Chennai District

On-site
INR 800,000 - 1,200,000
Data Architect
Data Architect

Virtusa • Maharashtra

On-site
INR 4,200,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

DigitalXNode • Hyderabad

On-site
INR 2,000,000 - 3,000,000
Data Engineer
Data Engineer

TOPS Infosolutions Pvt. Ltd. • Ahmedabad District

On-site
INR 800,000 - 1,500,000