Data Engineer

Dun & Bradstreet India

Hyderabad

Hybrid

INR 800,000 - 1,400,000

Full time

13 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Dun & Bradstreet India is seeking a Data Engineer I to join an agile team supporting Research and Managed Services. This hybrid role involves hands-on coding, data tool development, and maintaining data pipelines to feed AI-enabled and traditional workflows.

The role requires evaluating data sources for quality, coverage, privacy, and technical compatibility, and implementing AI-enabled workflows with LangChain or equivalent frameworks. Adaptability to learn new tools is essential.

Qualifications

  • Strong SQL and Python skills with ability to write and maintain code as a core daily activity.
  • Experience with Playwright, Selenium, and other web-data collection techniques.
  • Experience developing and supporting data-ingestion, transformation, and ETL/ELT workflows.

Responsibilities

  • Write, review, test, and maintain code to build data tools, automations, and ingestion/ transformation workflows.
  • Automate manual processes and develop data tools to improve efficiency, accuracy, quality, and throughput.
  • Develop and promote coding standards and participate in code reviews within the agile team.
  • Build and maintain web-scraping solutions, API integrations, and reusable data-processing components.
  • Support scalable ETL/ELT pipelines for structured and unstructured data.

Skills

SQL
Python
Web scraping
ETL/ELT
Data pipelines
BigQuery
GCS/S3
LangChain / AI workflows
Stakeholder management
Dashboards (Power BI/Tableau)

Tools

Playwright
Selenium
Power BI
Tableau
LangChain

Job description

The Data Engineer I work as part of an agile team supporting the organization's Research and Managed Services. This is a hands-on, hybrid, technical, and operational role. The team members will actively write and maintain code to build data tools, automations, and pipelines, while also performing day-to-day operational tasks that keep Research and Managed Services running reliably. The role includes identifying and evaluating internal and external data sources to feed AI-enabled and traditional research workflows. The team member will assess source relevance, quality, coverage, accessibility, reliability, compliance considerations, and applicability to defined business and technical use cases. The role drives best-in-class standards and continuous improvement and requires an adaptable professional who is willing and able to learn and adopt new technologies as they are introduced to the organization

Key Responsibilities:
Coding & Development
  • Write, review, test, and maintain code, including SQL and Python, to build data tools, automations, and ingestion and transformation workflows that support Research and Managed Services
  • Automate manual processes and develop data tools to improve efficiency, accuracy, quality, and throughput
  • Develop and promote coding standards and contribute to code reviews within the agile team
  • Build and maintain web-scraping solutions, API integrations, and reusable data-processing components
  • Support scalable ETL/ELT pipelines for structured and unstructured data
Source Evaluation & AI Enablement
  • Identify prospective sources that can feed AI solutions and Research and Managed Services workflows
  • Define and apply source-evaluation criteria covering relevance, authority, freshness, completeness, coverage, consistency, accessibility, legal or licensing constraints, privacy, security, and technical compatibility
  • Perform source profiling, sample validation, proof-of-concept testing, and comparative assessments before recommending onboarding
  • Document source decisions, metadata, lineage, ownership, limitations, refresh expectations, and approved use cases
  • Implement and support AI-enabled workflows using LangChain or equivalent orchestration frameworks, large language models, embeddings, retrieval-augmented generation, vector databases, and prompt-engineering approaches where applicable
  • Monitor source and AI-workflow performance and recommend remediation, replacement, or additional sources when quality or coverage falls below requirements
Operational Tasks
  • Perform day-to-day operational activities supporting Research and Managed Services, including monitoring, exception handling, data maintenance, and issue resolution
  • Perform database administration activities, including performance tuning and implementation of best practices
  • Implement new data-maintenance processes and provide end-to-end process ownership
  • Ensure data integrity by validating, reconciling, and regularly cleaning data
  • Investigate and resolve production incidents, pipeline failures, data-quality issues, and operational exceptions
  • Follow applicable data governance, security, and operational standards
  • Evaluate and implement new technology solutions, and proactively learn and adopt new tools, platforms, and methodologies introduced by the organization
  • Communicate with stakeholders and conduct knowledge-exchange sessions for technical and non-technical audiences
  • Develop and maintain data documentation, including data dictionaries, source assessments, data-flow diagrams, data mappings, runbooks, and data lineage
  • Collaborate with cross-functional teams across Data & Analytics, Technology, Research Services, Managed Services, Product, and Data Governance
  • Additional duties as assigned.
Key Skills:
  • Strong SQL and Python skills, with demonstrated ability to write and maintain code as a core part of daily work
  • Experience with Playwright, Selenium, and other web-data collection techniques
  • Experience developing and supporting data-ingestion, transformation, and ETL/ELT workflows
  • Ability to collect and interpret data from multiple sources, including web scraping and GCS/S3, and formats including delimited files, XML, JSON, and PDF
  • Working knowledge of data systems and databases used to maintain data pipelines
  • Experience with Power BI, Tableau, or other dashboard tools
  • Experience managing stakeholders and project plans
  • Proficiency in Microsoft Office Suite
  • Willingness and demonstrated ability to learn new technologies as they are introduced
  • BigQuery experience and knowledge of AWS and/or GCP
  • Hands-on experience implementing AI solutions using LangChain or an equivalent orchestration framework
  • Exposure large language models, prompt engineering, retrieval-augmented generation, embeddings, vector databases, AI agents, or graph databases
  • Knowledge of Data Operations methodologies, data management approaches, ServiceNow, and/or Jira
  • Experience with NoSQL technologies, SQL Server administration, R programming, web technologies, and data mapping from multiple sources.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Engineer
Senior AI Engineer

Latent View Analytics Limited • Chennai District

On-site
INR 3,000,000 - 4,200,000
Senior Data Engineer
Senior Data Engineer

BMC Software • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer (AI Data Platforms & Automation)
Senior Data Engineer (AI Data Platforms & Automation)

E Solutions • Dadri

On-site
INR 1,200,000 - 2,400,000
Data & Integration Analyst
Data & Integration Analyst

Bold Business • Maharashtra

On-site
INR 1,800,000 - 2,400,000
Data & Integration Analyst
Data & Integration Analyst

SCALIS • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior AI Engineer
Senior AI Engineer

PeopleStrong • Bengaluru

On-site
INR 2,400,000 - 4,200,000
Data Engineering Operations Lead
Data Engineering Operations Lead

TransUnion • Pune District

On-site
INR 3,500,000 - 7,000,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
AI Engineer/Lead AI Engineer
AI Engineer/Lead AI Engineer

Salesforce • Bengaluru, Hyderabad

On-site
INR 4,000,000 - 8,000,000
Data Management Specialist
Data Management Specialist

HTC Global Services • Hyderabad

On-site
INR 1,200,000 - 1,800,000