Data Engineer

Coretek Services India

Hyderabad

Hybrid

INR 2,800,000 - 4,000,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Coretek Services India in Hyderabad seeks an experienced Data Engineer to design, build, and maintain production-grade batch and streaming data pipelines that are idempotent, observable, and recoverable.

You'll model data for analytics, integrate various data sources, implement data quality checks, own pipelines in production, and collaborate with analysts, data scientists, and product engineers to turn complex requirements into robust data contracts while applying CI/CD and infrastructure as

Qualifications

  • 5+ years building production data pipelines.
  • Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
  • Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
  • Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
  • Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
  • Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
  • Git-based workflow and experience shipping through CI/CD.
  • Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
  • Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
  • Strong knowledge and experience in working with customers in a consultative approach in a technical environment.

Responsibilities

  • Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
  • Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
  • Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
  • Build data quality checks into pipelines and define how failures alert and elevate.
  • Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
  • Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing).
  • Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
  • Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
  • Partner with analysts, data scientists, and product engineers to turn ambiguous requirements into data contracts.
  • Maintain data dictionaries, lineage, and pipeline runbooks so consumers can find a dataset, understand what each field means and how current it is, and use it correctly without having to ask the team that built it.

Skills

Python
PySpark
SQL
Data modeling
CI/CD
Git
Azure data platform
Databricks
Spark UI
Kafka
Event Hubs
Delta Lake
Iceberg
Terraform
Docker
Kubernetes
GitHub

Tools

Azure Data Factory
Databricks
Synapse/Fabric
ADLS
Kafka
Event Hubs
Delta Lake
Iceberg
Terraform
Bicep
Docker
Kubernetes
GitHub
Azure DevOps

Job description

Role & responsibilities
  • Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
  • Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
  • Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
  • Build data quality checks (freshness, volume, uniqueness, referential integrity) into pipelines rather than bolting them on afterward, and define how failures alert and elevate.
  • Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
  • Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing) and make the tradeoffs explicit.
  • Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
  • Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
  • Partner with analysts, data scientists, and product engineers to turn ambiguous requirements into durable data contracts.
  • Maintain data dictionaries, lineage, and pipeline runbooks so consumers can find a dataset, understand what each field means and how current it is, and use it correctly without having to ask the team that built it.

Preferred candidate profile

  • 5+ years building production data pipelines.
  • Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
  • Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
  • Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
  • Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
  • Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
  • Git-based workflow and experience shipping through CI/CD.
  • Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
  • Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
  • Strong knowledge and experience in working with customers in a consultative approach in a technical environment.

Additional Qualifications
  • Streaming experience (Kafka, Event Hubs).
  • Lakehouse formats: Delta Lake, Iceberg.
  • Infrastructure as code (Terraform, Bicep) and containerization (Docker, Kubernetes).
  • Experience in a regulated environment (HIPAA, SOC 2, PCI, GDPR): auditability, encryption, data residency.
  • Experience building data platforms for ML or supporting feature pipelines.
  • Proven ability to manage multiple client projects and deliver high-quality results on time.
  • Experience in Azure DevOps or GitHub for source control and pipelines.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Coretek Services • Kondapur

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

Coretek • Kondapur

On-site
INR 1,200,000 - 1,800,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Group Data Engineer I
Group Data Engineer I

DP World • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

CloudBoson • Dadri

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

Infosys • Maharashtra

On-site
INR 1,800,000 - 3,200,000
Data Engineering Lead
Data Engineering Lead

Kumaran Systems • Hyderabad

On-site
INR 2,800,000 - 4,000,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Data Engineer
Data Engineer

Metlife • Pune District

Hybrid
INR 900,000 - 1,500,000