Python Data Processing Engineer

Infosys

Bengaluru

On-site

INR 600,000 - 800,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Infosys is seeking a Python - Data Processing professional in Bengaluru to design and implement robust data pipelines. You will build Python-based data processing components to handle structured and semi-structured data, ensuring data quality and end-to-end pipeline reliability.

The role requires 2–3 years of hands-on Python experience, strong SQL skills, and practical ETL experience. You will collaborate with cross‑functional teams and document pipeline design and data mappings for operational

Qualifications

  • Bachelor's degree in Engineering/Computer Science/IT or related field.
  • 2–3 years hands-on experience in Python for data processing.
  • Practical experience building ETL processes and data pipelines end-to-end.
  • Strong SQL skills including joins, aggregations, subqueries, and performance-aware query writing.
  • Ability to write clean, maintainable code and follow basic engineering practices (version control, reviews, testing mindset).

Responsibilities

  • Build and maintain Python-based data processing components for structured and semi-structured datasets.
  • Develop and support ETL workflows to ingest, transform, validate, and load data into target systems.
  • Design and optimize SQL queries for data extraction, transformation, reconciliation, and reporting needs.
  • Implement reliable data pipelines with proper logging, error handling, retries, and monitoring hooks.
  • Perform data quality checks, anomaly detection rules, and reconciliation to ensure accuracy and completeness.
  • Troubleshoot pipeline failures and performance bottlenecks; drive root‑cause analysis and permanent fixes.
  • Collaborate with cross‑functional teams to gather requirements and deliver incremental improvements.
  • Maintain clear technical documentation for pipeline design, data mappings, and operational runbooks.

Skills

Python data processing
ETL pipelines
SQL
Data pipelines end-to-end
Code quality
Version control

Education

Bachelor's degree in Engineering/CS/IT

Tools

Linux/Shell scripting
Pandas
Apache Airflow
Spark (PySpark)

Job description

Python - Data Processing
Primary skills
  • Python - Data Processing/Technology->Big Data - Data Processing->PySpark,Technology->OpenSystem->Python - OpenSystem->Python Minimum Qualifications:
Minimum Qualifications
  • Bachelor’s degree (or equivalent) in Engineering/Computer Science/IT or related field (BTech/BE/BSc or equivalent).
  • 2–3 years of hands‑on experience in Python for data processing and automation.
  • Practical experience building ETL processes and working with data pipelines end‑to‑end.
  • Strong SQL skills including joins, aggregations, subqueries, and performance‑aware query writing.
  • Ability to write clean, maintainable code and follow basic engineering practices (version control, reviews, testing mindset).
Preferred Qualifications
  • Experience designing scalable pipeline patterns (incremental loads, CDC concepts, partitioning, backfills).
  • Familiarity with Python data libraries and processing approaches (e.g., Pandas, batch processing patterns).
  • Exposure to orchestration/scheduling concepts and operationalizing pipelines for reliability and observability.
  • Experience working with large datasets and optimizing end‑to‑end pipeline performance (I/O, SQL tuning, compute efficiency).
  • Proven ability to collaborate with stakeholders, translate requirements into technical solutions, and deliver within timelines.
Good to have skills
  • Pandas
  • NumPy
  • Apache Airflow
  • Spark (PySpark)
  • Linux/Shell Scripting
Key Responsibilities
  • Build and maintain Python-based data processing components for structured and semi-structured datasets.
  • Develop and support ETL workflows to ingest, transform, validate, and load data into target systems.
  • Design and optimize SQL queries for data extraction, transformation, reconciliation, and reporting needs.
  • Implement reliable data pipelines with proper logging, error handling, retries, and monitoring hooks.
  • Perform data quality checks, anomaly detection rules, and reconciliation to ensure accuracy and completeness.
  • Troubleshoot pipeline failures and performance bottlenecks; drive root‑cause analysis and permanent fixes.
  • Collaborate with cross‑functional teams to gather requirements and deliver incremental improvements.
  • Maintain clear technical documentation for pipeline design, data mappings, and operational runbooks.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Python Data Engineer
Python Data Engineer

Lonvec Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Python Developer
Python Developer

Gemini Solutions • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Data Engineer - Python
Data Engineer - Python

IntraEdge • Bengaluru

On-site
INR 600,000 - 1,200,000
Python, PySpark, ETL Developer
Python, PySpark, ETL Developer

Infosys • Hyderabad

On-site
INR 2,000,000 - 4,200,000
Python /Pyspark Developer
Python /Pyspark Developer

CIEL HR • Bengaluru

On-site
INR 1,200,000 - 2,500,000
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)

Tata Consultancy Services • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,600,000
Pyspark
Pyspark

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Authorasist Hyderabad • Hyderabad, Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Senior ETL / Data Engineer
Senior ETL / Data Engineer

Durapid Technologies Pvt Ltd • India

On-site
INR 2,500,000 - 4,500,000
Data Engineer
Data Engineer

Staples India • Chennai District

On-site
INR 800,000 - 1,200,000