Big Data Engineer - Python/PySpark

Qcentrio

Mumbai

On-site

INR 2,500,000 - 4,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Qcentrio is seeking a senior Data Engineer to actively participate in all phases of the software development lifecycle, focusing on data pipelines and lake solutions. You will work with business and technology stakeholders to design scalable features and ensure data quality at scale.

The ideal candidate has 6+ years of experience, strong Python, SQL, Spark/PySpark skills, and hands-on AWS (S3, EMR, Databricks). Experience with Airflow and GitHub is preferred, and a Bachelor's in CS is required.

Qualifications

  • 5+ years of relevant experience developing Data and analytic solutions.
  • Experience building data lake solutions leveraging AWS, EMR, S3, Hive & PySpark.
  • Experience with relational SQL.
  • Experience with Python and scripting languages.

Responsibilities

  • Participate in all phases of the software development lifecycle from requirements gathering to rollout and support.
  • Solve complex business problems using disciplined development methodology.
  • Produce scalable, flexible, and maintainable data solutions using applicable technologies.
  • Analyze source and target data; map transformations to meet requirements.
  • Interact with clients and onsite coordinators during project phases.
  • Design and implement data features with business and tech stakeholders.
  • Improve data quality by identifying and solving data management issues.

Skills

Big Data
Python
SQL
Spark/PySpark
AWS Cloud

Education

Bachelor's degree in computer science

Tools

AWS
EMR
S3
Hive
PySpark
GitHub
Airflow
Databricks

Job description

Work Location : Pan India
Experience : 6+ Years
Notice Period : Immediate - 30 days
Mandatory Skills : Big Data, Python, SQL, Spark/Pyspark, AWS Cloud

JD and required Skills & Responsibilities :
  • Actively participate in all phases of the software development lifecycle, including requirements gathering, functional and technical design, development, testing, roll-out, and support.
  • Solve complex business problems by utilizing a disciplined development methodology.
  • Produce scalable, flexible, efficient, and supportable solutions using appropriate technologies.
  • Analyse the source and target system data. Map the transformation that meets the requirements.
  • Interact with the client and onsite coordinators during different phases of a project.
  • Design and implement product features in collaboration with business and Technology stakeholders.
  • Anticipate, identify, and solve issues concerning data management to improve data quality.
  • Clean, prepare, and optimize data at scale for ingestion and consumption.
  • Support the implementation of new data management projects and re-structure the current data architecture.
  • Implement automated workflows and routines using workflow scheduling tools.
  • Understand and use continuous integration, test-driven development, and production deployment frameworks.
  • Participate in design, code, test plans, and dataset implementation performed by other data engineers in support of maintaining data engineering standards.
  • Analyze and profile data for the purpose of designing scalable solutions.
  • Troubleshoot straightforward data issues and perform root cause analysis to proactively resolve product issues.
Required Skills :
  • 5+ years of relevant experience developing Data and analytic solutions.
  • Experience building data lake solutions leveraging one or more of the following AWS, EMR, S3, Hive & PySpark
  • Experience with relational SQL.
  • Experience with scripting languages such as Python.
  • Experience with source control tools such as GitHub and related dev process.
  • Experience with workflow scheduling tools such as Airflow.
  • In-depth knowledge of AWS Cloud (S3, EMR, Databricks)
  • Has a passion for data solutions.
  • Has a strong problem-solving and analytical mindset
  • Working experience in the design, Development, and test of data pipelines.
  • Experience working with Agile Teams.
  • Able to influence and communicate effectively, both verbally and in writing, with team members and business stakeholders
  • Able to quickly pick up new programming languages, technologies, and frameworks.
  • Bachelor's degree in computer science
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer
Big Data Engineer

Qcentrio • Jaipur

On-site
INR 1,500,000 - 2,100,000
Big Data Engineer
Big Data Engineer

Qcentrio • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Big Data Engineer
Big Data Engineer

Qcentrio • Pune District

Hybrid
INR 2,500,000 - 3,600,000
Big Data Engineer - Python/PySpark
Big Data Engineer - Python/PySpark

Qcentrio • Jaipur

On-site
INR 1,500,000 - 2,100,000
Big Data Engineer - Python/PySpark
Big Data Engineer - Python/PySpark

Qcentrio • Surat

On-site
INR 1,200,000 - 2,400,000
Big Data Engineer - Python/PySpark
Big Data Engineer - Python/PySpark

Qcentrio • Dadri

On-site
INR 1,400,000 - 2,200,000
Big Data Engineer - Python / PySpark
Big Data Engineer - Python / PySpark

Qcentrio • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Big Data Engineer
Big Data Engineer

Qcentrio • Gurugram District

On-site
INR 2,400,000 - 4,200,000
Big Data Engineer
Big Data Engineer

Qcentrio • Dadri

On-site
INR 1,500,000 - 2,100,000
Big Data Engineer -Pyspark, Spark, Hadoop, SQL
Big Data Engineer -Pyspark, Spark, Hadoop, SQL

Pi Square Technologies • Hyderabad, Pune District, Bengaluru

Hybrid
INR 1,800,000 - 3,000,000