Big Data Engineer - Python and Spark

Citi

Pune District

On-site

INR 600,000 - 900,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Citi is seeking a Data/Information Management Analyst trainee in Pune (India) to work on Big Data systems, Hive, Hadoop, and Spark with Python/Scala and SQL expertise. The role involves transforming unstructured code into optimized data processes and supporting data ingestion and governance tasks.

The candidate should have 0–2 years of experience in related tech and be capable of working in a team to deliver concise, data-driven solutions.

Qualifications

  • Master's/Engineering degree with 0–2 years in Big Data, Hive, Hadoop, Spark, and cloud data management.
  • Hands-on Unix scripting, Python and Scala with strong SQL experience.
  • Turn unstructured code into clean, best-class designs and efficient data processes.
  • Experience with large datasets, data warehouses, and data extraction via code.
  • Data preprocessing and application engineering skills.
  • At least 3 years of experience designing software systems with real-time and batch needs.
  • Experience with supervised/unsupervised ML techniques.
  • Exposure to data ingestion, ETL tools like Talend; Cloudera stack a plus.
  • Knowledge of data governance and data security.

Responsibilities

  • Collaborate with team to analyze data requirements and deliver solutions.
  • Resolve problems by applying technical experience and precedents.
  • Communicate findings clearly to stakeholders.

Skills

Python
Scala
SQL
Unix scripting
Big data technologies
Data preprocessing
Problem solving

Education

Master's / Engineering degree
Bachelor's / University degree

Tools

Talend
Cloudera stack
Pepper data

Job description

The Data/Information Mgt Analyst is a trainee professional role. Requires a good knowledge of the range of processes, procedures and systems to be used in carrying out assigned tasks and a basic understanding of the underlying concepts and principles upon which the job is based. Good understanding of how the team interacts with others in accomplishing the objectives of the area. Makes evaluative judgements based on the analysis of factual information. They are expected to resolve problems by identifying and selecting solutions through the application of acquired technical experience and will be guided by precedents. Must be able to exchange information in a concise way as well as be sensitive to audience diversity. Limited but direct impact on the business through the quality of the tasks/services provided. Impact of the job holder is restricted to own job.

Qualifications
  • Master's / Engineering Degree with 0- 2 years of experience in Big Data systems, Hive, Hadoop, Spark (Python/ scala) and cloud-based data management technologies
  • Hands-on experience in Unix Scripting, Python and Scala programing along with strong experience in SQL.
  • Comfortable working with completed unstructured, undocumented code and turning it around into best-class code redesigning costly compute and data processes and aligning to best development standards
  • Experienced in working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Well versed with necessary data preprocessing and application engineering skills
  • At least 3 years of experience designing software systems with intense computational needs across real time and batch process .
  • Experience and understanding of Supervised, unsupervised machine learning techniques
  • Exposure to data ingestion, ETL tools such as Talend, modeling tools, Performance Management tooling such as Pepper data, Cloudera stack will be a plus
  • Knowledge of data management, data governance, data security and regulatory practices
  • Ability to identify, clearly articulate and solve complex business problems and present them to the management in a structured and simpler form
  • Should have experience of working in onsite, offsite delivery model
  • Experience working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Previous related experience preferred
  • High attention to detail
Education
  • Bachelors/University degree or equivalent experience

This job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.

Job Family Group: Decision Management

Job Family: Data/Information Management

Time Type: Full time

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citigroup Inc. • Pune District

On-site
INR 600,000 - 1,000,000
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citi • Pune District

On-site
INR 3,500,000 - 6,000,000
Data Engineer - Assistant Vice President
Data Engineer - Assistant Vice President

Citi • Maharashtra

On-site
INR 1,500,000 - 2,100,000
Data Engineer – Big Data, Python, Databricks – Assistant Vice President
Data Engineer – Big Data, Python, Databricks – Assistant Vice President

Jobtailor • Chennai District

On-site
INR 1,200,000 - 1,800,000
Data Engineer - Assistant Vice President
Data Engineer - Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 1,200,000 - 1,600,000
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citigroup Inc. • Pune District

On-site
INR 3,000,000 - 5,000,000
PySpark Big Data Developer
PySpark Big Data Developer

Citi • Maharashtra

On-site
INR 1,100,000 - 1,800,000
Python data engineer - AI and ML applications - Vice president
Python data engineer - AI and ML applications - Vice president

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citi • Maharashtra

On-site
INR 3,000,000 - 4,500,000