Big Data Engineer - Python and Spark

JobCubby

Pune District

On-site

INR 900,000 - 1,200,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Citi is hiring a Data/Information Mgt Analyst trainee to design and optimize software systems handling large datasets and data warehouses. You will apply technical expertise to resolve complex business problems and align processes with best development standards.

The role requires 1–3 years in Big Data, Hive/Hadoop/Spark, and strong proficiency in Python/Scala/SQL. You will work on both real-time and batch processing, with exposure to ETL tools like Talend and data governance practices.

Qualifications

  • Engineering degree with 1–3 years in Big Data systems (Hive/Hadoop/Spark).
  • Strong Python, Scala, SQL and Unix scripting skills.
  • Experience with ML techniques and data management practices.
  • Experience designing software systems with high compute and data needs, across real‑time and batch processes.
  • Familiarity with data ingestion/ETL tools such as Talend and data governance concepts.
  • Excellent communication and stakeholder management; willingness to work in onsite/offsite delivery models.

Responsibilities

  • Design and redesign software systems with intense computational needs.
  • Handle large datasets and data warehouses; optimize compute and data processing.
  • Collaborate across teams to align processes with development standards; present findings succinctly.

Skills

Big Data
Hive
Hadoop
Spark
Python
Scala
SQL
Unix scripting
Machine learning
Data management
ETL concepts
Stakeholder communication

Education

Engineering degree
Bachelor level or equivalent

Tools

Talend
Pepper Data
Cloudera
Python
Scala

Job description

to apply - email only, no card. You can also save this posting or score it againstyour profile with AI.## About the roleThe role involves designing and redesigning software systems with intense computational needs while managing large datasets and data warehouses. The analyst will resolve complex business problems by applying technical experience and aligning processes with best development standards.## RequirementsCandidates must hold an engineering degree and possess 1-3 years of experience in Big Data systems, including Hive, Hadoop, and Spark. Strong proficiency in Python, Scala, SQL, and Unix scripting is required, along with experience in machine learning techniques and data management practices.## Full descriptionThe Data/Information Mgt Analyst is a trainee professional role. Requires a good knowledge of the range of processes, procedures and systems to be used in carrying out assigned tasks and a basic understanding of the underlying concepts and principles upon which the job is based. Good understanding of how the team interacts with others in accomplishing the objectives of the area. Makes evaluative judgements based on the analysis of factual information. They are expected to resolve problems by identifying and selecting solutions through the application of acquired technical experience and will be guided by precedents. Must be able to exchange information in a concise way as well as be sensitive to audience diversity. Limited but direct impact on the business through the quality of the tasks/services provided. Impact of the job holder is restricted to own job.Qualifications:* / Engineering Degree with more 1 - 3 years of experience in BigData systems, Hive, Hadoop, Spark (Python/ scala) and cloud based data management technologies* Hands-on experience in Unix Scripting, Python and Scala programing along with strong experience in SQL.* Comfortable working with completed unstructured, undocumented code and turning it around into best-in-class code redesigning costly compute and data processes and aligning to best development standards* Experienced in working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.* Well versed with necessary data preprocessing and application engineering skills* At least 3 years of experience designing software systems with intense computational needs across real time and batch process .* Experience and understanding of Supervised, unsupervised machine learning techniques* Exposure to data ingestion, ETL tools such as Talend, modeling tools, Performance Management tooling such as Pepper data, Cloudera stack will be a plus* Knowledge of data management, data governance, data security and regulatory practices* Ability to identify, clearly articulate and solve complex business problems and present them to the management in a structured and simpler form* Should have experience of working in onsite, offsite delivery model* Experience working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.* Experience in Credit Cards and Retail Banking* Should have excellent communication and inter-personal skills* Strong process/project management skills* Multiple stake holder management* Control orientated and Risk awarenessEducation:* Bachelors/University degree or equivalent experienceThis job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.-Job Family Group:Decision Management-Job Family:Data/Information Management-Time Type:Full time-Most Relevant SkillsPlease see the requirements listed above.-Other Relevant SkillsFor complementary skills, please see above and/or contact the recruiter.-Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.View Citi's EEO Policy Statement and the Know Your Rights poster.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citigroup • Pune District

On-site
INR 500,000 - 900,000
Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citi • Maharashtra

On-site
INR 600,000 - 900,000
Python Engineering AI Lead-Assistant Vice president
Python Engineering AI Lead-Assistant Vice president

Citibank (Switzerland) AG • Chennai District

Hybrid
Confidential
Senior Data Engineer - Assistant Vice President
Senior Data Engineer - Assistant Vice President

Citi • Chennai District

On-site
INR 2,800,000 - 5,200,000
Senior Data Engineer - Assistant Vice President
Senior Data Engineer - Assistant Vice President

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
Associate Data Engineer
Associate Data Engineer

JobCubby • Hyderabad

On-site
INR 1,200,000 - 2,000,000
Data Platform Engineer (AI-Enabled)
Data Platform Engineer (AI-Enabled)

Citibank (Switzerland) AG • Pune District

Hybrid
Confidential
Assistant Vice President - Data Analytics
Assistant Vice President - Data Analytics

Citi • Chennai District

On-site
INR 3,000,000 - 5,000,000
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,000,000 - 8,000,000
Senior Data Engineer - Assistant Vice President
Senior Data Engineer - Assistant Vice President

Citigroup Inc. • Chennai District

On-site
INR 1,800,000 - 2,400,000