Job Description
In this role, you will be expected to work with our lead data scientist and the team executive as well as our internal analytics, technology, product and policy partners to deploy advanced analytical solutions with the goal of reducing fraud losses, reducing false positive declines at the point‑of‑sale, improving client experience, and ensuring the bank minimizes its total cost of fraud.
Responsibilities
- Extensive experience in Python and PySpark for building, optimizing, and maintaining scalable data pipelines using Big Data technologies (e.g., Spark, Hadoop ecosystems).
- Design, develop, and manage end‑to‑end data ingestion and transformation workflows to process large, complex datasets efficiently and make them analytics‑ and model‑ready.
- Implement automation and data orchestration to eliminate manual, repetitive processes, including scheduling, monitoring, and recovery of data jobs using Python/PySpark and workflow tools.
- Build and maintain high‑quality, reliable data foundations that support downstream use cases such as reporting, analytics, machine learning, and fraud detection systems.
- Develop data validation, quality checks, and anomaly detection mechanisms to ensure accuracy, consistency, and trust in enterprise datasets.
- Architect and support real‑time and near‑real‑time data pipelines, enabling low‑latency access to critical data for operational decision‑making.
- Collaborate closely with data scientists, product owners, and business stakeholders to define data requirements and data strategy, ensuring the availability of customer‑level, unified data views across systems.
- In addition, the role will be expected to work with the lead data scientist and the team executive and stakeholders to help develop the data strategy for client protection to ensure the organization has the proper data to make the right decisions, with a priority on data availability in real‑time, and generating true customer‑level views able to make intelligent fraud decisions leveraging the entirety of our interactions with a customer.
Education
- Advanced degree, preferably in Statistics/Mathematics, Applied Sciences, Engineering from a premier institute.
Certification
- Technical certifications in SQL, GSQL, Cypher, Hive, Python, PySpark preferred.
Experience
- 4‑6 years of relevant experience in the field of analytics.
- 2+ years of experience in engineering, extensive experience in PySpark and other big data methods, experience owning and optimizing a codebase.
Foundational Skills
- Python, PySpark and SQL.
- Good to have knowledge of big‑data systems.
- Good to have graph databases and related query languages.
- U.S. financial services experience preferable.
- Understanding of business domains like Fraud/Compliance/Risk preferable.
- Bachelor's degree in a quantitative discipline such as mathematics, engineering, economics, finance, business, computer science. Master's degree or higher preferred. In lieu of a specific degree, advanced certifications in combination with strong experience will also be considered.
- The candidate must be at an advanced to expert level on SQL and basic knowledge of statistical knowledge.
- Candidate must have a proven track record of building and deploying analytical solutions that have resulted in material financial results and extensive management experience. Ability to work in a fast‑paced, dynamic environment is critical. Must have exceptional organizational, project management and communications skills.
Desired Skills
- Solid knowledge of Python or Spark and various commercial and model generation software. Should additionally have familiarity with other tools such as HUE, Hive, and other data gathering tools.
Work Timings
11:30 AM to 8:30 PM, 12:30 PM to 9:30 PM
Location
Mumbai