Analytics Engineer

Cutshort

Mumbai

On-site

INR 400,000 - 700,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cutshort is seeking a data/analytics engineer to own the end-to-end data-structuring layer across the organisation. You will transform large volumes of raw and unstructured data (SMS, device, bureau, and app logs) into clean, analytics-ready datasets powering risk analytics, fraud detection, and underwriting.

You will collaborate with data science, tech, and product teams to design data schemas, build batch and streaming pipelines on AWS, maintain a feature store, and ensure low-latency data for

Qualifications

  • Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
  • Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
  • Must have established a continuous learning cycle to expand parser coverage
  • Experience in Lending / NBFC / Fintech domain
  • Experience working with Bureau, SMS, Device, or Banking data
  • Strong Python and SQL (production level)
  • Experience handling unstructured data (SMS, logs, JSON, APIs)
  • Experience building data pipelines, schedulers, and cron jobs
  • Strong database design and data modelling skills
  • Ability to work in a startup environment with high ownership
  • Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift

Responsibilities

  • End-to-End Data Ownership
  • Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion
  • Work closely with Tech, Product, and Data Science to define what data should be captured
  • Maintain data documentation, data dictionaries, and schema governance
  • Ensure data quality, consistency, and version control
  • Unstructured Data Processing (Highest Priority)
  • Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
  • Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
  • Convert JSON, logs, and raw API responses into structured feature tables
  • Build regex/keyword-based parsers for financial SMS classification
  • Feature Implementation (From Risk & Data Science Team)
  • Implement feature creation logic provided by Risk/Data Science team
  • Translate business and policy logic into SQL/Python pipelines
  • Create reusable feature layers for underwriting, fraud, collections, and monitoring
  • Maintain a feature store for consistent model and policy usage
  • Lending Data Understanding (Domain-Specific Requirement)
  • Work with Bureau data
  • Structure SMS-derived financial variables (income, stress, EMI signals)
  • Work with Account Aggregator and bank transaction datasets
  • Understand fintech alternate data used in underwriting and fraud detection
  • Data Pipelines & Automation
  • Build and maintain ETL/ELT pipelines using Python & SQL
  • Create cron jobs for automated data ingestion and feature refresh
  • Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
  • Ensure low-latency pipelines for real-time underwriting use cases
  • Database Structuring & Storage Architecture
  • Structure clean datasets in PostgreSQL (analytics layer)
  • Manage raw data storage in DynamoDB / S3 data lake
  • Design normalized and denormalised tables for risk analytics
  • Optimise database performance for large-scale query workloads
  • Dashboards & Readable Data Layer
  • Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
  • Enable self-serve data access for Risk, Business, and Founders
  • Support ad-hoc analysis requirements from leadership
  • Cross-Functional Collaboration (Very Important)
  • The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.

Skills

Data engineering
SMS parsing
Unstructured data handling
Python
SQL
Data pipelines
Scheduling/cron jobs
Data modelling
Lending/fintech domain
Bureau/Banking data
AWS platform knowledge

Tools

AWS
Snowflake
Google BigQuery
Redshift
PostgreSQL
DynamoDB
APIs
Cron jobs tooling

Job description

  • Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
  • Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
  • Must have established a continuous learning cycle to expand parser coverage
  • Experience in Lending / NBFC / Fintech domain
  • Experience working with Bureau, SMS, Device, or Banking data
  • Strong Python and SQL (production level)
  • Experience handling unstructured data (SMS, logs, JSON, APIs)
  • Experience building data pipelines, schedulers, and cron jobs
  • Strong database design and data modelling skills
  • Ability to work in a startup environment with high ownership
  • Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Must-Have Skills
  • Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
  • Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
  • Must have established a continuous learning cycle to expand parser coverage
  • Experience in Lending / NBFC / Fintech domain
  • Experience working with Bureau, SMS, Device, or Banking data
  • Strong Python and SQL (production level)
  • Experience handling unstructured data (SMS, logs, JSON, APIs)
  • Experience building data pipelines, schedulers, and cron jobs
  • Strong database design and data modelling skills
  • Ability to work in a startup environment with high ownership
  • Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Good to Have
  • Experience in STPL, especially less than 25K ticket size
  • Experience with streaming (Kafka/Kinesis) and orchestration (Airflow or Step Functions)
  • Experience with feature stores and risk analytics datasets
  • Knowledge of regex, NLP basics for SMS parsing
  • Experience supporting real-time decision engines/underwriting systems
Role Summary

This role will be responsible for owning the end-to-end data-structuring layer across the organisation. The individual will transform large volumes of raw, unstructured, and semi-structured data (such as SMS, device, bureau, and app data) into clean, standardised, and analysis-ready datasets. These structured datasets will directly power risk analytics, fraud detection, marketing insights, collections strategy, and policy decisioning.

Key Objective of the Role

Ensure all raw lending data (SMS, Bureau, Device, AA, App logs) is captured, parsed, structured, and stored in a clean analytics-ready format inside databases (PostgreSQL, DynamoDB, AWS stack) so that the Risk and Data Science team can directly use it for feature creation, policy building, and portfolio monitoring.

Core Responsibilities
  • End-to-End Data Ownership
  • Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion

parsing -> structuring -> storage

  • Work closely with Tech, Product, and Data Science to define what data should be captured
  • Maintain data documentation, data dictionaries, and schema governance
  • Ensure data quality, consistency, and version control
  • Unstructured Data Processing (Highest Priority)
  • Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
  • Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
  • Convert JSON, logs, and raw API responses into structured feature tables
  • Build regex/keyword-based parsers for financial SMS classification
  • Feature Implementation (From Risk & Data Science Team)
  • Implement feature creation logic provided by Risk/Data Science team
  • Translate business and policy logic into SQL/Python pipelines
  • Create reusable feature layers for underwriting, fraud, collections, and monitoring
  • Maintain a feature store for consistent model and policy usage
  • Lending Data Understanding (Domain-Specific Requirement)
  • Work with Bureau data
  • Structure SMS-derived financial variables (income, stress, EMI signals)
  • Work with Account Aggregator and bank transaction datasets
  • Understand fintech alternate data used in underwriting and fraud detection
  • Data Pipelines & Automation
  • Build and maintain ETL/ELT pipelines using Python & SQL
  • Create cron jobs for automated data ingestion and feature refresh
  • Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
  • Ensure low-latency pipelines for real-time underwriting use cases
  • Database Structuring & Storage Architecture
  • Structure clean datasets in PostgreSQL (analytics layer)
  • Manage raw data storage in DynamoDB / S3 data lake
  • Design normalized and denormalised tables for risk analytics
  • Optimise database performance for large-scale query workloads
  • Dashboards & Readable Data Layer
  • Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
  • Enable self-serve data access for Risk, Business, and Founders
  • Support ad-hoc analysis requirements from leadership
  • Cross-Functional Collaboration (Very Important)
  • The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.
Tech Stack (Current Environment)
  • AWS Services
  • PostgreSQL (Primary analytics DB)
  • DynamoDB (Raw/NoSQL storage)
  • Python (Pandas, NumPy, ETL frameworks)
  • Advanced SQL
  • APIs, JSON, and Log Data Handling

Skills:- Data Structures, Python, SQL, Amazon Web Services (AWS) and PostgreSQL

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Analyst
Sr. Data Analyst

Time Hack Consulting • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Senior Data Scientist
Senior Data Scientist

Comviva • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Credit Risk Analyst
Credit Risk Analyst

One97 Communications Limited • Dadri

On-site
INR 1,000,000 - 1,500,000
Senior Data Analyst
Senior Data Analyst

JobItUs • Ahmedabad District

On-site
INR 1,000,000 - 1,500,000
Senior Risk Analyst
Senior Risk Analyst

Comviva • Gurugram District

On-site
INR 800,000 - 1,200,000
Senior Data Analyst
Senior Data Analyst

Sayyam Investments Private Limited • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior Data Analyst
Senior Data Analyst

Mantras2success Consultants • Ahmedabad District

On-site
INR 1,500,000 - 2,100,000
DATA & RISK PLATFORM - LEAD || Fintech || Mumbai
DATA & RISK PLATFORM - LEAD || Fintech || Mumbai

KSA INC • Mumbai

On-site
INR 3,500,000 - 5,000,000
Senior Data Scientist
Senior Data Scientist

Dun & Bradstreet Technologies & Data Services • Chennai District

On-site
INR 1,500,000 - 2,500,000
Data Scientist Executive
Data Scientist Executive

GSTSahay • Ahmedabad District

On-site
INR 1,200,000 - 1,800,000