Senior ML Engineer

IDfy

Mumbai

On-site

INR 1,200,000 - 1,800,000

Full time

45 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IDfy in Mumbai and Pune seeks a senior ML engineer to shape, train, and productionize models at scale. You will frame business requirements into precise objectives, design robust validation, and own end-to-end model lifecycle from training to retraining.

Expect rigorous evaluation, scale-focused optimization, and leadership of 3–6 junior engineers, with emphasis on reproducibility and cost-aware improvements.

Qualifications

  • Proven model-training experience across deep learning models.
  • Depth in either computer vision or NLP with literacy in the other.
  • Shipped at least two models into production and managed their health against drift.

Responsibilities

  • Frame ambiguous problems mathematically and define objectives, validation, and metrics.
  • Train models end to end and own lifecycle from framing to retraining.
  • Evaluate with rigorous metrics (ROC, PR, AUC; avoid leakage).
  • Optimize inference at scale: quantize, prune, and CPU-offload when justified.
  • Lead and mentor 3–6 junior engineers in math-driven problem solving.
  • Develop production ML systems with 24x7 availability and monitoring.
  • Reproduce and improve research; favor simpler, cheaper models with equal accuracy.
  • Work with LLMs: fine-tuning, embeddings, retrieval, and RAG; explain retriever issues.

Skills

Probability and statistics
Optimization
Linear algebra
Derivation not import
Deep learning experience
PyTorch
Model training from scratch
Production ML systems
Evaluation design

Tools

PyTorch
Kubernetes
GCP
AWS
OpenVINO
CTranslate2
Triton
vLLM

Job description

Job Description:

The scale you will operate at

We run 40+ production ML models across two large families, on a mixed CPU and GPU fleet.

Documents: Region of Interest detection, Photocopy Classifier, Text Tampering, Photo Tampering, Readability, OCR, Named Entity Recognition, PII Masking.

Faces: Face Detection, Face Quality, Sunglass Detection, NSFW, Face Mask Detection, Face Match, Liveness Detection, Deepfake Detection.

The Production Footprint
  • 2000 requests per second at peak
  • 25 TB of serving RAM
  • ~18,000 CPUs
  • 120 NVIDIA L4 GPUs
  • 2M+ verifications per day

At this scale, cost-to-serve is a first-class engineering constraint. A model that is 2 percentage points more accurate but 10x more expensive to run may be the wrong model. You will own that tradeoff with numbers, not opinions.

What you will do

Frame ambiguous problems mathematically. Turn a business requirement like "catch deepfakes in KYC" into a well-posed objective: the right positive class, the right operating point, the right validation protocol that does not leak, and a metric that survives class-prevalence shift.

Train models, end to end. Own the full lifecycle for one or more model families: problem framing, data strategy, architecture choice, training, evaluation, deployment, monitoring, and retraining. You are accountable for the model in production, not just the notebook.

Evaluate with rigor. Design evaluation that predicts production behavior. Confusion matrices, ROC and PR curves, TPR at a fixed low FPR, calibration error, and cross-distribution generalization. Know why AUC can lie about a model you operate at FPR = 1e-4.

Optimize for inference at scale. Quantize, prune, and re-architect models so they serve within latency, throughput, and cost budgets. Move workloads off GPU to quantized CPU inference where the numbers justify it, and prove the accuracy held with a production canary before, not after. Deploy and operate resilient production systems that run 24x7

Reproduce and improve on research. Read a paper, reproduce its results, strip it down to what actually matters for our constraints, and ship it. We value simplification that preserves accuracy. A smaller, cheaper model that matches a heavier one is treated as a genuine result here, not a compromise.

Lead. Mentor 3 to 6 junior engineers on math-driven problem solving, experiment design, and ML systems. Set the bar for source discipline and reproducibility. Review models, not just code.

What you bring

ML fundamentals, deeply held.

  • Probability and statistics: distributions, estimation, hypothesis testing, class prevalence and its effect on precision, sampling.
  • Optimization: gradient descent and its variants, loss landscapes, convergence behavior, why training diverges and how to fix it.
  • Linear algebra: enough to reason about embeddings, projections, and what a layer is actually computing.
  • You can derive, not just import. If asked why a loss is shaped the way it is, or what changes when you quantize a layer to INT8, you can work it out on a whiteboard.

Proven model-training experience.

  • You have trained deep learning models from scratch and fine-tuned pretrained networks, in PyTorch. Training loops, data pipelines, augmentation, and loss design are things you have written, not things you have called.
  • Depth in at least one of Computer Vision (detection, classification, metric or embedding learning, anti-spoofing) or NLP (sequence labeling, NER, transformer models), with working literacy in the other.
  • You have shipped at least two models into production, operated them at some scale and kept them healthy through drift and adversarial pressure.

Evaluation and calibration.

  • You design evaluations that do not overstate performance: leakage-free splits, cross-generator or cross-source generalization tests, and operating points chosen for the real cost asymmetry.
  • You understand calibration and can reason about decision thresholds under changing prevalence.

Production ML at scale.

  • 4+ years building and operating large-scale ML systems: real-time inference, monitoring, drift detection, and retraining.
  • Inference optimization: quantization (INT8, BF16), and hands-on with at least one of OpenVINO, CTranslate2, Triton, or vLLM. You reason about latency, throughput, and cost-to-serve as engineering targets.
  • Comfortable on Kubernetes and a major cloud (we run primarily on GCP and AWS).

Working with LLMs, the right way.

  • If you work with LLMs, you work with the model, not only the API: fine-tuning, embedding geometry, retrieval statistics, and transformer internals. You can build a RAG system and also explain why the retriever is failing at the embedding level.

What makes you stand out

  • Face recognition, liveness, or presentation-attack and deepfake detection experience.
  • Metric learning and large-scale vector search (millions of vectors and up).
  • OCR, document forensics, or tampering detection on real-world degraded documents.
  • Contributions to open-source ML, published papers, or teaching and speaking in the community.
  • Familiarity with Indian regulatory context: DPDP, RBI KYC modes, UIDAI, PCI DSS.

Why this role is rare

Most "AI" roles today are integration roles: call a hosted model, shape a prompt, ship. This is not that. Here you build the models that have real world impact on 2 million people a day, in an adversarial setting where accuracy and cost both have real consequences. You get production traffic at 2000 RPS, a fleet of 40+ models to learn from, and a mandate to make them faster, better and cheaper.

This is a place for people who are still excited by the math.

Job Location: Mumbai, Pune.

Requirements:

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer - 2
Machine Learning Engineer - 2

SatSure Analytics India • India

On-site
INR 1,500,000 - 4,000,000
AI/ML engineer
AI/ML engineer

Tata Consulting Services, PLC • Bengaluru, Pune District

On-site
INR 1,200,000 - 2,000,000
AI/ML Engineer
AI/ML Engineer

Creuto Cloud Private Limited • India

On-site
INR 800,000 - 1,600,000
ML Engineer · Mid–Senior
ML Engineer · Mid–Senior

Think Right Advisory Services Pvt. Ltd. • Bengaluru Urban

Hybrid
INR 4,000,000 - 7,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Jobtailor • Chennai District

On-site
INR 3,000,000 - 6,000,000
AI/ML Engineer
AI/ML Engineer

Jash Data Sciences Pvt. Ltd. • Pune District

On-site
INR 800,000 - 1,200,000
Competitive salary
Learning opportunities
Exposure to latest AI technologies
Machine Learning Engineering Senior Engineer
Machine Learning Engineering Senior Engineer

V2soft • Chennai District

On-site
INR 1,200,000 - 1,800,000
Lead AI ML Engineer
Lead AI ML Engineer

Netscribes • Maharashtra

On-site
INR 300,000 - 500,000
Technical Lead – ML Platform & MLOps
Technical Lead – ML Platform & MLOps

Myntra • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

IVY Mobility • Gurugram District

On-site
INR 2,500,000 - 4,000,000