Senior AI/ML Data Scientist or Engineer - LLM Fine-Tuning, Evaluation & Inference

Vamstar

India

On-site

INR 600,000 - 1,200,000

Full time

41 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Vamstar is seeking a senior ML data and platform engineer in India to own end-to-end ML data pipelines, evaluation workflows and model-training support. You will build ingestion, parsing, cleaning, normalisation, enrichment and deduplication of execution logs and documents, and convert raw data into structured supervised examples with context, tools and reference labels.

The role focuses on ML platform and MLOps, deployment, testing, monitoring and cost-aware optimisation.

Qualifications

  • 7+ years of Python experience.
  • 5+ years of AWS experience.
  • 7+ years across ML, data, backend or applied AI engineering.
  • Hands-on transformer fine-tuning using PyTorch, Hugging Face or equivalent tooling.

Responsibilities

  • Build pipelines to ingest, parse, clean, normalise, enrich and deduplicate execution logs and documents.
  • Convert raw logs into structured supervised examples with context, candidate tools, expected actions and reference labels.
  • Create versioned training, validation and holdout datasets with privacy, schema, label and quality checks.
  • Design leakage-aware splits using task- or session-level boundaries.
  • Build reproducible evaluations for tool-selection accuracy, reference accuracy, hallucinations, disagreements and failure slices.

Skills

Python
AWS
ML engineering
Transformer fine-tuning
PyTorch
Hugging Face
Model training
Experiment tracking
SQL
NoSQL
Spark
Kafka
Flink
Vector databases
RAG embeddings
Automation

Job description

Core responsibilities
ML data and evaluation
  • Build pipelines to ingest, parse, clean, normalise, enrich and deduplicate execution logs and documents.
  • Convert raw logs into structured supervised examples with context, candidate tools, expected actions and reference labels.
  • Create versioned training, validation and holdout datasets with privacy, schema, label and quality checks.
  • Design leakage-aware splits using task- or session-level boundaries.
  • Build reproducible evaluations for tool-selection accuracy, reference accuracy, hallucinations, disagreements and failure slices.
  • Trace misleading metrics back to inputs, labels, dataset versions and model versions.
  • Benchmark open-weight models such as Qwen, Llama and Mistral on quality, training effort, long-context behaviour and inference cost.
  • Run supervised fine-tuning, knowledge distillation and parameter-efficient fine-tuning experiments.
  • Apply sound experiment design: clear baselines, controlled comparisons, reproducible configurations and error analysis.
  • Identify overfitting, leakage, sampling bias, label noise and uneven task coverage.
  • Maintain versioned datasets, prompts, training configurations, checkpoints and evaluation code.
  • Recommend which experiments justify additional training or inference cost.
ML platform and MLOps
  • Take models through deployment, testing, monitoring and controlled release.
  • Benchmark serving configurations for throughput, latency, GPU memory, concurrency and cost at an agreed quality level.
  • Evaluate vLLM or comparable runtimes, INT8 quantisation, checkpoint compatibility and KV-cache requirements.
  • Measure cost per successful task or action, not just cost per token.
  • Implement model versioning, CI/CD, shadow inference, disagreement logging, feature-flag releases and rollback.
  • Monitor data drift, quality regressions, serving errors, latency and cost.
LLM workflow engineering
  • Build pipelines involving document parsing, chunking, embeddings, vector indexing and retrieval.
  • Implement function or tool calling, multi-step workflows, retries, fallbacks and state handling.
  • Design observability across prompts, retrieved context, model outputs, tool calls and downstream failures.
  • Write production-grade, tested Python for large-scale document and execution-data processing.
  • 7+ years of Python experience.
  • 5+ years of AWS experience.
  • 7+ years across ML, data, backend or applied AI engineering, with production delivery experience.
  • Strong supervised ML fundamentals, including validation design, overfitting, leakage, sampling bias and error analysis.
  • Hands-on transformer fine-tuning using PyTorch, Hugging Face or equivalent tooling.
  • Direct ownership of a completed model training, fine-tuning, distillation or benchmarking project.
  • Experience preparing ML/LLM datasets and building model evaluation code.
  • Experience with experiment tracking, model versioning, CI/CD, deployment, monitoring and rollback.
  • Working knowledge of GPU inference, quantisation, batching, concurrency, tail latency and cost optimisation.
  • Experience with large-scale data processing using Spark, Kafka, Flink or comparable technologies.
  • Strong SQL and practical experience with relational and NoSQL systems.
  • Production experience with RAG, embeddings, vector databases, function calling or agentic workflows.
  • Ability to work autonomously and own outcomes without detailed specifications.
Preferred background
  • Startup or similarly high-ownership environment.
  • AWS S3, Glue, IAM, EC2 GPU instances, CloudWatch and model registries.
  • vLLM, NVIDIA GPU optimisation or AWS G5-family instances.
  • Qwen, Llama or Mistral model families.
  • LoRA, QLoRA, knowledge distillation or long-context evaluation.
  • Enterprise or regulated environments where privacy, auditability and governance matter.
What to include with your application
  • A model you trained, fine-tuned, distilled or benchmarked, including the dataset, validation approach, metrics and your contribution.
  • An ML data, evaluation, deployment or inference system you owned, including scale, failure handling and measurable results.
  • Your current Indian city, earliest start date and interview availability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM
Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM

Carelon Global Solutions India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
AI Developer
AI Developer

Salvo Software • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Senior ML/AI Engineer
Senior ML/AI Engineer

Syren Cloud Inc. • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior/Principal Local Llm & Generative Ai Platform Engineer
Senior/Principal Local Llm & Generative Ai Platform Engineer

Parallelwireless • Maharashtra

On-site
INR 3,000,000 - 5,500,000
AI Developer
AI Developer

Salvo Software LLC • India

On-site
INR 900,000 - 1,500,000
Senior AI / ML Architect – Model Training, Fine-Tuning & Hosting
Senior AI / ML Architect – Model Training, Fine-Tuning & Hosting

CriticalRiver Inc. • Hyderabad

On-site
INR 4,000,000 - 8,000,000
Machine Learning Engineer
Machine Learning Engineer

Tranzeal • Bengaluru

On-site
INR 3,500,000 - 7,500,000
LLM Engineer (Large Language Models)
LLM Engineer (Large Language Models)

Fospe UK Ltd • Bengaluru

Hybrid
INR 2,500,000 - 5,200,000
Competitive compensation with bonuses
Hybrid work at Bangalore Innovation Cn
Health, dental, wellness insurance
+3
Senior MLOps Engineer
Senior MLOps Engineer

Tranzeal • Bangalore Rural, Bengaluru

On-site
INR 1,800,000 - 3,000,000
AI/ML Ops Engineer
AI/ML Ops Engineer

ZK Technologies • Pune District

Hybrid
INR 3,000,000 - 5,400,000