Data Scientist
Experience Level: Mid to Senior (8-12 years)
Employment Type: Full-Time
Location: HYd,Bangalore,pune,Chennai,Delhi,Kolkatta
About the Role
We are seeking a Data Scientist who can independently own the full data science workflow from problem framing and data exploration through modeling, validation, deployment, and monitoring. The ideal candidate brings strong depth in developing, fine-tuning, and domain-adapting Large Language Models (LLMs) and Vision-Language Models (VLMs), alongside solid grounding in traditional predictive/statistical modeling.
Key Responsibilities
- Fine-tune and adapt LLMs/VLMs/Multimodal Models (instruction tuning, domain adaptation, alignment) using parameter-efficient/SFT/DAPT methods.
- Fine-tune vision and multimodal models (ViT, CLIP, diffusion-based architectures) for classification, detection, and segmentation tasks for vision AI and Reasoning/CoT/QnA/Decision making task for Language models.
- Curate, clean, and construct high-quality training and evaluation datasets, including synthetic data generation and annotation pipelines.
- Design and run experiments to optimize hyperparameters, training strategies, and model performance under compute constraints.
- Build and maintain reproducible training pipelines with proper versioning, logging, and experiment tracking.
- Define and implement evaluation frameworks and benchmarks (accuracy, perplexity, task-specific and human-in-the-loop metrics).
- Apply quantization, distillation, and optimization techniques to reduce inference cost and latency for production.
- Collaborate with ML engineers to deploy tuned models and monitor for drift, degradation, and regressions.
- Research and prototype emerging fine-tuning and alignment techniques (RLHF, DPO, RAG-based adaptation).
- Document methodology, findings, and best practices to scale fine-tuning workflows across teams.
Required Skills & Experience
- Strong foundation in statistics, machine learning, and predictive modeling techniques, with hands-on project experience.
- Hands-on experience developing and fine-tuning LLMs and VLMs, including parameter-efficient fine-tuning methods (LoRA/QLoRA), instruction tuning, and domain/task adaptation.
- Understanding of foundation model architectures (transformers, attention mechanisms) and the training/adaptation lifecycle from pretraining through fine-tuning.
- Experience curating and preparing training/fine-tuning datasets, including data quality, labeling, and synthetic data techniques.
- Experience with distributed and multi-GPU training, mixed precision, and memory optimization
- Familiarity with experiment tracking and MLOps, AgentOps tooling (Weights & Biases, MLflow, Phoenix etc )
- Proficiency in Python and core ML/DL libraries (PyTorch, Hugging Face Transformers, scikit-learn).
- Experience with model evaluation methodologies for generative models (benchmarking, hallucination/bias assessment) and predictive models (A/B testing, standard ML metrics).
- Working knowledge data pipelines, and cloud-based data/ML platforms (Azure, GCP, or AWS).
- Strong problem-solving skills with the ability to work independently across the full project lifecycle, from discovery to deployment.