Senior Data Scientist, NLP & LLM Fine-Tuning
Who weare
Xenoss isanAI engineering and integration services company. Wehelp medium tolarge enterprises runAI transformation end-to-end: situation analysis and goal framing, data discovery and preparation, pipeline building, model development, retraining pipeline design, deployment, and support.
Webuild abroad range ofAIsolutions: user behaviour prediction, content generation, NLP, audience segmentation, pathfinding, AIassistants, edge computer vision, fraud detection, and more.
Wework with prominent companies such asMicrosoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica.
Weare inthe top 100 software companies onthe Inc.5000list.
What isthe project
Weare looking for aSenior Data Scientist for along-term In-Call Assistant initiative with aworld-leading financial services company.
The project focuses onbuilding areal-time conversationalAI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.
You will primarily work onthe signal detection and trigger layer: turning live conversation streams into structured signals, confidence scores, and routing decisions under strict latency requirements.
The broader solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, and compliance guardrails.
The data islarge and messy. Itincludes over one million speech-to-text call transcripts with PIIredaction, ASR errors, and nogold labels.
What you will do
Text data processing
- Build scalable pipelines toclean and normalise noisy ASR call transcripts.
- Handle PII-redacted text with typed placeholders. Detect and measure redaction errors.
- Segment conversations into turns. Repair turns broken byinterruptions and overlap.
- Remove near-duplicate and low-quality samples (MinHash/LSH, embedding similarity, heuristic and model-based quality filters).
- Mine and cluster objections and agent responses with embeddings and topic models.
- Build train, validation, and test sets with time-based splits and noleakage.
Supervision and labeling
- Build training targets from outcome signals: next customer reaction, call conversion, and agent performance.
- Create labels withLLM-assistedannotation. Validate them with sales subject-matter experts.
- Apply weak supervision and label-noise detection toweakly labeled data.
- Build preference datasets: chosen/rejected pairs for DPO and good/bad pools for KTO.
- Balance data across ataxonomy ofaround 50objection types.
LLM fine-tuning and post-training
- Run and improve supervised fine-tuning (SFT) with LoRA/QLoRA and full-parameter methods.
- Run preference optimization: DPO, KTO, ORPO, SimPO, and iterative on-policy DPO.
- Run reinforcement learning atturn level with GRPO (and PPO where useful).
- Train reward models for response quality and conversion likelihood.
- Design composite rewards (reward model + LLM judge + rule checks). Detect and limit reward hacking withKL control and reward ensembles.
- Run best-of-N sampling and reranking asastrong baseline.
- Train onasingle node of8× H100 80GB GPUs with DeepSpeed ZeRO-3or FSDP.
- Prototype onsmaller models, then scale tothe 100B+ MoE model.
Evaluation
- Design offline evaluation for generated responses when nogold answer exists.
- Build LLM-as-judge rubrics: relevance, empathy, clarity, factual accuracy, compliance, and how easy the response istosay aloud.
- Use pairwise comparison with position swap. Calibrate judges against expert labels.
- Measure agreement between judges and experts (Cohen’s kappa, Krippendorff’s alpha).
- Estimate business impact from historical data with causal methods. Control for agent skill, channel, and customer segment.
- Report results with bootstrap confidence intervals, across time windows and data slices.
- Run ablations: input context, summary vsfull transcript, customer profile onvsoff.
- Check factual accuracy and compliance. Flag invented fees, terms, ornumbers.
Collaboration
- Work closely with the AISolution Architect, the other data scientist, and client MLteams.
- Present experiments and results clearly totechnical and business stakeholders.
You are expected tobedeeply hands-on indata, training, and evaluation.
Scope ofownership
- Own data quality for training and evaluation.
- Own ameasurable, stable evaluation framework that the client trusts.
- Deliver improved model weights that beat the current SFT model and the best-of-N baseline.
- Document every experiment: setup, data, metrics, and decisions.
What you should bring
Must have
- 5+years ofhands-on experience inMLor NLP. Atleast 2years with LLMs.
- Strong Python. Clean, tested, reproducible code.
- Hands-on fine-tuning ofopen-weight LLMs of7Bparameters ormore (for example Llama, Qwen, Mistral, GPT-OSS).
- Practical experience with SFT and atleast one preference method (DPO, KTO, ORPO, orsimilar).
- Multi-GPU training experience with DeepSpeed orFSDP.
- Hands-on experience with Hugging Face Transformers, TRL, and PEFT.
- Experience processing large, noisy text datasets (1M+documents): cleaning, deduplication, filtering.
- Experience building supervision from weak, noisy, orincomplete labels.
- Experience evaluating generative models without gold labels (LLM-as-judge, human evaluation, pairwise comparison).
- Strong statistics: confidence intervals, significance testing, bias and confounding.
- Systematic error analysis and structured experiment design.
- Clear communication with technical teams and domain experts.
Nice tohave
- RLfor LLMs: GRPO, PPO, RLHF, orRLAIF. Reward modeling.
- Fine-tuning Mixture-of-Experts models ormodels of70B+ parameters.
- Knowledge ofchat templates and response formats (for example the GPT-OSS harmony format).
- Experience with vLLMor SGLang for large-scale generation.
- Causal inference oruplift modeling onobservational data.
- ConversationalAI, contact-centre, orsales-call data.
- Speech and ASR pipelines. Working with ASR errors.
- PII redaction orprivacy-preserving data work.
- Financial services domain and compliance constraints.
- Google Cloud (Vertex AI, BigQuery).
- Publications, open-source contributions, orpublic fine-tuned models.
Operating model
- Engagement: full-time, long-term B2B contract
- Location: remote, EU-based
- Time-zone overlap: atleast 4working hours with the New York team
- Infrastructure: client environment only. Noexternal training ordata processing
- Data residency: all work stays within the client perimeter
- Delivery mode: offline model improvement first, then acontrolled live pilot and production evolution