Senior AI Research Scientist

Mindfire Solutions

Khordha

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mindfire Solutions seeks a Senior AI Research Scientist with deep expertise in foundation models, large language models, diffusion models, and advanced neural architectures. The role blends theory, experimentation, and production collaboration to advance model quality, latency, memory, and energy efficiency across diverse hardware.

You will shape the research roadmap, lead technical directions, and mentor a team while collaborating with product and engineering teams to translate breakthroughs

Qualifications

  • Strong background in modern foundation models and neural-network architectures.
  • Experience researching large language models, diffusion and flow-based generative models.
  • Proven ability to translate theory into reproducible prototypes and production collaborations.
  • Experience evaluating model quality, latency, memory use, and energy efficiency.

Responsibilities

  • Design, implement, and evaluate new neural-network architectures for language, multimodal generation, and agentic workloads.
  • Research strengths and limitations of Transformers, SSMs, diffusion models, and hybrid designs.
  • Develop architectures that combine attention, recurrence, memory, and modular components where appropriate.
  • Explore efficient long-context methods, memory, retrieval, and adaptive computation.
  • Translate research papers into reproducible prototypes and production-relevant experiments.

Skills

Foundation models
Large language models
Neural-network architectures
Diffusion models
Mamba-style architectures
Research leadership
Prototype to production

Job description

We are seeking a Senior AI Research Scientist with deep expertise in modern foundation models and advanced neural-network architectures. The ideal candidate understands large language models, diffusion and flow-based generative models, state-space models (SSMs), Transformer-SSM hybrids, mixture-of-experts systems, synthetic-data training, and distributed neural-network inference.

This is a hands‑on research role for someone who can move comfortably between mathematical theory, rigorous experimentation, prototype implementation, and production collaboration. You will investigate new architectures and training methods while helping build systems that operate efficiently across GPUs, CPUs, NPUs, personal computers, edge devices, private servers, and distributed clusters.

Your work should lead to measurable improvements in model quality, reasoning, latency, throughput, memory use, energy efficiency, privacy, reliability, and total cost of operation. You will have significant influence over product's long-term technical direction and research roadmap.

Key Research Areas
  • Large language models and multimodal foundation models
  • Transformer alternatives and attention-efficient architectures
  • State-space models, including selective SSMs and Mamba-style architectures
  • Hybrid Transformer-SSM, recurrent, sparse, and modular model designs
  • Diffusion, discrete diffusion, flow matching, and multimodal generative systems
  • Mixture-of-experts models, expert routing, modular networks, and conditional computation
  • Synthetic data, model-generated supervision, self-training, and knowledge distillation
  • Distributed inference across heterogeneous and intermittently available devices
  • Memory-efficient inference, KV-cache management, long-context execution, and model sharding
  • Quantization, sparsity, pruning, low-rank adaptation, and dynamic adapter routing
  • Agentic models, tool use, planning, reasoning, and multi-agent coordination
  • Continual learning, personalization, privacy-preserving learning, and edge AI
Core Responsibilities
  • Design, implement, and evaluate new neural-network architectures for language, reasoning, multimodal generation, and agentic workloads.
  • Research the strengths and limitations of Transformers, SSMs, diffusion models, recurrent architectures, mixture-of-experts models, and hybrid designs.
  • Develop architectures that combine attention, state-space mechanisms, recurrence, memory, sparse routing, retrieval, and modular components where appropriate.
  • Explore efficient long-context methods, external and recurrent memory, adaptive computation, speculative execution, and improved reasoning techniques.
  • Investigate diffusion and flow-based approaches for text, image, audio, video, structured data, and multimodal generation.
  • Translate promising research papers and mathematical concepts into reproducible prototypes and production-relevant experiments.
Synthetic Data and Model-Generated Training
  • Design scalable pipelines for creating high-quality synthetic examples, reasoning traces, preferences, critiques, simulations, and task-specific training data.
  • Develop teacher-student, self-training, rejection-sampling, curriculum-learning, process-supervision, and knowledge-distillation approaches.
  • Evaluate alignment and post-training methods such as supervised fine-tuning, preference optimization, reinforcement learning, and AI-generated feedback.
  • Build filtering, scoring, deduplication, provenance, contamination-detection, and quality-control systems for synthetic datasets.
  • Study and reduce the risks of feedback loops, bias amplification, hallucinations, overfitting, reward hacking, and model collapse caused by poorly controlled synthetic data.
  • Establish methods for combining synthetic, public, licensed, customer-authorized, and human-generated data while maintaining privacy and traceability.
Distributed Inference and Neural Systems
  • Develop algorithms that partition, route, and execute neural-network workloads across heterogeneous devices and infrastructure.
  • Research tensor, pipeline, expert, sequence, and context parallelism for inference in resource-constrained and geographically distributed environments.
  • Design efficient methods for model sharding, layer placement, distributed KV-cache management, cache-aware scheduling, and dynamic workload migration.
  • Develop routing and scheduling strategies that account for memory capacity, bandwidth, latency, thermal limits, energy use, device availability, privacy rules, and workload priority.
  • Create resilient inference methods that tolerate node loss, unreliable connectivity, changing resource availability, and partial system failure.
  • Explore decentralized or collaborative neural networks in which multiple devices jointly execute models without requiring all data or model components to reside in one location.
  • Advance compression and acceleration methods, including quantization, sparsity, pruning, distillation, speculative decoding, optimized kernels, and hardware-aware model design.
Research Evaluation and Scientific Rigor
  • Form clear hypotheses, define baselines, design ablation studies, and run statistically sound experiments.
  • Create evaluation suites covering accuracy, reasoning, robustness, safety, privacy, latency, throughput, memory, energy use, resilience, and cost.
  • Identify where conventional benchmarks fail to predict real-world performance and develop task-relevant evaluations.
  • Analyze quality-versus-efficiency tradeoffs and produce evidence that guides architecture and product decisions.
  • Maintain reproducible research code, experiment records, model cards, dataset documentation, and technical reports.
  • Monitor relevant research and clearly communicate which developments are promising, immature, or unsuitable for systems.
Technical Leadership and Product Collaboration
  • Help define the company's research roadmap, technical strategy, and standards for scientific quality.
  • Collaborate with engineering teams to move successful research from prototype to reliable production systems.
  • Work with product leaders to connect research goals to customer needs, deployment constraints, and measurable business or social outcomes.
  • Mentor researchers and engineers, review experimental designs, and raise the technical quality of the broader team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Research Scientist
Senior AI Research Scientist

Mindfire Solutions • India

On-site
INR 3,500,000 - 7,000,000
Senior AI Research Scientist
Senior AI Research Scientist

Mindfire • New Delhi

On-site
INR 3,500,000 - 7,000,000
Lead Principal Scientist -AI Research
Lead Principal Scientist -AI Research

Technoworkz Technologies • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Staff AI Scientist
Senior Staff AI Scientist

I00M05 Wipro GE Healthcare Private Limited • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Senior AI Engineer_ (Backend)
Senior AI Engineer_ (Backend)

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

Mulya Technologies • Delhi

On-site
INR 4,000,000 - 7,000,000
AI ML Research Engineer
AI ML Research Engineer

Litmus7 Systems Consulting • Thiruvananthapuram

On-site
INR 1,200,000 - 2,100,000
Research Engineer | Title: AI Research Engineer
Research Engineer | Title: AI Research Engineer

RiDiK • Bangalore Rural

On-site
INR 900,000 - 1,500,000
Senior Research Scientist
Senior Research Scientist

adaption • India

Hybrid
INR 3,000,000 - 6,000,000
Flexible work: In-person collaboration
Adaption Passport: Travel stipend
Lunch Stipend
+1
Lead AI/ML Engineer
Lead AI/ML Engineer

Optum • Hyderabad

On-site
INR 3,500,000 - 5,500,000