Senior AI Research Scientist

Mindfire

New Delhi

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Frontier Model Architecture Research seeks researchers to design and evaluate neural-network architectures for language, reasoning, and multimodal workloads. You will explore Transformers, SSMs, diffusion, and mixture-of-experts, with an emphasis on modular, memory-augmented designs.

You will translate research into reproducible prototypes and production-ready experiments, contribute to evaluation suites, and collaborate across teams to advance scalable, privacy-conscious AI systems.

Responsibilities

  • Design, implement, and evaluate new neural-network architectures for language, reasoning, multimodal generation, and agentic workloads.
  • Research strengths and limitations of Transformers, SSMs, diffusion models, recurrent architectures, mixture-of-experts models, and hybrid designs.
  • Develop architectures that combine attention, state-space mechanisms, recurrence, memory, sparse routing, retrieval, and modular components where appropriate.
  • Explore efficient long-context methods, external and recurrent memory, adaptive computation, speculative execution, and improved reasoning techniques.
  • Investigate diffusion and flow-based approaches for text, image, audio, video, structured data, and multimodal generation.
  • Translate promising research papers and mathematical concepts into reproducible prototypes and production-relevant experiments.
  • Design scalable pipelines for creating high-quality synthetic examples, reasoning traces, preferences, critiques, simulations, and task-specific training data.
  • Develop teacher-student, self-training, rejection-sampling, curriculum-learning, process-supervision, and knowledge-distillation approaches.
  • Evaluate alignment and post-training methods such as supervised fine-tuning, preference optimization, reinforcement learning, and AI-generated feedback.
  • Build filtering, scoring, deduplication, provenance, contamination-detection, and quality-control systems for synthetic datasets.
  • Study and reduce the risks of feedback loops, bias amplification, hallucinations, overfitting, reward hacking, and model collapse caused by poorly controlled synthetic data.
  • Establish methods for combining synthetic, public, licensed, customer-authorized, and human-generated data while maintaining privacy and traceability.
  • Create evaluation suites covering accuracy, reasoning, robustness, safety, privacy, latency, throughput, memory, energy use, resilience, and cost.
  • Maintain reproducible research code, experiment records, model cards, dataset documentation, and technical reports.
  • Monitor relevant research and clearly communicate which developments are promising, immature, or unsuitable for systems.
  • Help define the companys research roadmap, technical strategy, and standards for scientific quality.
  • Collaborate with engineering teams to move successful research from prototype to reliable production systems.
  • Work with product leaders to connect research goals to customer needs, deployment constraints, and measurable business or social outcomes.
  • Mentor researchers and engineers, review experimental designs, and raise the technical quality of the broader team.
  • Contribute to patents, peer-reviewed publications, open research, technical demonstrations, grant proposals, and strategic partnerships when appropriate.
  • Explain complex research clearly to technical teams, customers, partners, investors, and nontechnical stakeholders.

Job description

Frontier Model Architecture Research
  • Design, implement, and evaluate new neural-network architectures for language, reasoning, multimodal generation, and agentic workloads.
  • Research the strengths and limitations of Transformers, SSMs, diffusion models, recurrent architectures, mixture-of-experts models, and hybrid designs.
  • Develop architectures that combine attention, state-space mechanisms, recurrence, memory, sparse routing, retrieval, and modular components where appropriate.
  • Explore efficient long-context methods, external and recurrent memory, adaptive computation, speculative execution, and improved reasoning techniques.
  • Investigate diffusion and flow-based approaches for text, image, audio, video, structured data, and multimodal generation.
  • Translate promising research papers and mathematical concepts into reproducible prototypes and production-relevant experiments.
Synthetic Data and Model-Generated Training
  • Design scalable pipelines for creating high-quality synthetic examples, reasoning traces, preferences, critiques, simulations, and task-specific training data.
  • Develop teacher-student, self-training, rejection-sampling, curriculum-learning, process-supervision, and knowledge-distillation approaches.
  • Evaluate alignment and post-training methods such as supervised fine-tuning, preference optimization, reinforcement learning, and AI-generated feedback.
  • Build filtering, scoring, deduplication, provenance, contamination-detection, and quality-control systems for synthetic datasets.
  • Study and reduce the risks of feedback loops, bias amplification, hallucinations, overfitting, reward hacking, and model collapse caused by poorly controlled synthetic data.
  • Establish methods for combining synthetic, public, licensed, customer-authorized, and human-generated data while maintaining privacy and traceability.
Distributed Inference and Neural Systems
  • Develop algorithms that partition, route, and execute neural-network workloads across heterogeneous devices and infrastructure.
  • Research tensor, pipeline, expert, sequence, and context parallelism for inference in resource-constrained and geographically distributed environments.
  • Design efficient methods for model sharding, layer placement, distributed KV-cache management, cache-aware scheduling, and dynamic workload migration.
  • Improve inference across mixed hardware, including data-center GPUs, consumer GPUs, CPUs, NPUs, Apple Silicon, integrated graphics, edge devices, and browser-based runtimes where appropriate.
  • Develop routing and scheduling strategies that account for memory capacity, bandwidth, latency, thermal limits, energy use, device availability, privacy rules, and workload priority.
  • Create resilient inference methods that tolerate node loss, unreliable connectivity, changing resource availability, and partial system failure.
  • Explore decentralized or collaborative neural networks in which multiple devices jointly execute models without requiring all data or model components to reside in one location.
  • Advance compression and acceleration methods, including quantization, sparsity, pruning, distillation, speculative decoding, optimized kernels, and hardware-aware model design.
Research Evaluation and Scientific Rigor
  • Form clear hypotheses, define baselines, design ablation studies, and run statistically sound experiments.
  • Create evaluation suites covering accuracy, reasoning, robustness, safety, privacy, latency, throughput, memory, energy use, resilience, and cost.
  • Identify where conventional benchmarks fail to predict real-world performance and develop task-relevant evaluations.
  • Analyze quality-versus-efficiency tradeoffs and produce evidence that guides architecture and product decisions.
  • Maintain reproducible research code, experiment records, model cards, dataset documentation, and technical reports.
  • Monitor relevant research and clearly communicate which developments are promising, immature, or unsuitable for systems.
Technical Leadership and Product Collaboration
  • Help define the companys research roadmap, technical strategy, and standards for scientific quality.
  • Collaborate with engineering teams to move successful research from prototype to reliable production systems.
  • Work with product leaders to connect research goals to customer needs, deployment constraints, and measurable business or social outcomes.
  • Mentor researchers and engineers, review experimental designs, and raise the technical quality of the broader team.
  • Contribute to patents, peer-reviewed publications, open research, technical demonstrations, grant proposals, and strategic partnerships when appropriate.
  • Explain complex research clearly to technical teams, customers, partners, investors, and nontechnical stakeholders.
  • Experience with distributed training or inference systems and parallel-computing frameworks.
  • Hands-on work with model parallelism, expert parallelism, distributed KV caches, inference schedulers, or heterogeneous clusters.
  • Experience with architectures such as Transformers, selective SSMs, Mamba-style systems, mixture-of-experts models, diffusion models, or hybrid combinations.
  • Knowledge of inference engines and optimization stacks such as vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, Triton, CUDA, ROCm, MLX, ONNX Runtime, WebGPU, or similar technologies.
  • Experience with quantization, low-rank adaptation, dynamic adapters, pruning, sparsity, speculative decoding, or custom kernels.
  • Experience creating or governing synthetic datasets at scale, including quality scoring, safety filtering, provenance, and contamination controls.
  • Familiarity with alignment and post-training techniques, including preference optimization, reinforcement learning, AI feedback, red teaming, and safety evaluation.
  • Experience with multimodal models spanning language, images, audio, video, documents, or structured data.
  • Knowledge of privacy-preserving, federated, decentralized, on-device, or edge machine learning.
  • A record of leading ambiguous research projects and helping others turn exploratory work into reliable systems.

Disclaimer: This job description has been sourced from a public domain and may have been modified by Naukri.com to improve clarity for our users. We encourage job seekers to verify all details directly with the employer via their official channels before applying.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Research Scientist
Senior AI Research Scientist

Mindfire Solutions • Khordha

On-site
INR 3,000,000 - 6,000,000
Senior AI Research Scientist
Senior AI Research Scientist

Mindfire Solutions • India

On-site
INR 3,500,000 - 7,000,000
AI Research Scientist
AI Research Scientist

Iflowtech Solutions • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

EnCharge AI • Delhi

On-site
INR 4,000,000 - 7,000,000
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

EnCharge AI • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

Mulya Technologies • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Scientist
AI Scientist

Epergne Solutions • Maharashtra

On-site
INR 3,500,000 - 7,000,000
AI Research Engineer
AI Research Engineer

BuildxPartners • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Yotta Data Services Private Limited • Mumbai

On-site
INR 3,500,000 - 6,000,000
AI Scientist
AI Scientist

Epergne Solutions • Delhi

On-site
INR 3,000,000 - 6,000,000