AI Research Engineer (Pre-training - LLM & Multi-Modal)

Tether

Bengaluru

On-site

INR 1,500,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tether is seeking a member of the AI model team to drive innovation in architecture development for state-of-the-art models. Your work will enhance intelligence and introduce new capabilities in AI, emphasizing large-scale pre-training for LLMs and Multi-Modal architectures on distributed servers.

Ideal candidates will have a significant background in Computer Science, preferably a PhD in NLP or Machine Learning, with hands-on experience in model development and large-scale training frameworks. Join us to push boundaries in AI performance!

Qualifications

  • Degree in Computer Science; ideally a PhD in NLP, Machine Learning, or a related field.
  • Experience with large-scale LLM or Multi-Modal pre-training runs.
  • Knowledge of transformer modifications for enhanced model performance.

Responsibilities

  • Conduct foundational pre-training for LLMs and Multi-Modal models.
  • Design and scale innovative architectures and tokenizers.
  • Source and curate large-scale datasets for pre-training.

Skills

Large Language Models (LLMs)
Multi-Modal Architectures
Pre-training Optimization
PyTorch
Hugging Face

Education

PhD in NLP or Machine Learning

Tools

NVIDIA GPUs
Distributed Training Frameworks

Job description

About the job

As a member of the AI model team, you will drive innovation in architecture development for cutting‑edge models of various scales, including small, large, and multi‑modal systems. Your work will enhance intelligence, improve efficiency, and introduce new capabilities to advance the field.

You will have a deep expertise in Large Language Model (LLM) and Multi‑Modal architectures, a strong grasp of pre‑training optimization, and a hands‑on, research‑driven approach. Your mission is to explore and implement novel techniques and algorithms that lead to groundbreaking advancements: multi‑modal data curation and alignment, strengthening baselines, and identifying and resolving existing pre‑training bottlenecks to push the limits of cross‑modal AI performance.

Responsibilities
  • Large‑Scale Pre‑Training: Conduct foundational pre‑training for LLMs and Multi‑Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers equipped with multi‑nodes and thousands of NVIDIA GPUs.

  • Architecture & Alignment Innovation: Design, prototype, and scale innovative architectures, tokenizers, and cross‑modal alignment layers to enhance model intelligence and multi‑modal understanding.

  • Data Strategy: Source, filter, and curate massive‑scale textual and multi‑modal datasets, establishing robust data pipelines for efficient pre‑training.

  • Experimental Research: Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency.

  • Optimization & Debugging: Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi‑modal alignment stability during long training runs.

  • System Scalability: Contribute to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms.

Qualifications
  • A degree in Computer Science or related field. Ideally a PhD in NLP, Machine Learning, or a related field, with a solid track record in AI R&D and publications in A* conferences.

  • Hands‑on experience contributing to large‑scale LLM or Multi‑Modal pre‑training runs on large, distributed servers equipped with thousands of NVIDIA GPUs, ensuring scalability and impactful advancements in model performance.

  • Familiarity and practical experience with large‑scale, distributed training frameworks, libraries, and tools.

  • Deep knowledge of state‑of‑the‑art transformer and non‑transformer modifications aimed at enhancing intelligence, efficiency, and scalability.

  • Strong expertise in PyTorch and Hugging Face libraries with practical experience in model development, continual pre‑training, and deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Artificial Intelligence Researcher
Artificial Intelligence Researcher

AMD • Bengaluru

On-site
INR 1,800,000 - 2,400,000
AI Engineer (LLMs, Agentic Systems & Model Training)
AI Engineer (LLMs, Agentic Systems & Model Training)

Kayana | Ordering & Payment Solutions • Mumbai

On-site
INR 1,200,000 - 2,000,000
Competitive salary and benefits
Opportunity to work with cutting-edge AI systems
Collaborative environment
+1
AIMD: AI Model Developer (SLM Specialist)
AIMD: AI Model Developer (SLM Specialist)

Indian AI Research Organisation (IAIRO) • India

On-site
INR 2,000,000 - 3,200,000
Machine Learning Specialist
Machine Learning Specialist

Recro • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Research Engineer | Title: AI Research Engineer
Research Engineer | Title: AI Research Engineer

RiDiK • Bangalore Rural

On-site
INR 900,000 - 1,500,000
Senior/Lead AI ML Engineer
Senior/Lead AI ML Engineer

Hiringhood • Vadodara

On-site
INR 2,500,000 - 4,500,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Yotta Data Services Private Limited • Mumbai

On-site
INR 3,500,000 - 7,000,000
Software Engineer - AI Software & Platform
Software Engineer - AI Software & Platform

United States Digital Space LLC • Karnataka

On-site
INR 2,500,000 - 4,500,000
Senior AI Engineer
Senior AI Engineer

Genzeon Global • Pune District

On-site
INR 2,500,000 - 5,000,000
Senior ML/AI Engineer
Senior ML/AI Engineer

Syren Cloud Inc. • Hyderabad

On-site
INR 1,800,000 - 3,000,000