AI Research Engineer (Multi-Modal & Vision)

tether

Dubai

On-site

AED 250,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

tether is seeking a member for its AI model team to innovate in training and optimizing vision-language models. This role covers the entire model development lifecycle, focusing on real-world deployment to drive both capability and efficiency.

Successful candidates will possess a strong background in multimodal learning, with responsibilities ranging from developing high-quality datasets to publishing research findings. A passion for pushing the boundaries of multimodal AI is essential.

Qualifications

  • Strong experience with multimodal post-training workflows.
  • Hands-on experience with parameter-efficient fine-tuning.
  • Demonstrated ability to build vision-language models with measurable results.

Responsibilities

  • Conduct end-to-end research and engineering on vision-language models.
  • Design and implement post-training pipelines.
  • Drive model efficiency and deployability in resource-constrained environments.

Skills

Multimodal post-training workflows
Parameter-efficient fine-tuning
Reinforcement learning from human feedback
Model evaluation
Open-source contributions
Research publications

Education

Degree in Computer Science or Machine Learning
MS/PhD preferred

Job description

About the Job

As a member of the AI model team, you will drive innovation in training and optimizing vision-language models with a focus on real-world deployment. Your work will span the full model development lifecycle - from data curation and training pipeline design to model evaluation and optimization - with the goal of building models that are both highly capable and practical to deploy at scale.

You will work across a wide spectrum of multimodal architectures integrating text and vision, applying state-of-the-art research to improve model quality, efficiency, and domain-specific performance. We expect you to bring a research‑driven mindset combined with strong engineering discipline - someone who can identify the right technique for a given problem, implement it rigorously, and measure its impact clearly.

You will work closely with a small, high-caliber team where your contributions will have direct and meaningful impact. If you are passionate about pushing the boundaries of what multimodal AI can achieve in production environments, this is your opportunity.

Responsibilities
  • Conduct end-to-end research and engineering on vision-language models, covering training, evaluation, and optimization across the full model development lifecycle.
  • Design and implement post-training pipelines including supervised fine-tuning, knowledge distillation, and reinforcement learning from human feedback.
  • Develop and maintain high-quality multimodal datasets, including data curation, filtering, and balancing for domain-specific tasks.
  • Drive model efficiency and deployability, adapting models for resource-constrained environments using compression and optimization techniques.
  • Design and implement evaluation frameworks and benchmarks to measure model performance, robustness, and real-world task success.
  • Build and scale training workflows across distributed GPU infrastructure.
  • Identify and resolve bottlenecks in training pipelines to achieve state-of-the-art model quality on target benchmarks.
  • Contribute to and leverage open-source ecosystems including models, datasets, and tooling to accelerate development.
  • Stay current with the latest research in multimodal learning and vision-language systems, translating relevant findings into practical improvements.
  • Publish research findings in top-tier AI conferences and journals where applicable.
Qualifications
  • Degree in Computer Science, Machine Learning, or a related field; MS/PhD preferred.
  • Strong experience with multimodal post-training workflows including supervised fine-tuning, knowledge distillation, and reinforcement learning from feedback.
  • Hands‑on experience with parameter-efficient fine‑tuning and distributed training frameworks.
  • Demonstrated ability to build and improve vision-language models with measurable results on standard benchmarks or real-world tasks.
  • Experience adapting models for resource‑constrained environments.
  • Proven open-source contributions in multimodal AI on GitHub or HuggingFace.
  • Publications at top AI conferences (NeurIPS, ICML, ICLR, CVPR, ECCV etc.).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vision-Language AI Research Engineer
Vision-Language AI Research Engineer

tether • Dubai

On-site
AED 250,000 - 350,000
Research Scientist - World Modeling
Research Scientist - World Modeling

Institute of Foundation Models • Abu Dhabi

On-site
AED 80,000 - 120,000
AI Engineer
AI Engineer

Nayeducation • Abu Dhabi

On-site
AED 120,000 - 150,000
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

Recenso • Abu Dhabi

On-site
Opportunity to lead AI innovation
Research-driven environment
Leadership exposure across teams
Multimodal AI Engineer: Vision-Language in Production
Multimodal AI Engineer: Vision-Language in Production

Tether.io • United Arab Emirates

On-site
AED 293,000 - 441,000
Vision-Language AI Research Engineer
Vision-Language AI Research Engineer

Technology Innovation Institute • United Arab Emirates

On-site
AED 400,000 - 700,000
Research Engineer - NLP
Research Engineer - NLP

Institute of Foundation Models • Abu Dhabi

On-site
Research Scientist
Research Scientist

Institute of Foundation Models • Abu Dhabi

On-site
Senior AI Engineer
Senior AI Engineer

Technology Innovation Institute • United Arab Emirates

On-site
AED 300,000 - 450,000
AI Engineer
AI Engineer

Technology Innovation Institute • United Arab Emirates

On-site
AED 257,000 - 331,000