Senior Research Engineer - Multimodal & Video Foundation Model

Tether.io

United Kingdom

On-site

GBP 60,000 - 80,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading AI technology company is seeking a Senior Research Engineer to drive innovation in multimodal and video foundation models. The successful candidate will design AI architectures, develop training pipelines, and collaborate cross-functionally to translate research into real-world applications. This role demands a strong background in Computer Science, expertise in Python and PyTorch, and experience with various data modalities. A PhD is a plus but not required.

Qualifications

  • Expertise in Python & PyTorch, with experience in the full development pipeline.
  • Experience with large-scale text data or multimodal data.
  • Hands-on experience in developing or benchmarking language models.

Responsibilities

  • Pioneer multimodal and video-centric research contributing to prototypes.
  • Design and implement novel AI architectures for multimodal models.
  • Collaborate with teams to translate research into production-grade solutions.

Skills

Python
PyTorch
Machine Learning
Computer Vision

Education

Bachelor’s degree in Computer Science or related field

Job description

Overview

Senior Research Engineer - Multimodal & Video Foundation Model

As a member of the AI model team, you will drive innovation in architecture development for cutting-edge models of various scales, including small, large, and multi-modal systems. Your work will enhance intelligence, improve efficiency, and introduce new capabilities to advance the field.

Responsibilities
  • Pioneer multimodal and video-centric research that moves fast and breaks ground, contributing directly to usable prototypes and scalable systems.
  • Design and implement novel AI architectures for multimodal language models, integrating text, visual, and audio modalities.
  • Engineer scalable training and inference pipelines optimized for large-scale multimodal datasets and distributed GPU systems across thousands of GPUs.
  • Optimize systems and algorithms for efficient data processing, model execution, and pipeline throughput.
  • Build modular tools for preprocessing, analyzing, and managing multimodal data assets (e.g., images, video, text).
  • Collaborate cross-functionally with research and engineering teams to translate cutting-edge model innovations into production-grade solutions.
  • Prototype generative AI applications showcasing new capabilities of multimodal foundation models in real-world products.
  • Develop benchmarking tools to rigorously evaluate model performance across diverse multimodal tasks.
Qualifications
  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience
  • Expertise in Python & PyTorch, including practical experience working with the full development pipeline from data processing & data loading to training, inference, and optimization.
  • Experience working with large-scale text data, or (bonus) interleaved data spanning audio, video, image, and/or text.
  • Direct hands-on experience in developing or benchmarking at least one of the following topics: LLMs, Vision Language Models, Audio Language Models, generative video models
Nice to have skills
  • PhD in Computer Vision, Machine Learning, NLP, Computer Science, Applied Statistics, or a closely related field
  • Demonstrated expertise in computer vision, video generation foundation model and/or multimodal research.
  • First-author publications at leading AI conferences such as CVPR, ICCV, ECCV, ICML, ICLR, NeurIPS etc.
Important information for candidates
  • Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles: Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
  • Double-check email addresses. All communication from us will come from emails ending in @tether.to or @tether.io
  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
Job details
  • Seniority level: Not Applicable
  • Employment type: Full-time
  • Job function: Information Technology
  • Industries: Technology, Information and Internet
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Engineer (Pre-training - LLM & Multi-Modal) - 100% Remote Worldwide
AI Research Engineer (Pre-training - LLM & Multi-Modal) - 100% Remote Worldwide

Tether.io • United Kingdom

On-site
GBP 90,000 - 150,000
AI Research Engineer (Pre-training - LLM & Multi-Modal) - 100% Remote Worldwide
AI Research Engineer (Pre-training - LLM & Multi-Modal) - 100% Remote Worldwide

Tether • United Kingdom

Remote
GBP 111,000 - 141,000
Senior Multimodal & Video Foundation AI Engineer
Senior Multimodal & Video Foundation AI Engineer

Tether.io • United Kingdom

On-site
GBP 60,000 - 80,000
Computer Vision Engineer
Computer Vision Engineer

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

SLAMcore • United Kingdom

Remote
GBP 120,000 - 180,000
Remote AI Research Engineer - LLM & Multimodal Architect
Remote AI Research Engineer - LLM & Multimodal Architect

Tether.io • United Kingdom

On-site
GBP 90,000 - 150,000
Senior AI Research Scientist - Multimodal Models
Senior AI Research Scientist - Multimodal Models

IC Resources • Greater London

On-site
GBP 70,000 - 120,000
AI Research Scientist
AI Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000
Remote AI Research Engineer — LLM & Multimodal
Remote AI Research Engineer — LLM & Multimodal

Tether • United Kingdom

Remote
GBP 111,000 - 141,000
Research Engineer, Pretraining Scaling - London
Research Engineer, Pretraining Scaling - London

Anthropic • Greater London

On-site
GBP 250,000 - 435,000
Equity benefits
Visa sponsorship