Computer Vision Research Scientist

IC Resources

Greater London

On-site

GBP 70,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IC Resources in London is seeking a Research Scientist – Computer Vision to contribute to pioneering multimodal AI research and real-world product applications. You will design ViT architectures and scale training for large multimodal models, working with an international team on state-of-the-art models spanning vision, language, and multimodal tasks.

Ideal candidates have a PhD or equivalent in a technical field, strong Python and PyTorch experience, and a proactive research mindset.

Qualifications

  • Degree in Computer Science, Mathematics, Statistics, Artificial Intelligence, or a related technical discipline.
  • Strong Python programming skills with hands-on experience using PyTorch and modern deep learning frameworks.
  • Excellent algorithmic thinking, mathematical reasoning, and problem-solving ability.
  • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
  • A proactive, self-motivated approach and enthusiasm for tackling challenging research problems.

Responsibilities

  • Design and develop Vision Transformer (ViT) and multimodal model architectures with enhanced reasoning, efficiency, and scalability.
  • Advance research in multimodal representation learning, alignment techniques, and long-context modelling.
  • Investigate scalable training approaches for large multimodal foundation models.
  • Improve model performance, robustness, and generalisation across diverse tasks.
  • Process and curate large-scale multimodal datasets comprising images, video, audio, and text.
  • Build robust pipelines for data cleaning, filtering, annotation, and quality assurance.
  • Maintain reproducible datasets through effective versioning and documentation.
  • Optimise data sampling strategies and improve dataset quality through iterative evaluation and feedback.
  • Develop distributed training systems for large-scale multimodal models.
  • Optimise GPU utilisation, resource scheduling, and training efficiency.
  • Contribute to training and inference frameworks that support scalable model development.
  • Improve the reliability, performance, and scalability of AI infrastructure.
  • Apply advanced multimodal AI capabilities to intelligent products and user-facing applications.
  • Work closely with engineering and product teams to bring research innovations into production.
  • Contribute to the continuous improvement and deployment of cutting-edge AI technologies.

Skills

Python programming
Deep learning
Algorithmic thinking
Communication
Collaboration

Education

Degree in Computer Science / Mathematics / Statistics / AI

Tools

PyTorch

Job description

A global technology company at the forefront of artificial intelligence and advanced computing. With a strong commitment to research and innovation, the organisation operates internationally, investing in cutting edge AI technologies and collaborating with leading academic and industry partners to develop next generation intelligent systems.

Our London-based AI research team is expanding and is seeking a Research Scientist – Computer Vision to contribute to pioneering research in multimodal artificial intelligence. Working alongside a highly experienced international team, you will help develop state-of-the-art models spanning computer vision, multimodal learning, and foundation models, with opportunities to translate research into impactful real-world applications.

This is a permanent, full-time position based in Central London.

Key Responsibilities
  • Design and develop Vision Transformer (ViT) and multimodal model architectures with enhanced reasoning, efficiency, and scalability
  • Advance research in multimodal representation learning, alignment techniques, and long-context modelling
  • Investigate scalable training approaches for large multimodal foundation models
  • Improve model performance, robustness, and generalisation across diverse tasks
Data & Model Development
  • Process and curate large-scale multimodal datasets comprising images, video, audio, and text
  • Build robust pipelines for data cleaning, filtering, annotation, and quality assurance
  • Maintain reproducible datasets through effective versioning and documentation
  • Optimise data sampling strategies and improve dataset quality through iterative evaluation and feedback
Systems & Infrastructure
  • Develop distributed training systems for large-scale multimodal models
  • Optimise GPU utilisation, resource scheduling, and training efficiency
  • Contribute to training and inference frameworks that support scalable model development
  • Improve the reliability, performance, and scalability of AI infrastructure
Research Translation
  • Apply advanced multimodal AI capabilities to intelligent products and user-facing applications
  • Work closely with engineering and product teams to bring research innovations into production
  • Contribute to the continuous improvement and deployment of cutting-edge AI technologies
Person specification
Essential
  • Degree in Computer Science, Mathematics, Statistics, Artificial Intelligence, or a related technical discipline
  • Strong Python programming skills with hands-on experience using PyTorch and modern deep learning frameworks
  • Excellent algorithmic thinking, mathematical reasoning, and problem-solving ability
  • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams
  • A proactive, self-motivated approach and enthusiasm for tackling challenging research problems
Desirable
  • Publications at leading AI or computer vision conferences (e.g. CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR)
  • Experience training or fine-tuning large-scale vision, language, or multimodal models
  • Contributions to open-source AI projects or research experience within industry or academic laboratories

Please contact Charles Duran for more information.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Computer Vision Research Scientist
Computer Vision Research Scientist

IC Resources Recruitment • Greater London

On-site
GBP 70,000 - 110,000
Computer Vision Scientist | Multimodal AI / Vision-Language Models / Deep Learning
Computer Vision Scientist | Multimodal AI / Vision-Language Models / Deep Learning

European Tech Recruit • Greater London

On-site
GBP 90,000 - 120,000
Computer Vision Engineer
Computer Vision Engineer

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • City Of London

On-site
GBP 85,000 - 120,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • Greater London

On-site
GBP 110,000 - 140,000
Computer Vision Research Engineer
Computer Vision Research Engineer

Block MB • Greater London

On-site
GBP 50,000 - 80,000
High autonomy
Direct influence over the research roadmap
Opportunity to publish at top venues
Multimodal Vision Research Scientist
Multimodal Vision Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000
Multimodal Vision AI Scientist: Research & Impact in London
Multimodal Vision AI Scientist: Research & Impact in London

IC Resources Recruitment • Greater London

On-site
GBP 70,000 - 110,000
Multimodal Vision Engineer — AI Systems & Research
Multimodal Vision Engineer — AI Systems & Research

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
Multimodal Vision Scientist: Vision-Language & DL Innovator
Multimodal Vision Scientist: Vision-Language & DL Innovator

European Tech Recruit • Greater London

On-site
GBP 90,000 - 120,000