Computer Vision Engineer

microTECH Global LTD

Greater London

On-site

GBP 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

microTECH Global LTD in London invites applications for a permanent, full-time AI Research Engineer role. You will contribute to frontier multimodal model breakthroughs from our King’s Cross hub, collaborating with an international research team to translate insights into production-ready systems.

Requirements include a Bachelor’s degree or higher in Computer Science, Mathematics, Statistics, or related fields, strong Python skills, and practical experience with PyTorch.

Qualifications

  • Bachelor’s degree or higher in a technical field required.
  • Strong Python and PyTorch experience needed.
  • Proven ability to develop algorithms and reason over complex models.

Responsibilities

  • Develop ViT and multimodal model architectures with improved reasoning and efficiency.
  • Advance multimodal alignment, representation learning, and long-context modeling.
  • Process large-scale multimodal data across images, videos, audio, and text.
  • Build pipelines for data cleaning, filtering, annotation, and quality control.
  • Engineer training, inference, and serving infrastructure for scalable models.
  • Collaborate with product and engineering teams to deploy and iterate models.

Skills

Python programming
PyTorch
Algorithm development
Cross-functional collaboration
Mathematical reasoning

Education

Bachelor’s degree or above in Computer Science, Mathematics, Statistics, or related technical disciplines

Tools

PyTorch

Job description

Backed by an international research team and abundant computing resources, the center focuses on core research directions including multimodal understanding and generation, vision-language large models, and embodied intelligence. This is a permanent, full-time position located in the tech hub of King's Cross, London.

Key Responsibilities:
Frontier Technical Breakthroughs
  • Develop ViT and multimodal large model architectures with improved reasoning and efficiency
  • Advance multimodal alignment, representation learning, and long-context modeling
  • Optimize model architectures for generalization and performance
Data Ecosystem Construction
  • Process large-scale multimodal data across images, videos, audio, and text
  • Build pipelines for data cleaning, filtering, annotation, and quality control
  • Construct and maintain datasets with versioning and reproducibility
  • Optimize data mixtures and sampling strategies for model training
  • Improve data quality through feedback-driven curation loops
MLLM Systems & Infrastructure
  • Optimize GPU utilization, cluster efficiency, and resource scheduling
  • Engineer training, inference, and serving infrastructure
  • Improve scalability, stability, and performance of model systems
Business Value Delivery
  • Integrate multimodal capabilities into assistant and content generation scenarios
  • Translate research into production and user-facing applications
  • Collaborate with product and engineering teams to deploy and iterate models
Person Specification:
  • Academic Background: Bachelor’s degree or above in Computer Science, Mathematics, Statistics, or related technical disciplines.
  • Technical Skills: Proficient in Python programming with strong hands-on experience in PyTorch and deep learning frameworks.
  • Core Competencies: Strong algorithm development and implementation skills, solid mathematical and logical reasoning ability, and excellent cross-functional communication and collaboration skills.
  • Traits: Self-driven and highly motivated toward advancing artificial intelligence (AI), with strong resilience and the ability to tackle challenging technical problems.
  • Strong track record of publications in top-tier AI or computer vision conferences (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR)
  • High-impact open-source projects or internship experience in leading technology companies within CV, NLP, or multimodal domains
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Computer Vision Research Scientist
Computer Vision Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000
Computer Vision Scientist | Multimodal AI / Vision-Language Models / Deep Learning
Computer Vision Scientist | Multimodal AI / Vision-Language Models / Deep Learning

European Tech Recruit • Greater London

On-site
GBP 90,000 - 120,000
Computer Vision Research Scientist
Computer Vision Research Scientist

IC Resources Recruitment • Greater London

On-site
GBP 70,000 - 110,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • Greater London

On-site
GBP 110,000 - 140,000
Research Scientist/Engineer - Multimodal AI & LLM
Research Scientist/Engineer - Multimodal AI & LLM

Adecco • City Of London

On-site
GBP 85,000 - 120,000
Computer Vision Researcher (VLM)
Computer Vision Researcher (VLM)

DeepRec.ai • Greater London

Hybrid
GBP 80,000 - 100,000
Multimodal Vision Engineer — AI Systems & Research
Multimodal Vision Engineer — AI Systems & Research

microTECH Global LTD • Greater London

On-site
GBP 70,000 - 110,000
Computer Vision Research Engineer
Computer Vision Research Engineer

Block MB • Greater London

On-site
GBP 50,000 - 80,000
High autonomy
Direct influence over the research roadmap
Opportunity to publish at top venues
Computer Vision Engineer Lead
Computer Vision Engineer Lead

Connect-AI • England

Hybrid
GBP 127,000 - 150,000
Equity share options
25 days of annual leave plus bank holidays
Comprehensive private medical coverage
+1
Multimodal Vision Research Scientist
Multimodal Vision Research Scientist

IC Resources • Greater London

On-site
GBP 70,000 - 120,000