Backed by an international research team and abundant computing resources, the center focuses on core research directions including multimodal understanding and generation, vision-language large models, and embodied intelligence. This is a permanent, full-time position located in the tech hub of King's Cross, London.
Key Responsibilities:
Frontier Technical Breakthroughs
- Develop ViT and multimodal large model architectures with improved reasoning and efficiency
- Advance multimodal alignment, representation learning, and long-context modeling
- Optimize model architectures for generalization and performance
Data Ecosystem Construction
- Process large-scale multimodal data across images, videos, audio, and text
- Build pipelines for data cleaning, filtering, annotation, and quality control
- Construct and maintain datasets with versioning and reproducibility
- Optimize data mixtures and sampling strategies for model training
- Improve data quality through feedback-driven curation loops
MLLM Systems & Infrastructure
- Optimize GPU utilization, cluster efficiency, and resource scheduling
- Engineer training, inference, and serving infrastructure
- Improve scalability, stability, and performance of model systems
Business Value Delivery
- Integrate multimodal capabilities into assistant and content generation scenarios
- Translate research into production and user-facing applications
- Collaborate with product and engineering teams to deploy and iterate models
Person Specification:
- Academic Background: Bachelor’s degree or above in Computer Science, Mathematics, Statistics, or related technical disciplines.
- Technical Skills: Proficient in Python programming with strong hands-on experience in PyTorch and deep learning frameworks.
- Core Competencies: Strong algorithm development and implementation skills, solid mathematical and logical reasoning ability, and excellent cross-functional communication and collaboration skills.
- Traits: Self-driven and highly motivated toward advancing artificial intelligence (AI), with strong resilience and the ability to tackle challenging technical problems.
- Strong track record of publications in top-tier AI or computer vision conferences (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR)
- High-impact open-source projects or internship experience in leading technology companies within CV, NLP, or multimodal domains