AI Research Engineer (World Models & Foundation Models)
Location: Remote / Hybrid (Flexible)
Experience: 5-10 Years (Exceptional candidates with strong research/open-source contributions will also be considered)
About the Role
We are looking for an exceptional AI Research Engineer to help design and build the next generation of world models and foundation models for video understanding and intelligent agents.
This is a research engineering role, not an application development role. You'll work on inventing and implementing novel neural architectures, training models from scratch, and reproducing cutting‑edge research from leading AI conferences.
If you enjoy reading research papers, experimenting with new ideas, and pushing the boundaries of deep learning, this role is for you.
Key Responsibilities
- Design and implement novel neural network architectures for world models and foundation models.
- Build Transformer architectures completely from scratch without relying solely on existing libraries.
- Develop and improve attention mechanisms, transformer blocks, and latent representation models.
- Train large-scale deep learning models from scratch using distributed GPU infrastructure.
- Research, reproduce, and extend recent papers from NeurIPS, ICLR, CVPR, ICML, and other leading AI conferences.
- Develop world models capable of learning predictive representations from video data.
- Design latent‑space prediction models instead of conventional pixel‑space prediction.
- Rapidly prototype, evaluate, and iterate on new research ideas.
- Collaborate closely with research teams to explore new approaches in representation learning and generative modeling.
- Continuously evaluate emerging AI research and translate theoretical concepts into working implementations.
Required Technical Skills
Transformer Architecture (Mandatory)
Strong understanding of Transformer internals, including:
- Multi-Head Self-Attention
- Cross-Attention
- Causal Attention
- Masked Attention
- Rotary Position Embeddings (RoPE)
- Positional Embeddings
- Feed Forward Networks (FFN)
- LayerNorm & RMSNorm
- Residual and Skip Connections
- KV Cache
Candidates must be capable of implementing a complete Transformer architecture from scratch using PyTorch.
Vision & Representation Learning
Experience with:
- Vision Transformer (ViT)
- CLIP
- SigLIP
- DINO / DINOv2
- MAE
- VideoMAE
- Representation Learning
- Latent Embeddings
- Vector Quantization (VQ)
- VQ-VAE
- VQGAN
- Discrete Latent Spaces
- Learned Tokenizers
Video Understanding
Hands‑on experience with:
- Video Tokenization
- Temporal Embeddings
- Spatio‑Temporal Attention
- Video Transformers
- Latent Video Representations
- Predictive Video Modeling
Experience with architectures such as:
- VideoPoet
- MAGVIT
- Cosmos
- Genie
- Sora-inspired architectures
is highly desirable.
World Models
Strong understanding or implementation experience with:
- World Models (Ha & Schmidhuber)
- Dreamer (V1, V2, V3)
- MuZero
- Gato
- RT-2
- JEPA
- V-JEPA
- V-JEPA 2
- Genie
Candidates should understand the principles behind latent world modeling rather than only fine‑tuning existing large language models.
Model Training
Experience training large-scale models using:
- PyTorch
- CUDA
- Distributed Training
- FSDP
- DeepSpeed
- Mixed Precision Training
- Gradient Checkpointing
- FlashAttention
Experience optimizing GPU memory and large-scale training pipelines is preferred.
Mathematics & Machine Learning
Strong understanding of:
- Linear Algebra
- Probability & Statistics
- Information Theory
- Optimization Algorithms
- Gradient Descent & Backpropagation
- Attention Mathematics
Computer Vision
Knowledge of:
- CNN Fundamentals
- Vision Transformers
- Optical Flow
- Feature Extraction
- Video Understanding
Programming Skills
Required:
Nice to Have:
- Triton
- CUDA Kernel Programming
- C++ for Deep Learning Systems
Research Experience
Candidates should be comfortable:
- Reading research papers from NeurIPS, ICLR, ICML, CVPR, ECCV, and ICCV.
- Reproducing state‑of‑the‑art models from published research.
- Modifying existing architectures.
- Designing and validating new model architectures.
- Running rigorous experiments and analyzing results.
Preferred Research Background
Hands‑on implementation experience with some of the following papers is highly desirable:
Transformers
- Attention Is All You Need
- FlashAttention
- RoFormer (RoPE)
Vision Models
- Vision Transformer (ViT)
- CLIP
- DINO / DINOv2
- MAE
World Models
- World Models (2018)
- Dreamer V3
- MuZero
- Genie
- V-JEPA
- V-JEPA 2
- VideoPoet
Latent Models
Qualifications
- Master's or PhD in Computer Science, Artificial Intelligence, Machine Learning, Robotics, Computer Vision, or a related field preferred.
- Bachelor's degree with exceptional research or open-source contributions will also be considered.
- Strong publication record or significant open-source implementations is a plus.
Ideal Candidate You are someone who:
- Loves reading AI research papers.
- Enjoys building models from first principles.
- Thinks beyond existing architectures and proposes novel ideas.
- Is comfortable working in ambiguous research environments.
- Learns quickly through experimentation.
- Can rapidly prototype and validate new concepts.
- Has a strong mathematical foundation and research mindset.
What We're Looking For
This role is suited for a researcher who can contribute to the design of entirely new AI architecturesnot someone focused solely on fine-tuning or deploying existing models.
We're looking for an engineer who can bridge theoretical research and practical implementation to help build the next generation of intelligent world models.
If you're passionate about advancing the state of AI through original research and engineering excellence, we'd love to hear from you.