Core AI Engineer Specialist
Experience: 7+ years
Location: Remote
Role Overview:
Experience req can be 7+ years with hands-on work in the pre-transformer era - people who have built models from scratch before 2017 and not just fine-tuned existing models. Good understanding of classical ML/DL algorithms and maths behind it.
We value what you have built over where you studied. A GitHub history, a Kaggle record, merged pull requests to a training library, a paper you reimplemented and evaluated properly all of these count for more with us than the name on your degree. If your CV is thin and your repositories are not, apply anyway. We read the repositories first.
Key Responsibilities:
- Build models from scratch architectures, training loops, custom losses, data pipelines in PyTorch , not in a wrapper.
- Post-train language models: supervised fine-tuning, knowledge distillation under a real capacity gap, PEFT, and RL from preferences or verifiable rewards.
- Implement papers end to end, from an arXiv PDF to working code, and evaluate them rigorously enough that we can act on the result including when it is negative.
- Adapt general models to specific domains without destroying what they already knew.
- Cut GPU cost parallelism placement, memory, quantisation, serving-time arithmetic and measure the saving rather than asserting it.
- Track the field and adopt what is worth adopting, quickly, and tell us what is not.
What we require:
- Deep learning fundamentals. Backpropagation, optimisation and transformer architectures attention, tokenisation, embeddings understood mechanically, not as vocabulary. You should be able to explain why attention is scaled by 1/d, what a normalisation layer does to gradient flow, and what actually happens in an optimiser step.
- Hands-on PyTorch or Keras. You build and train models from scratch — not fine-tuning through an API, not Trainer with a config file.
- LLM post-training, in practice — SFT (data construction, packing, masking, and what goes wrong); knowledge distillation (forward vs reverse KL, on-policy distillation, capacity gaps); PEFT ( LoRA and QLoRA , quantisation, adapter placement and rank selection); RLHF and its descendants — DPO, GRPO, PPO, RLOO : what each optimises and how each fails.
- Engineering. Solid Python . CUDA and GPU-training basics — what is actually running on the device, and why your step time is what it is. Distributed training (DDP / FSDP) is a plus, not a requirement.
- Proven ability to implement papers end to end — arXiv to working code — and to evaluate rigorously. Rigorous means seeds, baselines, controls and an honest statement of the noise floor.
- Speed of adoption. This field moves faster than any curriculum. We need people who read, try, discard and keep the residue — and who can tell the difference between a real advance and a well-marketed one.
Mandatory Tech Stack:
- Languages: Python, C++, CUDA
- Deep Learning: PyTorch, NumPy, SciPy
- LLM / Transformers: Hugging Face Transformers, TRL, Tokenizers
- Fine-tuning & Post-training: PEFT, LoRA, QLoRA, SFT, DPO, GRPO, PPO, RLOO, Knowledge Distillation
- Training: DDP, FSDP, DeepSpeed, Mixed Precision, Gradient Checkpointing
- GPU & Performance: CUDA, NVIDIA GPUs, GPU Profiling, Quantisation, Memory & Compute Optimisation
- Inference & Serving: vLLM, TensorRT-LLM
- Data & Experimentation: Hugging Face Datasets, Pandas, Weights & Biases / MLflow
- Engineering: Linux, Git, GitHub, Docker
Core stack: Python + PyTorch + CUDA + Transformers/TRL + PEFT + LLM Post-training + Distributed Training + GPU Optimisation + vLLM.
Would you like to grow forward together?