Our client is an ambitious, well-funded AI start-up company building next-generation intelligent products designed for real-world use. They are seeking a Technical Lead of Machine Learning to bridge cutting-edge research with production engineering, ensuring advanced machine learning models become reliable, scalable, and high-performing systems that deliver tangible impact.
This position is responsible for driving the execution of production machine learning systems, transforming research concepts into robust, deployable solutions. Working across research, infrastructure, and product engineering, the successful candidate will lead the development of scalable ML platforms capable of operating efficiently under real-world production constraints.
Key Responsibilities
- Lead the end-to-end development of machine learning systems, including data pipelines, training workflows, evaluation frameworks, inference architecture, and production deployment.
- Fine-tune and optimise large language models using modern techniques such as LoRA, QLoRA, SFT, DPO, and model distillation.
- Design, build, and operate scalable inference infrastructure with a focus on latency, reliability, and cost efficiency.
- Develop and maintain high-quality data pipelines for both synthetic and real-world training datasets.
- Implement comprehensive evaluation frameworks covering model performance, robustness, safety, and bias in collaboration with research teams.
- Optimise production deployments through GPU acceleration, memory efficiency improvements, latency reduction, and scalable infrastructure design.
- Work closely with application engineering teams to integrate machine learning capabilities into backend services, desktop applications, and mobile platforms.
- Make pragmatic engineering decisions that enable rapid iteration while maintaining production quality.
- Deliver machine learning systems that perform reliably within real-world operational constraints, including performance, cost, safety, and reliability.
What Success Looks Like
- Research innovations are successfully translated into production-ready machine learning solutions with measurable performance improvements.
- Training pipelines, inference systems, and deployment workflows are reliable, efficient, and maintainable.
- Production issues are identified and resolved quickly, minimising impact on users.
- Engineering teams are well-supported, aligned, and able to deliver high-impact work efficiently.
- Continuous improvements to models and infrastructure are measurable, safe, and enhance the overall user experience.
- Python
- PyTorch and/or JAX
- GPU-based training and inference platforms
Ideal Background
- Proven experience designing, building, and deploying production machine learning systems used at scale.
- Strong experience working with large-scale machine learning or foundation models and an understanding of their practical limitations and failure modes.
- Excellent software engineering skills with a focus on writing robust, production-quality code.
- Experience building scalable ML infrastructure and deploying models into production environments.
- A pragmatic, hands-on approach with the ability to take ownership of complex technical challenges.
- Strong communication and collaboration skills, with experience working in small, high-performing engineering teams.