Get more replies from employers
Send a job-specific resume in minutes.
G-Research in London is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within the NLP Engineering team. You’ll optimise the inference stack for cost-efficiency and speed, collaborating with researchers to bring ideas to life.
This hands-on role involves profiling workloads, improving deployment across GPU architectures, and building reliable tooling for ML workloads, while communicating across research, infra and engineering teams.
G-Research in London is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within the NLP Engineering team. You’ll optimise the inference stack for cost-efficiency and speed, collaborating with researchers to bring ideas to life.
This hands-on role involves profiling workloads, improving deployment across GPU architectures, and building reliable tooling for ML workloads, while communicating across research, infra and engineering teams.