Senior Machine Learning Engineer (Large Systems)

EngineersOfAI

Bristol

On-site

GBP 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Graphcore, based in Bristol, is seeking a Senior Machine Learning Engineer for the Applied AI team. You will develop and optimize AI models for our specialized hardware, scaling to thousands of accelerators, and collaborate with Software and Research teams to push AI compute forward.

You should have strong ML fundamentals, Python or C++, and experience with PyTorch or JAX. You’ll design and run ML experiments, identify performance bottlenecks, and help deliver efficient kernels and distributed

Qualifications

  • Bachelor/Master’s/PhD or equivalent in ML, CS, Maths, Data Science, or related field.
  • Proficiency in deep learning frameworks like PyTorch/JAX.
  • Strong Python or C++ software development skills.
  • Expertise in deep learning from model training to optimisation and evaluation.
  • Capable of designing, executing and reporting from ML experiments.
  • Developed deep understanding of performance bottlenecks and how to overcome them.
  • Ability to move quickly in a dynamic environment.
  • Enjoy cross-functional work collaborating with other teams.
  • Strong communicator – able to explain complex technical concepts to different audiences.

Responsibilities

  • Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators.
  • Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews.
  • Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency.
  • Design and conduct experiments on novel AI methods, implement them and evaluate results.
  • Collaborate with Research, Software, and Product teams to define, build, and test Graphcore’s next generation of AI hardware.
  • Engage with the AI community and keep in touch with the latest developments in AI.

Skills

Python
C++
PyTorch/JAX
ML Optimization
Experiment Design
Performance Bottlenecks
Cross-functional
Communication

Education

Bachelor/Master/PhD

Job description

About Graphcore

At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long‑term investment, we are delivering key technology into the fast‑growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.

Job Summary

As a Senior Machine Learning Engineer in the Applied AI team at Graphcore, you will contribute to advancing AI technology by developing and optimising AI models tailored to our specialised hardware. You will work on large‑scale systems where performance is critical to the success of our projects. Working closely with the Software Development and Research teams, you will play a critical role in identifying opportunities to innovate and differentiate Graphcore’s technology. We seek engineers with strong technical skills and an understanding of AI model implementation at scale, eager to make a tangible impact in this rapidly evolving field.

The Team

The Applied AI team’s role is to be proxies for our customers; we need to understand the latest AI models, applications, and software to ensure that Graphcore’s technology works seamlessly with the AI ecosystem and at scale. We build reference applications, contribute to key software libraries e.g. optimising kernels for efficiency on our hardware, and collaborate with the Research team to develop and publish novel ideas in domains such as efficient compute, model scaling and distributed training and inference of AI models for multiple modalities and applications.

If you’re excited about advancing the next generation of AI models on cutting‑edge hardware, we’d love to hear from you!

Responsibilities and Duties
  • Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators.
  • Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews.
  • Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency.
  • Design and conduct experiments on novel AI methods, implement them and evaluate results.
  • Collaborate with Research, Software, and Product teams to define, build, and test Graphcore’s next generation of AI hardware.
  • Engage with the AI community and keep in touch with the latest developments in AI.
Candidate Profile
Essential
  • Bachelor/Master’s/PhD or equivalent experience in Machine Learning, Computer Science, Maths, Data Science, or related field.
  • Proficiency in deep learning frameworks like PyTorch/JAX.
  • Strong Python or C++ software development skills.
  • Expertise in deep learning from model training to optimisation and evaluation.
  • Capable of designing, executing and reporting from ML experiments.
  • Developed deep understanding of performance bottlenecks and how to overcome them.
  • Ability to move quickly in a dynamic environment.
  • Enjoy cross‑functional work collaborating with other teams.
  • Strong communicator – able to explain complex technical concepts to different audiences.
Desirable
  • Experience in one or more of:
    • MLOps for Kubernetes‑based clusters
    • Building production systems with large language models
    • Efficient computing based on low‑precision arithmetic.
  • Experience writing C++/Triton/CUDA kernels for performance optimisation of ML models.
  • Experience in distributed training or inference of ML models across 64+ accelerators.
  • Familiarity with HPC systems and networking including Infini.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

EngineersOfAI • Cambridge

On-site
GBP 110,000 - 140,000
Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

EngineersOfAI • Greater London

On-site
GBP 110,000 - 140,000
Senior Machine Learning Engineer (Large Systems) Bristol, UK
Senior Machine Learning Engineer (Large Systems) Bristol, UK

graphcore • Bristol

On-site
GBP 60,000 - 80,000
Competitive salary
Flexible working
Generous annual leave policy
+7
Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

Graphcore • Cambridge, City Of London, Bristol

Hybrid
GBP 110,000 - 150,000
Flexible working
Private medical insurance
Generous annual leave policy
+4
AI Research Engineer
AI Research Engineer

EngineersOfAI • Cambridge

On-site
GBP 90,000 - 130,000
AI Research Engineer
AI Research Engineer

EngineersOfAI • Bristol

On-site
GBP 90,000 - 130,000
AI Research Engineer
AI Research Engineer

EngineersOfAI • Greater London

On-site
GBP 65,000 - 90,000
Senior Software Engineer - AI Compute Libraries & Performance
Senior Software Engineer - AI Compute Libraries & Performance

Graphcore • West of England

On-site
GBP 90,000 - 130,000
Flexible working
Private medical insurance
Dental plan
+6
Senior Software Engineer - AI Compute Libraries & Performance
Senior Software Engineer - AI Compute Libraries & Performance

EngineersOfAI • Bristol

On-site
GBP 70,000 - 120,000
Infrastructure and MLOps Engineer
Infrastructure and MLOps Engineer

Graphcore • Cambridge

On-site
GBP 90,000 - 130,000
Flexible working
Private medical insurance
Dental plan
+6