Senior Software Engineer, ML Ops

Isomorphic Labs

Greater London

On-site

GBP 75,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A biotech AI company in Greater London is seeking a Senior or Principal Software Engineer to ensure the reliability and scalability of their ML platform. The successful candidate will be responsible for leading platform reliability strategies, managing GPU/TPU infrastructure, and optimizing inference services. Applicants should have proven experience in architecting AI workloads and expertise in Google Cloud Platform. The position follows a hybrid model with in-office attendance expected three days a week.

Qualifications

  • Proven experience in architecting and managing large-scale AI/ML workloads in production.
  • Expertise in cloud compute design, specifically within Google Cloud Platform (GCP).
  • Significant experience deploying and managing complex workloads within Kubernetes.

Responsibilities

  • Own the end-to-end strategy for platform reliability focusing on accelerator (GPU/TPU) infrastructure.
  • Lead reliability work for the global job scheduler and ensure safe validation of infrastructure upgrades.
  • Architect and optimize next-generation inference services for high-throughput performance.

Skills

Architecting AI/ML workloads
Cloud compute design (GCP)
Kubernetes expertise
NVIDIA GPU knowledge
Programming skills

Job description

Senior or Principal Software Engineer, ML Platform (Stability & Infrastructure)


Your Impact

We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a pivotal role in ensuring the reliability and scalability of the foundations that make this possible.


As a Principal Engineer, you will lead the efforts to harden our systems, ensuring our groundbreaking AI is built on an unshakeable base, working closely with the research team and the Applied ML teams to ensure the infrastructure is stable, reliable and can operate with more data and larger models as we grow.


What You Will Do


  • You will own the end-to-end strategy for platform reliability, with a specific focus on our accelerator (GPU/TPU) infrastructure and workload orchestration. You will move between high-level architectural design and hands‑on systems engineering to eliminate friction in the researcher experience.

  • Lead the reliability work for our global job scheduler. You will design and implement a robust \"test harness\" to safely validate infrastructure upgrades without impacting live research.

  • Architect and optimize our next‑generation inference services. You will solve core scaling limits, ensuring high‑throughput performance and feature parity across our model serving stack.

  • Overhaul our logging and monitoring systems to provide radical visibility. You will build proactive alerting and telemetry that identifies systemic failures before they impact research workflows.

  • Improve our internal CI/CD stability, targeting a significant reduction in failure rates and significantly faster feedback loops for the engineering organization.

  • Contribute to core technical decisions on tooling and architectural design while partnering with science, product, and operations teams to align infrastructure with biotech R&D cycles.


Skills And Qualifications

Essential


  • Proven experience in architecting and managing large‑scale AI/ML workloads in a production environment.

  • Expertise in cloud compute design, specifically within Google Cloud Platform (GCP).

  • Significant experience deploying and managing complex workloads within Kubernetes (GKE).

  • Professional familiarity with NVIDIA GPU generations and the intricacies of high‑performance compute.

  • Strong programming skills and a \"reliability‑first\" approach to software development.


Nice to Have


  • A career history that spans both ML Software Engineering and Infrastructure SRE roles.

  • Experience leading multi‑disciplinary projects and navigating complex stakeholder requirements in a fast‑paced environment.

  • Familiarity with workload scheduling, ML efficiency research, and hardware benchmarking.

  • Experience with Google TPU generations and specialized ML‑driven R&D cycles.


Hybrid Working

It’s hugely important for us to share knowledge and build strong relationships with each other, and we find it easier to do this if we spend time together in person. This is why we follow a hybrid model, and you would require you to be able to come into the office 3 days a week (currently Tuesday, Wednesday, and one other day depending on which team you’re in). If you have additional needs that would prevent you from following this hybrid approach, we’d be happy to talk through these if you’re selected for an initial screening call.


Equal Employment Opportunity

We are committed to equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy or related condition (including breastfeeding) or any other basis protected by applicable law. If you have a disability or additional need that requires accommodation, please let us know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer, Research Singapore
Senior Infrastructure Engineer, Research Singapore

Linuxconfig • Greater London

Hybrid
GBP 90,000 - 140,000
Equity options
Enhanced pension contribution
Private medical insurance
+1
Software Engineer
Software Engineer

Experis • Greater London

On-site
GBP 138,000 - 153,000
ML Research Engineer, London
ML Research Engineer, London

Isomorphic Labs • Greater London

Hybrid
GBP 65,000 - 95,000
Principal Machine Learning Infrastructure Engineer
Principal Machine Learning Infrastructure Engineer

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Visa Hunt • Greater London

Hybrid
GBP 110,000 - 170,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Principal ML Platform Engineer
Principal ML Platform Engineer

Synthesia • Greater London

On-site
GBP 70,000 - 100,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
ML Ops Engineer
ML Ops Engineer

Anaplan • Greater London

On-site
GBP 90,000 - 150,000