Principal ML Platform Engineer - Reliability & Scale
Isomorphic Labs
Greater London
Hybrid
GBP 75,000 - 100,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A biotech AI company in Greater London is seeking a Senior or Principal Software Engineer to ensure the reliability and scalability of their ML platform. The successful candidate will be responsible for leading platform reliability strategies, managing GPU/TPU infrastructure, and optimizing inference services. Applicants should have proven experience in architecting AI workloads and expertise in Google Cloud Platform. The position follows a hybrid model with in-office attendance expected three days a week.
Qualifications
Proven experience in architecting and managing large-scale AI/ML workloads in production.
Expertise in cloud compute design, specifically within Google Cloud Platform (GCP).
Significant experience deploying and managing complex workloads within Kubernetes.
Responsibilities
Own the end-to-end strategy for platform reliability focusing on accelerator (GPU/TPU) infrastructure.
Lead reliability work for the global job scheduler and ensure safe validation of infrastructure upgrades.
Architect and optimize next-generation inference services for high-throughput performance.
Skills
Architecting AI/ML workloads
Cloud compute design (GCP)
Kubernetes expertise
NVIDIA GPU knowledge
Programming skills
Job description
A biotech AI company in Greater London is seeking a Senior or Principal Software Engineer to ensure the reliability and scalability of their ML platform. The successful candidate will be responsible for leading platform reliability strategies, managing GPU/TPU infrastructure, and optimizing inference services. Applicants should have proven experience in architecting AI workloads and expertise in Google Cloud Platform. The position follows a hybrid model with in-office attendance expected three days a week.