Staff Software Engineer, Platform

scaleai

Greater London

On-site

GBP 120,000 - 180,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Scale AI's Applied Intelligence Systems (AIS) team in London is seeking a Staff Software Engineer to own hard problems end to end, from backend context services to the evaluation decisions that prove usefulness.

You will design core memory and retrieval primitives, work across distributed systems and applied ML, and set technical standards adopted by the team in collaboration with SWE and MLE peers.

Qualifications

  • 8+ years of engineering experience building production systems end to end.
  • Direct experience with ML and information retrieval problems: embeddings, vector search, retrieval augmented generation, fine-tuning, or memory architectures.
  • Comfort owning both software engineering and applied ML breadth, with strong judgment to make calls.
  • Proficiency in Python and production ML infra: containers, cloud platforms, CI/CD.
  • Track record of setting technical standards and influencing decisions beyond the team.

Responsibilities

  • Own large, ambiguous problems in context and memory end to end, from design through production.
  • Architect primitives for retrieval, storage, and cross-session memory with knowledge bases and vector stores.
  • Design and maintain evaluation rubrics and benchmarks for memory and retrieval quality.
  • Set technical standards and patterns adopted across the team.
  • Collaborate with SWE and MLE peers and other AIS workstreams on memory-context intersections.
  • Debug production issues related to context, retrieval, and cross-tenant memory at scale.

Job description

Staff Software Engineer

London, UK


About the role

Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities. This role owns the context and memory capabilities within AIS, including their correctness, performance, and evaluation.


We are looking for a Staff Engineer who can own hard technical problems end to end, from the software systems that serve context to agents through to the evaluation choices that determine whether that context is actually useful. This role suits someone comfortable moving between distributed systems engineering and applied ML, because the team is built the same way, with software engineers and ML engineers working the same roadmap rather than two separate tracks.


What you'll do


  • Own large, ambiguous problems in context and memory end to end, from design through production, including the backend systems, the retrieval and memory algorithms, and the evaluation that proves they work.

  • Architect the core primitives that let agents retrieve, store, and reason over long-running and cross-session context, including knowledge base retrieval, vector stores, and memory strategies.

  • Design and maintain the evaluation methodology for memory and retrieval quality, including the rubrics and benchmarks that catch regressions before customers do.

  • Set technical standards that other engineers on the team adopt, whether that is an architectural pattern, an eval practice, or an approach to failure handling under partial system failure.

  • Partner with SWE and MLE peers on the same roadmap, and coordinate with AIS's other workstreams (orchestration, evaluation and oversight, systems optimisation) where memory and context intersect their scope.

  • Debug and resolve the most severe production issues tied to context and memory, including incorrect retrieval, stale or leaked context across tenants, and degraded relevance at scale.


What we look for


  • 8+ years of engineering experience, with a multi-year track record owning production systems end to end, not just implementing scoped work handed to you by another team.

  • Direct experience with the ML and information retrieval problems underneath modern AI systems, such as embeddings, vector search, retrieval augmented generation, fine-tuning, or agent memory architectures, and the judgement to choose between them for a given problem.

  • Comfort owning both sides of the stack. You do not need to be equally deep in both software engineering and applied ML, but you need enough range in each to make good calls without waiting for a specialist.

  • Proficiency in Python, and experience with the infrastructure that production ML and agentic systems run on (containers, cloud platforms, and CI/CD).

  • Track record of setting technical standards that outlived the project they were built for, and of influencing decisions beyond your own team through design reviews, technical writing, or direct mentorship of software and ML engineers.


Ideally you'd have


  • Experience scaling products at hyper growth startups

  • Experience with agent orchestration frameworks and multi-agent systems in production.

  • Contributions to open-source agentic AI projects.


Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.


About Us:

At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications.


We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.


We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at ----- Please see the United States Department of Labor's Know Your Rights poster for additional information.


We comply with the United States Department of Labor's Pay Transparency provision.


We collect, retain and use personal data for our professional business purposes, including notifying you of job opportunities that may be of interest and sharing with our affiliates. We limit the personal data we collect to that which we believe is appropriate and necessary to manage applicants’ needs, provide our services, and comply with applicable laws. Any information we collect in connection with your application will be treated in accordance with our internal policies and programs designed to protect personal data. Please see our privacy policy for additional information.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Platform
Staff Software Engineer, Platform

Scale AI • Greater London

On-site
GBP 120,000 - 160,000
Staff Software Engineer, Platform
Staff Software Engineer, Platform

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 180,000
Machine Learning Engineer, Platform
Machine Learning Engineer, Platform

scaleai • Greater London

On-site
GBP 90,000 - 120,000
Frontier Agents Engineer
Frontier Agents Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Staff Applied AI Engineer London, UK Apply →
Staff Applied AI Engineer London, UK Apply →

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 180,000
Software Engineer, Platform
Software Engineer, Platform

Scale AI • Greater London

On-site
GBP 90,000 - 150,000
Staff Applied AI Engineer
Staff Applied AI Engineer

Scale AI • Greater London

On-site
GBP 120,000 - 180,000
AI Infrastructure Engineer, Sandbox Platform
AI Infrastructure Engineer, Sandbox Platform

scaleai • Greater London

Hybrid
GBP 110,000 - 160,000
Machine Learning Engineer, Platform
Machine Learning Engineer, Platform

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 150,000
Engineering Manager, Infrastructure London, UK Apply →
Engineering Manager, Infrastructure London, UK Apply →

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 160,000