Staff Software Engineer, Platform

United States Digital Space LLC

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Scale is seeking a Staff Software Engineer in London to own end-to-end context and memory systems for agentic AI platforms. You will design and optimize backend systems, retrieval, and memory strategies, while guiding evaluation methods and technical standards.

Collaboration with SWE and ML teams is essential to ensure reliability at scale. You will work across memory, retrieval, and evaluation, balancing software engineering with applied ML expertise, and contributing to production-grade agent

Qualifications

  • 8+ years of engineering experience, with a multi-year track record owning production systems end to end.
  • Experience with ML and information retrieval problems in modern AI systems, including embeddings, vector search, retrieval augmented generation, fine-tuning, or memory architectures.
  • Proficiency in Python and infrastructure for production ML/agentic systems (containers, cloud platforms, CI/CD).
  • Ability to set technical standards and influence decisions beyond own team through design reviews or mentoring.

Responsibilities

  • Own large, ambiguous problems in context and memory end to end, from design to production.
  • Architect core primitives for context retrieval, storage, and reasoning across long-running sessions.
  • Design and maintain evaluation methodologies for memory and retrieval quality.
  • Set technical standards and patterns adopted by the team, including failure handling.
  • Collaborate with SWE and ML engineering peers along the roadmap across multiple workstreams.
  • Debug and resolve severe production issues related to context and memory, including cross-tenant leakage.

Skills

Python
Distributed systems
ML systems
Information retrieval
Vector search
Embeddings
Cloud platforms
CI/CD

Tools

Containers
Vector stores
Big data tooling

Job description

**Staff Software Engineer**London, UK

About the role

Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities. This role owns the context and memory capabilities within AIS, including their correctness, performance, and evaluation.

We are looking for a Staff Engineer who can own hard technical problems end to end, from the software systems that serve context to agents through to the evaluation choices that determine whether that context is actually useful. This role suits someone comfortable moving between distributed systems engineering and applied ML, because the team is built the same way, with software engineers and ML engineers working the same roadmap rather than two separate tracks.

What you'll do
  • Own large, ambiguous problems in context and memory end to end, from design through production, including the backend systems, the retrieval and memory algorithms, and the evaluation that proves they work.
  • Architect the core primitives that let agents retrieve, store, and reason over long-running and cross-session context, including knowledge base retrieval, vector stores, and memory strategies.
  • Design and maintain the evaluation methodology for memory and retrieval quality, including the rubrics and benchmarks that catch regressions before customers do.
  • Set technical standards that other engineers on the team adopt, whether that is an architectural pattern, an eval practice, or an approach to failure handling under partial system failure.
  • Partner with SWE and MLE peers on the same roadmap, and coordinate with AIS's other workstreams (orchestration, evaluation and oversight, systems optimisation) where memory and context intersect their scope.
  • Debug and resolve the most severe production issues tied to context and memory, including incorrect retrieval, stale or leaked context across tenants, and degraded relevance at scale.
What we look for
  • 8+ years of engineering experience, with a multi-year track record owning production systems end to end, not just implementing scoped work handed to you by another team.
  • Direct experience with the ML and information retrieval problems underneath modern AI systems, such as embeddings, vector search, retrieval augmented generation, fine-tuning, or agent memory architectures, and the judgement to choose between them for a given problem.
  • Comfort owning both sides of the stack. You do not need to be equally deep in both software engineering and applied ML, but you need enough range in each to make good calls without waiting for a specialist.
  • Proficiency in Python, and experience with the infrastructure that production ML and agentic systems run on (containers, cloud platforms, and CI/CD).
  • Track record of setting technical standards that outlived the project they were built for, and of influencing decisions beyond your own team through design reviews, technical writing, or direct mentorship of software and ML engineers.
Ideally you'd have
  • Experience scaling products at hyper growth startups
  • Experience with agent orchestration frameworks and multi-agent systems in production.
  • Contributions to open-source agentic AI projects.

***PLEASE NOTE:Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.*

About Us:

*At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst& Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications.*

*We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.*

*We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at ----- Please see the United States Department of Labor's Know Your Rights poster for additional information.*

*We comply with the United States Department of Labor's Pay Transparency provision.*

*PLEASE NOTE: We collect, retain and use personal data for our professional business purposes, including notifying you of job opportunities that may be of interest and sharing with our affiliates. We limit the personal data we collect to that which we believe is appropriate and necessary to manage applicants’ needs, provide our services, and comply with applicable laws. Any information we collect in connection with your application will be treated in accordance with our internal policies and programs designed to protect personal data. Please see our privacy policy for additional information.*

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Platform
Staff Software Engineer, Platform

Scale AI • Greater London

On-site
GBP 120,000 - 160,000
Staff Software Engineer, Platform
Staff Software Engineer, Platform

scaleai • Greater London

On-site
GBP 120,000 - 180,000
Machine Learning Engineer, Platform
Machine Learning Engineer, Platform

scaleai • Greater London

On-site
GBP 90,000 - 120,000
Staff Engineer: Context & Memory for AI Agents
Staff Engineer: Context & Memory for AI Agents

scaleai • Greater London

On-site
GBP 120,000 - 180,000
AI Infrastructure Engineer, Sandbox Platform
AI Infrastructure Engineer, Sandbox Platform

Scale AI • Greater London

On-site
GBP 120,000 - 170,000
Machine Learning Engineer, Platform
Machine Learning Engineer, Platform

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 150,000
Staff Applied AI Engineer
Staff Applied AI Engineer

Scale AI • Greater London

On-site
GBP 120,000 - 180,000
Staff Engineer: Context & Memory for Agentic AI
Staff Engineer: Context & Memory for Agentic AI

United States Digital Space LLC • Greater London

On-site
GBP 120,000 - 180,000
Staff Applied AI Engineer London, UK Apply →
Staff Applied AI Engineer London, UK Apply →

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 180,000
AI Infrastructure Engineer, Sandbox Platform
AI Infrastructure Engineer, Sandbox Platform

scaleai • Greater London

Hybrid
GBP 110,000 - 160,000