Founding Engineer, Agent Systems

HelmGuard

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Daily team lunch and specialty coffee
Roof terrace with King's Cross views
On-site showers
AI tooling and API budgets

Job summary

HelmGuard in London is seeking an experienced backend/AI platform engineer to own the agent platform, orchestration, evals and reliability work that turns model calls into product features customers trust.

You will push frontier APIs to production, maintain strong quality standards, and collaborate across teams. Room for research-flavoured work exists as interests align.

Qualifications

  • Experience with backend engineering in TypeScript or comparable, with 1-2+ years shipping production LLM features
  • Experience with agent frameworks, tool calling, and multi-step orchestration
  • Production evals chops: dataset curation, LLM-as-judge failure modes, regression testing under model swaps
  • Strong systems thinking: async, queues, idempotency
  • Comfort being the named owner of AI quality, including saying no when needed

Responsibilities

  • Agent scaffolding: tool use, context management, sandboxing, prompt-injection defence
  • Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping
  • Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning
  • The internal standard for what "good enough to ship" means for AI features here

Skills

Backend engineering (TypeScript)
LLM features production
Agent frameworks
Tool calling
System design

Tools

Anthropic API
OpenAI API

Job description

About HelmGuard

We're building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. As more decisions and transactions run through agents, the volume of risk to manage and trust to establish is growing very fast. Today both functions are fragmented, split across internal teams, point products, and outside consultants, and rebuilt from scratch whenever someone needs an answer. HelmGuard brings them onto one platform and runs them continuously: our agents sit on top of proprietary data and reassess as conditions change, rather than at fixed checkpoints. So our customers spend their time deciding and acting on what matters, not assembling the evidence to get there.


Hundreds of billions of dollars are spent across risk management and trust building annually. These funds are going to be reallocated to agent-native solutions in the next five years, and we will capture that spend.


We've grown to seven-figure revenue within months of product launch, on the back of multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Our founders come from Palantir and academic institutions: Oxford, Stanford, and ETH. We're backed by leading UK and US institutional investors and exceptional angels from Meta, Isomorphic Labs, Palantir, SpaceXAI, and more.


We're hiring across founding-team roles for people who want outsize impact, the influence over direction and culture that comes only from joining this early, and pre-Series A equity upside.


Your Impact

We already have the best agent scaffolding and orchestration in the trust and risk space. You'll make it the best of anyone shipping enterprise agents, in any vertical. You'll embody the AI-native services thesis our customers are betting on: agents that don't assist with workflows but become the system of action for them.


The Role

You own the agent platform: the orchestration, evals, and reliability work that turns model calls into product features customers trust. The bar is not that the demo works but rather that a domain expert reading our agent's output considers it at the level of a peer. You own the technical delivery to make that possible.


This isn't a research role at its core: we consume frontier APIs and make them production-grade. We push them hard, though, hard enough that we recently found and reported a bug in the Anthropic API that took their engineers weeks to reproduce. At that level, the line between using these models and studying them gets thin, so if research-flavoured work pulls at you, there's room to follow it.


What You Will Do


  • Agent scaffolding: tool use, context management, sandboxing, prompt-injection defence


  • Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping


  • Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning


  • The internal standard for what \"good enough to ship\" means for AI features here



What you bring



  • Experience with backend engineering in TypeScript or comparable, with 1-2+ years shipping production LLM features


  • Experience with agent frameworks, tool calling, and multi-step orchestration


  • Production evals chops: dataset curation, LLM-as-judge failure modes, regression testing under model swaps


  • Strong systems thinking: async, queues, idempotency


  • Comfort being the named owner of AI quality, including saying no when needed



Nice to have



  • Anthropic, OpenAI, or open-weight APIs in production at scale


  • Prompt-injection or agent-security work


  • Background in compliance, audit, or any domain where correctness is fuzzy and stakes are high



Culture and Values

We value a diversity of perspectives and experiences. We also hold a small set of core beliefs that reflect how we operate, and share them transparently with candidates so the fit is clear from the outset.


Put Customers First. Our customers buy outcomes from us, not features. We judge every decision by whether it delivers on that promise.


Take Ownership. Founding-stage means problems don't come pre-scoped. You see something that needs doing, scope it, ship it, own the outcome. We expect this from everyone, and provide you the backing to execute on it.


Work Hard, With Gratitude. This is the most consequential window for building enterprise software in a generation. We work hard because the opportunity is rare, and we do it with gratitude for the moment, for the people we get to build it with, and for the customers willing to bet on us this early.


Say the Silly Thing. The best ideas usually start out sounding half-baked, so we'd rather you say the silly thing than sit on it. We want you opinionated and willing to argue, and just as willing to change your mind when someone makes a better case. Disagreement here is a contribution, not a risk.


Working at HelmGuard

Location. King's Cross, London (Gridiron building). We're built around in-person collaboration and expect most days to be in-office, with flexibility for the days that need it.


Compensation. Top decile for the London market, with meaningful EMI-eligible options.


Perks. Daily team lunch and specialty coffee, a roof terrace overlooking King's Cross, on-site showers for those who enjoy active commuting, and serious per-engineer AI tooling and API budgets.


Interview process. Three stages: behavioural phone screen, technical phone screen, and a paid on-site work trial. Target turnaround is under two weeks from first conversation.


Tech Stack. TypeScript, Node.js, React, Tailwind, OpenAPI, Express, Azure (Container Apps, Service Bus, Front Door, Entra ID), Postgres, Terraform, GitHub Actions, Docker. Anthropic-first AI with in-house evals and scaffolding. Claude Code throughout.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Product Engineer
Founding Product Engineer

TechTree • Greater London

On-site
GBP 120,000 - 190,000
Daily team lunch
Specialty coffee
Rooftop terrace
+2
Founding Product Engineer
Founding Product Engineer

HelmGuard • Greater London

On-site
GBP 90,000 - 140,000
Team lunch
Coffee perks
AI tooling budget
Founding Engineer, Infrastructure & Platform
Founding Engineer, Infrastructure & Platform

HelmGuard • Greater London

On-site
GBP 90,000 - 120,000
Daily team lunch and specialty coffee
Rooftop terrace overlooking King's X
On-site showers
+1
Forward Deployed Engineer
Forward Deployed Engineer

HelmGuard • Greater London

On-site
GBP 110,000 - 170,000
EMI options
Team lunch
Roof terrace
+2
Founding Engineer, Agent Systems
Founding Engineer, Agent Systems

TechTree • Greater London

On-site
GBP 60,000 - 90,000
Staff+ Software Engineer (Claude Managed Agents)
Staff+ Software Engineer (Claude Managed Agents)

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Dental and vision insurance
Relocation support
+3
Founding Engineer, Infrastructure & Platform
Founding Engineer, Infrastructure & Platform

TechTree • Greater London

On-site
GBP 100,000 - 180,000
Technical Product Manager
Technical Product Manager

Dragonfly • Greater London

On-site
GBP 85,000 - 120,000
Competitive salary
Equity
Private health insurance
+3
Forward Deployed Engineer
Forward Deployed Engineer

TechTree • Greater London

On-site
GBP 95,000 - 150,000
Product Engineer
Product Engineer

Semaloop • Greater London

On-site
GBP 80,000 - 120,000
Salary + equity
Holidays
Medical & dental