Senior Software Engineer - Agent Safety / Evals - AI Foundations

Kraken

Greater London

On-site

GBP 100,000 - 140,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Kraken is searching for a Senior Software Engineer to join the Agent Safety / Evals team. You will own substantial parts of our evaluation, safety, and reliability stack, turning technical direction into production systems and mentoring others.

You will build internal safety services, implement guardrails, strengthen AI security and observability, and operate services in AWS while collaborating with platform, techops, and security partners.

Qualifications

  • Proven experience owning complex components or services end-to-end.
  • Production Python experience is preferred.
  • Experience with AI evaluation and safety.
  • Security and governance mindset.
  • Cloud experience with AWS.

Responsibilities

  • Build internal agent-safety systems.
  • Develop robust evaluation frameworks.
  • Implement guardrails and governance controls.
  • Strengthen AI security and observability.
  • Support verification and red-teaming.
  • Operate services in AWS.
  • Raise the engineering bar.

Skills

Senior software engineering
Python
AWS
Security & governance
Reliability & observability

Education

Bachelor's degree in Computer Science or equivalent

Job description

Help us use technology to make a big green dent in the universe! Kraken powers some of the most innovative global developments in energy. We create the technology that redefines utilities and unlocks a new energy system of the future. By optimising renewable generation, building a more intelligent grid, and empowering utilities to deliver an exceptional customer experience, our operating system is transforming the industry worldwide. It's an incredibly exciting time to work in energy. Join us on our mission to improve the lives of ONE BILLION humans within the decade and shape a cleaner, better future for everyone. AI is a key investment area for Kraken Technologies as we look to expand our existing capabilities. A crucial part of this is broadening the foundational infrastructure that enables teams across the organisation to use AI effectively and accelerate our mission. You'll work in the Agent Safety team. We build the shared platforms, harnesses, and guardrails that enable engineering and product teams to safely, reliably, and deterministically use machine learning and generative AI agents for internal systems and workflows across the business. This is a delivery-focused team that sits at the intersection of engineering, security, and internal enablement.

Where you'll fit in

We're hiring a Senior Software Engineer to join our newly formed Agent Safety / Evals team. As AI agents take on more autonomous tasks across Kraken's internal workflows, you will help ensure they do so securely, predictably, and within clearly defined operational boundaries. This is a hands-on senior individual-contributor role. Working with the Lead Software Engineer and the broader AI team, you will own substantial parts of our evaluation, safety, and reliability stack - turning technical direction into production systems, contributing to architecture decisions, and raising engineering standards through thoughtful collaboration and mentoring.

What you'll own
  • Build internal agent-safety systems: Design and implement services that improve the reliability, security, and determinism of LLMs and autonomous agents, keeping them within expected operational boundaries
  • Develop robust evaluation frameworks: Build scalable harnesses and evals for internal workflows and AI skills. Define meaningful test cases, interrogate output quality, and improve the reproducibility of the systems used to measure it
  • Implement guardrails and governance controls: Develop pre- and post-generation controls and translate agreed governance requirements into maintainable software across internal tools and platforms
  • Strengthen AI security and observability: Engineer authentication and permission patterns for internal agents, improve monitoring and tracing, and contribute to operational playbooks for AI-specific anomalies
  • Support verification and red teaming: Design tests and participate in continuous verification and red-teaming exercises to identify prompt injection, data-access, and unpredictable-behaviour risks before they reach production
  • Operate services in AWS: Deploy, run, and support high-throughput, low-latency safety services; make sound architecture, reliability, and cost trade-offs; and work effectively with platform, techops, and security partners
  • Raise the engineering bar: Contribute to technical decisions, review designs and code, mentor other engineers, and help establish pragmatic patterns that can be reused across AI Foundations
What you bring to the party
  • Strong senior-level software engineering: Proven experience owning complex components or services end to end, from design and implementation through testing, deployment, and operation
  • Deep engineering fundamentals: Strong judgement around system design, concurrency, security, testing, and architecture trade-offs. Production Python experience is preferred
  • Practical AI evaluation and safety experience: A critical understanding of LLM behaviour and experience building or using evaluation harnesses to measure output quality and reliability
  • Security and governance mindset: Experience with areas such as threat modelling, authentication and authorisation, red teaming, data-access controls, or guardrails for internal platforms
  • Cloud experience: Comfortable running services in AWS, owning reliability and scalability, and collaborating with platform, techops, and security teams
  • Clear communication and collaboration: Able to explain technical trade-offs, challenge constructively, and turn complex safety and evaluation findings into practical action
What Success Looks Like
  • Reliable delivery: You ship well-tested safety and evaluation capabilities that are adopted by internal engineering and product teams
  • High-quality evals: Evaluation harnesses produce meaningful, reproducible signals that reflect the real-world performance and safety of Kraken's internal AI tooling
  • Resilient infrastructu
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Software Engineer - Agent Safety
Lead Software Engineer - Agent Safety

Whereby • Greater London

On-site
GBP 110,000 - 150,000
Senior AI Safety Engineer — Evaluation & Guardrails
Senior AI Safety Engineer — Evaluation & Guardrails

Kraken • Greater London

On-site
GBP 100,000 - 140,000
Lead Software Engineer, AI Safety & Evals
Lead Software Engineer, AI Safety & Evals

Whereby • Greater London

On-site
GBP 110,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Octopus Energy • London

On-site
GBP 60,000 - 80,000
Founding Engineer, Agent Systems
Founding Engineer, Agent Systems

TechTree • Greater London

On-site
GBP 60,000 - 90,000
Principal AI Engineer
Principal AI Engineer

Intellias • Greater London

On-site
GBP 90,000 - 130,000
Research Engineer (AI Safety)
Research Engineer (AI Safety)

Axiōma Search • Greater London

On-site
GBP 90,000 - 180,000
Technical Programme Manager
Technical Programme Manager

Kraken • Greater London

On-site
GBP 60,000 - 80,000
Senior Software Engineer
Senior Software Engineer

The Portfolio Group • Greater London

On-site
GBP 90,000 - 120,000
Senior AI Solutions Engineer
Senior AI Solutions Engineer

Ocho People • Belfast City District

On-site
GBP 90,000 - 120,000