- The AI and Data Platforms team designs, builds and runs the platforms that enable our organisation to adopt AI and data capabilities. We own the full platform lifecycle, from architecture and engineering through to operations and continuous improvement
- As an AI Platform Engineer, you’ll extend our platform capabilities, embed AI and automation into platform operations, and help teams across the business adopt AI solutions safely and effectively
- This is a hands‑on engineering role that combines feature delivery with responsibility for resilience, observability, governance and measurable outcomes
- You’ll work closely with Product Owners, Architects, Engineers, Security teams and business stakeholders to deliver secure, scalable and reliable platform capabilities
- Design, build and deliver new platform capabilities across a multi‑cloud environment
- Contribute to solution design, technical decisions and implementation
- Build production‑ready services, APIs, automation and platform components
- Apply modern engineering practices, including automated testing, CI/CD, infrastructure as code, security by design and observability
- Take platform capabilities from initial design through to production use and ongoing improvement
- Apply AI and agentic capabilities to improve platform operations and engineering workflows
- Build automation and agents that support environment provisioning, onboarding, access management, fault diagnosis and remediation
- Automate the operational lifecycle of environments, workspaces, resources, agents and permissions
- Develop self‑service capabilities that reduce manual effort while maintaining appropriate controls
- Evaluate emerging AI capabilities and recommend practical approaches to adoption
- Design solutions that remain reliable and predictable when failures occur
- Build for known failure modes using appropriate retry, isolation, recovery and service degradation patterns
- Test recovery processes and use the results to strengthen platform resilience
- Implement monitoring, alerting, telemetry and operational dashboards
- Define and track service measures that reflect user needs and platform performance
- Investigate and resolve operational issues, continuously improving platform reliability and user experience
- Build reporting, telemetry and analytics that provide clear visibility of platform usage, cost, performance and outcomes
- Support the definition and tracking of KPIs, OKRs and service measures
- Develop reporting that demonstrates adoption and business value to technical stakeholders and senior audiences
- Use operational and user data to guide platform decisions and continuous improvement
- Work with engineering, data and business teams to design, integrate and support AI use cases
- Help teams move AI use cases into secure and reliable production environments
- Develop reusable patterns, standards, documentation and self‑service capabilities
- Support user onboarding, enablement and go‑live activities
- Share knowledge and promote consistent engineering practices across teams
- Gather feedback and use it to improve the platform and user experience
- Make sure platform capabilities align with enterprise security, governance, privacy and risk requirements
- Support the explainability, auditability and transparency of AI‑enabled capabilities
- Apply appropriate controls and guardrails throughout the AI lifecycle
- Work with data governance, cataloguing, lineage and access‑management capabilities where required
- Contribute to a strong control environment and support the responsible adoption of AI
What Success Looks Like
- Delivering secure, scalable platform capabilities used by multiple business teams
- Moving AI use cases into production with appropriate reliability, observability and governance
- Replacing routine manual activity with controlled automation and self‑service capabilities
- Improving monitoring and alerting so the team can identify and address issues before users report them
- Providing clear reporting on platform usage, cost, performance, adoption and outcomes
- Increasing organisational capability through reusable assets, documentation and knowledge sharing
- Continuously improving the reliability, efficiency and user experience of our AI and data platforms
Benefits
- 8% non‑contributory pension plus up to 5% matching
- Life Assurance & Death in Service
- Support, feedback and development programmes to help you reach your full potential
- Flexible culture and ways of working
- Focus on health & wellbeing
- 30 days holiday as standard (excluding bank holidays), with the ability to purchase an additional 5 days
- Time off when you need it includes carer’s leave, faith leave, grandparent leave and career breaks
- New parents receive up to 26 weeks of full pay while on leave
- 4-week back‑to‑work transition for new parents: work 80% of your contractual hours, be paid for 100%
- New parents can take family leave in flexible blocks within the first year of birth/adaoption
A track record of delivering technology capabilities into production and supporting their ongoing operationExperience developing and supporting cloud-native applications, APIs, services or platformsAn interest in AI, automation and agentic systems, combined with sound judgement about where they add valueExperience using logs, metrics, traces and telemetry to improve platform performance and reliabilityAbility to communicate complex technical concepts clearly to technical and non-technical audiencesExperience using infrastructure as code, CI/CD and automated delivery practicesStrong software engineering or scripting capability, with experience working in production codebasesDemonstrable experience building and operating cloud-based platforms and servicesA collaborative approach to working with engineering, data, security and business stakeholdersStrong understanding of cloud fundamentals, including networking, identity and access management, security and operational controlsEvidence of improving the dependability of production systems through practical engineering changesExperience designing solutions with resilience, scalability, observability and operability in mindExperience building AI-powered applications, agents, copilots or intelligent automationExperience with Retrieval-Augmented Generation, prompt engineering, vector search or enterprise knowledge retrievalExperience with major public cloud platforms and associated AI servicesExperience with data platform technologies and governance concepts, including metadata, lineage, cataloguing and access controlExperience delivering solutions in a regulated or controlled environmentKnowledge of containerisation, orchestration and API-first architectureFamiliarity with responsible AI, model governance and AI safety controlsExperience with MLOps, AI evaluation, monitoring and the operational support of AI solutions