Principal Platform Engineer

Trayport

Greater London

On-site

GBP 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Diversity & inclusion
Accommodations for applicants

Job summary

Trayport is hiring a Principal Platform Engineer to be a senior technical anchor in Platform/Operations. The role is 70% hands-on engineering, designing, building, and operating infrastructure that keeps a global trading platform running, while mentoring the wider team.

Core focus includes leading AI adoption within operations using LLM-based tooling and automation to reduce toil and improve incident response. You will also contribute as a technical leader across cross-team projects.

Qualifications

  • Operations or SRE background with highly available production platforms.
  • Deep hands-on AWS (VPC, networking, EKS) and Azure (AKS, Key Vault, networking).
  • Strong networking fundamentals: TCP/IP, routing, firewalls, load balancing, hybrid connectivity.
  • Production Kubernetes experience at scale, including day-2 operations.
  • Infrastructure-as-code and automation as default (Terraform, Ansible); scripting in Python/Go/Bash.
  • Demonstrable use of AI tooling to improve engineering or operational workflows.
  • Credibility and communication skills to mentor senior engineers and influence design decisions.
  • Understanding of different database technologies.

Responsibilities

  • Design, build, and operate cloud infrastructure across AWS, Azure, and on-prem.
  • Own network design and implementation across cloud and on-prem environments.
  • Ensure high availability with multi-region/AZ architectures and DR planning.
  • Develop and maintain infrastructure-as-code and CI/CD pipelines as products.
  • Define and monitor SLOs, error budgets, and observability standards.
  • Lead post-incident reviews and drive reliability improvements.

Skills

Cloud platforms
SRE
Kubernetes
IaC
Automation
Networking
Python/Go/Bash
AI tooling

Tools

Terraform
Ansible
Copilot

Job description

We're hiring a Principal Platform Engineer to be a senior technical anchor in our Platform/Operations function. This is a 70% hands-on engineering role: you'll design, build, and operate the infrastructure that keeps a global trading platform running, while acting as a technical mentor and design authority across the wider team.

A core part of this role is leading our adoption of AI within operations - using LLM-based tooling, agentic workflows, and automation to reduce toil, accelerate incident response, and raise the operational leverage of the whole department. We're not looking for someone to write a strategy deck about AI; we're looking for someone who will build with it.

What you'll do
Platform & cloud engineering (the 70%)
  • Design, build, and operate infrastructure across AWS (VPC, networking, EKS), Azure (AKS, AKV, networking), with an understand of on-prem environments
  • Own core networking design and implementation across within Cloud environments - connectivity, routing, DNS, load balancing, firewalls, and private links between cloud and on-prem
  • Engineer for high availability: multi-region and multi-AZ architectures, failover design, capacity planning, and disaster recovery for a platform where downtime has direct market impact
  • Build and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes platforms as products consumed by engineering teams
SRE & reliability
  • Define and drive SLOs, error budgets, and observability standards (metrics, logging, tracing) across the platform
  • Take part in post-incident reviews and help with the resulting reliability work to completion
  • Continuously reduce toil through automation - if we've done it manually twice, you're already scripting it
AI-driven operations
  • Identify, prototype, and productionise AI-assisted workflows across the department: incident triage and summarisation, runbook automation, log/alert analysis, change-risk assessment, internal knowledge tooling
  • Use AI-assisted engineering tools (Copilot, or equivalents) as a first-class part of your own workflow, and coach the team to do the same safely and effectively
  • Establish sensible guardrails for AI use in a regulated, availability-critical environment, knowing when automation should act and when it should recommend
Technical leadership (the 30%)
  • Act as a mentor to engineers across the Platform and Operations teams, raising the bar on design, code, and operational practice
  • Be a design authority on cross-team projects: review architectures, challenge assumptions, and ensure new services are built to be operable, observable, and resilient from day one
  • Contribute to the technical roadmap for the platform function, balancing reliability investment against delivery
What we're looking for
Must have
  • A background in Operations or SRE running highly available, redundant production platforms - you understand failure domains, graceful degradation, and what "five nines" costs
  • Deep hands-on experience with AWS (VPC design, networking, EKS) and Azure (AKS, Key Vault, networking) - genuinely multi-cloud, not one cloud plus a certification
  • Strong Networking fundamentals: TCP/IP, routing concepts, firewalls, load balancing, hybrid connectivity (Direct Connect / ExpressRoute, VPNs)
  • Production Kubernetes experience at scale, including day-2 operations (upgrades, capacity, security, multi-cluster)
  • Infrastructure-as-code and automation as a default working style (Terraform, Ansible, or similar; strong scripting in Python, Go, or Bash)
  • Demonstrable, practical use of AI tooling to improve engineering or operational workflows - you can show us something you've automated, accelerated, or de-toiled with it
  • The credibility and communication skills to mentor senior engineers and influence design decisions without formal authority
  • Understanding of different database technologies
Nice to have
  • Experience in trading, exchanges, market data, fintech, or another latency- and availability-sensitive domain

Trayport is committed to creating and sustaining a collegial work environment in which all individuals are treated with dignity and respect and one which reflects the diversity of the community in which we operate. We provide accommodations for applicants and employees who require it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Platform Engineer
Principal Platform Engineer

TMX Group • Greater London

On-site
GBP 140,000 - 210,000
AI Platform Engineer
AI Platform Engineer

Wave Group • Greater London

Hybrid
GBP 100,000 - 115,000
Equity compensation
Hybrid work
Lead Platform Engineer
Lead Platform Engineer

Dex • Greater London

On-site
GBP 110,000 - 150,000
Senior Platform Engineer
Senior Platform Engineer

Permutable • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Platform Engineer
Senior Platform Engineer

Provn • Glasgow

On-site
GBP 65,000 - 95,000
Senior Platform Engineer: AI-Driven Ops & Reliability
Senior Platform Engineer: AI-Driven Ops & Reliability

Trayport • Greater London

On-site
GBP 120,000 - 170,000
Diversity & inclusion
Accommodations for applicants
Senior AWS Platform Engineer
Senior AWS Platform Engineer

ReVybe IT Recruitment Limited • City Of London

Hybrid
GBP 76,000 - 90,000
Hybrid work in London office (2 days)
Benefits
AI-Driven Principal Platform Engineer - Multi-Cloud SRE
AI-Driven Principal Platform Engineer - Multi-Cloud SRE

TMX Group • Greater London

On-site
GBP 140,000 - 210,000
AI DevOps Engineer
AI DevOps Engineer

Bluecrestcapitalmanagement • Greater London

On-site
GBP 90,000 - 150,000
Platform Engineer - Bruin
Platform Engineer - Bruin

eFinancialCareers • Greater London

On-site
GBP 90,000 - 125,000