Platform Engineer - LLM Inference Infrastructure

Verda

Greater London

On-site

GBP 110,000 - 140,000

Full time

43 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Cash & equity
Healthcare
Lunch
Wellbeing
Global team

Job summary

Verda, a Helsinki-born, London-based AI cloud company, seeks a senior backend engineer to own Go services and the Kubernetes platform. You’ll write the Go code, Helm charts, and runbooks, and verify production behavior in a Kubernetes-native stack.

You’ll implement multi-tenancy, usage metering, and billing pipelines, build operators and control planes, and raise the bar on reliability, observability, and code quality across the platform.

Qualifications

  • 4+ years of experience in backend or cloud platform roles.

Responsibilities

  • Backend services in Go - platform API, usage API, metering and billing pipeline.
  • Kubernetes-native platform work - GitOps deployment, CRDs, controllers, progressive rollout, network policy, secret and certificate management.
  • Multi-tenancy and access control - tenant provisioning, per-project access, credential handling across environments.
  • Usage metering and billing correctness - per-tenant token accounting and reconciliation to ensure accurate billing.
  • Observability - logs, metrics and traces to aid incident response.

Skills

Go
Kubernetes
SQL
PostgreSQL
English writing

Tools

Helm charts
RBAC
GitOps
Argo CD
Flux
controller-runtime
kubebuilder
Operator SDK
KServe

Job description

At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.

We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join Verda while it's still being built - not once it's finished.

About The Role

We run a multi-tenant LLM inference platform: customers send requests to an OpenAI-compatible gateway, and we handle routing, tenancy, access control, usage metering and billing on top of our own GPU fleet.

You would own backend services and the Kubernetes platform they run on. This is not a role where infrastructure is someone else's problem, you write the Go service, the Helm chart, the network policy and the runbook, and you are the one who verifies it in production.

We are moving deliberately toward a Kubernetes-native architecture: less imperative tooling and hand-run scripts, more declarative APIs, custom resources and controllers that reconcile state. If you have wanted to build operators and control planes rather than consume them, that is the direction of this role.

Your Responsibilities
  • Backend services in Go - the platform API, the customer-facing usage API, the metering and billing pipeline. Small, focused services with real correctness requirements.
  • Kubernetes-native platform work - GitOps deployment, custom resources and controllers, progressive rollout, network policy, secret and certificate management. Moving what is currently scripted into something that reconciles.
  • Multi-tenancy and access control - tenant provisioning, per-project model access, credential handling across environments.
  • Usage metering and billing correctness - a pipeline that turns raw requests into per-tenant token accounting, plus the reconciliation that proves what we bill matches what actually happened. This is money, and it has to be right.
  • Observability - logs, metrics and traces that answer questions during an incident rather than after it.
Your key competencies
  • 4+ years of experience
  • Strong Go. You have shipped and maintained production Go services, and you are comfortable with concurrency, context propagation and error handling that fails loudly instead of silently.
  • Real Kubernetes depth. Not just kubectl apply. You understand the control loop, know why a pod is not ready without guessing, and have written Helm charts, network policies and RBAC that you then had to debug.
  • SQL and relational data modelling. PostgreSQL specifically. You can reason about transactions, indexes, migration safety.
  • You verify your work. You do not report something as working because it deployed and the health check is green. You go and prove it, and you say plainly what you did not test.
  • Clear written English. Design notes, runbooks, incident write-ups. Much of our engineering context lives in writing.
Nice to have
  • Building Kubernetes operators / controllers (controller-runtime, CRDs, kubebuilder, Operator SDK)
  • GitOps at scale - Argo CD or Flux, ApplicationSets, multi-cluster
  • LLM serving internals - vLLM, SGLang, TensorRT-LLM, KServe, or similar; GPU scheduling, batching, KV cache behaviour
  • Distributed messaging (NATS, Kafka) and event-driven pipelines
  • Traefik or Envoy/Istio at the ingress layer
  • Time-series and log stores (VictoriaMetrics/VictoriaLogs, Prometheus, ClickHouse)
  • Frontend competence (React + TypeScript) - our operator console is ours to maintain, and being able to fix it end to end is valuable
  • Billing, metering or payments systems, or anything else where being wrong is expensive
  • Python, for the gateway extension layer
Why Verda
  • Cash and equity compensation along with various fringe benefits (healthcare, lunch, wellbeing, and more).
  • Profitable operations with rapid, sustained growth.
  • 40+ nationalities, with 6 different ones on the management team.
  • A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.
Practicalities
  • Work mode: Based in Helsinki / London or remote in Europe
  • Level: Senior
  • Employment type: Full time and permanent
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Product Marketing Manager - Platform and Developer
Senior Product Marketing Manager - Platform and Developer

Verda • Greater London

On-site
GBP 90,000 - 130,000
Equity compensation
Healthcare
Lunch
+2
Junior Growth Sales Representative
Junior Growth Sales Representative

Verda • Greater London

Hybrid
GBP 40,000 - 60,000
Cash and equity compensation
Healthcare
Wellbeing and more
Product Manager, Serverless & Containers
Product Manager, Serverless & Containers

Verda • Greater London

Hybrid
GBP 60,000 - 94,000
Healthcare
Lunch
Wellbeing
Office Operations Specialist (UK)
Office Operations Specialist (UK)

Verda • Greater London

On-site
GBP 32,000 - 48,000
Equity and cash compensation
Healthcare
Wellbeing benefits
Kubernetes and Container Platform Security Specialist
Kubernetes and Container Platform Security Specialist

Verda • Greater London

Hybrid
GBP 120,000 - 180,000
Cash and equity compensation
Fringe benefits
Senior Backend Developer, IAM
Senior Backend Developer, IAM

Verda • Greater London

Hybrid
GBP 90,000 - 120,000
Lead Mechanical Engineer
Lead Mechanical Engineer

Verda • Greater London

Hybrid
GBP 90,000 - 130,000
Healthcare
Lunch
Wellbeing
Modular Product Line Engineer, EMEA
Modular Product Line Engineer, EMEA

Verda • Greater London

Hybrid
GBP 90,000 - 140,000
Cash and equity
Healthcare
Lunch
+1
Content Marketing Lead
Content Marketing Lead

Verda • Greater London

On-site
GBP 85,000 - 120,000
Architect
Architect

Verda • Greater London

Hybrid
GBP 70,000 - 120,000
Cash compensation
Equity
Healthcare
+2