Site Reliability Engineer

Helsing

City Of London

On-site

GBP 60,000 - 80,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

A defence AI company in the UK is seeking a Site Reliability Engineer to support high-security environments. In this role, you will design and manage Kubernetes infrastructure, ensuring system reliability through observability frameworks and collaboration with security teams. The ideal candidate has expertise in cloud-native technologies and scripting, contributing to impactful AI solutions. This is a full-time position at a mid-senior level.

Qualifications

  • Experience with cloud-native workloads in on-premises or air-gapped environments.
  • High level of personal integrity and attention to detail.
  • Software engineering mindset with a passion for productivity.

Responsibilities

  • Design, implement, and manage Kubernetes infrastructure.
  • Create observability frameworks using Grafana and Prometheus.
  • Collaborate with Security teams for supply chain security.

Skills

Scripting
GitOps workflows
Kubernetes expertise
Cloud-native technologies
Observability stack
Networking concepts
MLOps platforms
Infrastructure as code
System administration
Data and telemetry pipelines

Tools

Terraform
Ansible
Grafana
Prometheus
Kubeflow
Helm
Istio
OpenTelemetry

Job description

Who We Are

Helsing is a defence AI company. Our mission is to protect our democracies. We aim to achieve technological leadership so that open societies can continue to make sovereign decisions and control their ethical standards. As democracies, we believe we have a special responsibility to be thoughtful about the development and deployment of powerful technologies like AI. We take this responsibility seriously. We are an ambitious and committed team of engineers, AI specialists and customer‑facing programme managers. We are looking for mission‑driven people to join our European teams – and apply their skills to solve the most complex and impactful problems. We embrace an open and transparent culture that welcomes healthy debates on the use of technology in defence, its benefits, and its ethical implications.

The Role

Much of our work takes place in high‑security on‑premise environments, and we are looking for a Site Reliability Engineer to support our high security environments. Your role as a Site Reliability Engineer will be to design, implement, and manage our on‑premise Kubernetes infrastructure. We are looking for engineers with a strong work ethic and prioritisation skills. We value team players who communicate clearly, share knowledge generously, and collaborate effectively to move their team — and our mission—forward.

Day‑to‑Day
  • Design and build cloud‑native infrastructure platforms on‑premises, focusing on Kubernetes‑based solutions that enable our development teams to operate services at scale.
  • Create robust observability frameworks using Grafana, Prometheus, and distributed tracing to ensure system reliability and performance.
  • Architect and implement secure, multi‑tenant Kubernetes clusters with strong access controls, policy‑as‑code governance, and zero‑trust networking between red and black network domains. Develop operators and controllers to automate infrastructure provisioning and compliance.
  • Build and maintain MLOps platforms enabling AI researchers to deploy, monitor, and scale machine learning models in production.
  • Collaborate closely with our Security teams to implement supply chain security, container scanning, and runtime protection across our cloud‑native stack.
Key Skills
  • Scripting: experience in either Python, Go, Rust or Bash/ Shell for automation and tooling.
  • Experience with GitOps workflows and CI/CD automation.
  • Kubernetes Expertise: deep experience operating production Kubernetes clusters, writing custom controllers/operators, and implementing service mesh architectures (Istio/Linkerd).
  • Cloud‑Native Technologies: hands‑on experience with CNCF ecosystem, e.g. including Helm, ArgoCD, Flux and container runtime security tools like Falco.
  • Observability Stack: expert‑level knowledge of Grafana, Prometheus, Loki, Tempo, and OpenTelemetry. Experience building custom dashboards, alerts, and SLI/SLO frameworks.
  • Networking: expert understanding of networking concepts, protocols and security.
  • MLOps Platforms: experience with Kubeflow, MLflow, or similar platforms.
  • Infrastructure as Code: proficiency with Terraform, Ansible, and Kubernetes manifest templating. Experience with policy‑as‑code tools like OPA/Gatekeeper.
  • System Administration: deep understanding of Linux/Unix system administration and highly available, distributed systems.
  • Comfortable building out data and telemetry pipelines for debugging and future‑proofing solutions.
Should Apply If You
  • Have a high level of personal integrity, reliability, and attention to detail.
  • Have a software engineering mindset with a passion for building platforms and tools that multiply developer productivity.
  • Have experience running cloud‑native workloads in on‑premises or air‑gapped environments.
  • Are willing to relocate to Munich, London, or Paris.

Helsing is an equal opportunities employer. We are committed to equal employment opportunity regardless of race, religion, sexual orientation, age, marital status, disability or gender identity. Please do not submit personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, data concerning your health, or data concerning your sexual orientation. Helsing's Candidate Privacy and Confidentiality Regime can be found here.

Seniority Level

Mid‑Senior level

Employment Type

Full‑time

Job Function

Engineering and Information Technology

Industries

Software Development

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Deployed AI Engineer
Deployed AI Engineer

Helsing • Greater London

On-site
GBP 60,000 - 95,000
Competitive compensation and stock options
Relocation support
Social and education allowances
+2
Software Engineer - Backend
Software Engineer - Backend

Helsing • Greater London

On-site
GBP 65,000 - 85,000
Competitive compensation and stock options
Relocation support
Regular company events across Europe
Staff Software Engineer
Staff Software Engineer

Helsing • Greater London

Hybrid
GBP 120,000 - 160,000
Relocation support
Education allowances
Regular company events
+2
Field Engineer
Field Engineer

Helsing • Greater London

On-site
GBP 40,000 - 60,000
Competitive compensation
Stock options
Relocation support
+1
People Operations Lead
People Operations Lead

Helsing • Greater London

On-site
GBP 90,000 - 125,000
Relocation support
Learning allowance
Health & wellness programs
+2
Field Engineer Deployed Engineering London
Field Engineer Deployed Engineering London

helsing.ai • Greater London

Hybrid
GBP 55,000 - 75,000
Relocation support
Stock options
Social and education allowances
+2
Technical Recruiter
Technical Recruiter

Helsing • Greater London

On-site
GBP 40,000 - 60,000
Relocation support
Learning allowance
Health & wellness
+3
Frontend Engineer for Impactful Defense AI Platform
Frontend Engineer for Impactful Defense AI Platform

Helsing • Greater London

On-site
GBP 50,000 - 75,000
Competitive compensation and stock options
Focus on outcomes, not time-tracking
Team Lead - Training and Mission Success - International
Team Lead - Training and Mission Success - International

Helsing • Greater London

Hybrid
GBP 90,000 - 115,000
Competitive salary and VSOP options
Relocation support: up to €2,500 and 4
weeks temporary accommodation
+7
(Senior) Product Marketing Manager Marketing & Communications Munich - Berlin
(Senior) Product Marketing Manager Marketing & Communications Munich - Berlin

helsing.ai • Greater London

Hybrid
GBP 65,000 - 85,000
Relocation support
VSOP options
Social and education allowances
+1