Site Reliability Engineer: On-Prem Kubernetes & MLOps
Helsing
City Of London
On-site
GBP 60,000 - 80,000
Full time
14 days+
Application generator
A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Job summary
A defence AI company in the UK is seeking a Site Reliability Engineer to support high-security environments. In this role, you will design and manage Kubernetes infrastructure, ensuring system reliability through observability frameworks and collaboration with security teams. The ideal candidate has expertise in cloud-native technologies and scripting, contributing to impactful AI solutions. This is a full-time position at a mid-senior level.
Qualifications
Experience with cloud-native workloads in on-premises or air-gapped environments.
High level of personal integrity and attention to detail.
Software engineering mindset with a passion for productivity.
Responsibilities
Design, implement, and manage Kubernetes infrastructure.
Create observability frameworks using Grafana and Prometheus.
Collaborate with Security teams for supply chain security.
Skills
Scripting
GitOps workflows
Kubernetes expertise
Cloud-native technologies
Observability stack
Networking concepts
MLOps platforms
Infrastructure as code
System administration
Data and telemetry pipelines
Tools
Terraform
Ansible
Grafana
Prometheus
Kubeflow
Helm
Istio
OpenTelemetry
Job description
A defence AI company in the UK is seeking a Site Reliability Engineer to support high-security environments. In this role, you will design and manage Kubernetes infrastructure, ensuring system reliability through observability frameworks and collaboration with security teams. The ideal candidate has expertise in cloud-native technologies and scripting, contributing to impactful AI solutions. This is a full-time position at a mid-senior level.