Infrastructure Engineer

Jobtailor

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

WRITER is seeking an experienced Infrastructure Engineer to build and run scalable, fault-tolerant systems for a high-traffic enterprise generative AI platform. You will balance SRE, DevOps, and platform work, and automate tasks using Python or Go.

Candidates should have 5+ years in infrastructure/DevOps, hands-on containerisation, and experience with Helm/Terraform or Pulumi across major clouds (AWS preferred). Hybrid work with three days in the office is expected.

Qualifications

  • 5+ years of experience in infrastructure engineering, DevOps, or similar role focused on large-scale, high-availability production systems.
  • Experience running containerisation in production.
  • Experience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred).
  • Proficiency in Python or Go for automation and tooling.
  • Daily workflow includes agentic tooling (Claude Code, Droid, Codex, or internal) as a hard requirement.

Responsibilities

  • Build resilient, scalable, fault-tolerant infrastructure for a high-traffic enterprise generative AI platform.
  • Move between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shift.
  • Automate operational tasks and infrastructure management with Python or Go.
  • Design and operate infrastructure across AWS, GCP, and Azure.
  • Collaborate with product, security, and engineering peers on system design from conception through launch.
  • Lead incident response, post-mortems, and root-cause analyses.

Skills

Infrastructure Engineering
DevOps Practices
Python or Go Automation
Kubernetes
Helm
Terraform
Pulumi
AWS
GCP
Azure

Tools

Agentic Tooling
Terraform
Helm

Job description

  • Build resilient, scalable, fault-tolerant infrastructure for WRITER's high-traffic enterprise generative AI platform
  • Move between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shift
  • Automate operational tasks and infrastructure management with Python or Go
  • Design and operate infrastructure across AWS, GCP, and Azure
  • Work with Kubernetes, Helm, Terraform, and cloud and AI tooling
  • Use AI agents to investigate incidents, draft Terraform and Helm changes, write runbooks, scaffold tooling, and review pull requests
  • Encode recurring infrastructure tasks as reusable internal skills for human and agent teammates
  • Lead incident response, post-mortems, and root-cause analyses
  • Own reliability, performance, and efficiency of core services end-to-end
  • Define and uphold SLOs and error budgets and carry the on-call pager
  • Balance immediate reliability work with long-term platform, observability, cost, and reliability investments
  • Collaborate with product, security, and engineering peers on system design from conception through launch
  • Report to the director of engineering
Requirements
  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
  • Experience running containerisation in production
  • Experience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)
  • Good proficiency in Python or Go for automation and tooling
  • Daily workflow already includes agentic tooling such as Claude Code, Droid, Codex, or internal skills; this is a hard requirement
  • Demonstrated ability to challenge the status quo, identify systemic weaknesses, and propose innovative solutions to complex reliability problems
  • Ability to reason from constraints and failure modes and articulate tradeoffs in business terms
  • Ability to make reversible decisions, write rollback plans, work with monitoring and logging stacks, and stress systems safely
  • Excellent communication, collaboration, and problem-solving skills
  • Strong ownership and accountability for mission-critical systems
  • At least one end-to-end 0-to-1 infrastructure build with an attached outcome metric
  • Willingness and ability to work in person in the office 3 days per week
  • Legally authorized to work in the country where the job is located
Core Competencies

Demonstrates expertise in building and operating resilient, scalable infrastructure for high‑traffic enterprise platforms, with a strong focus on automation using Python or Go. Proven ability to lead incident response and uphold reliability standards while collaborating effectively with cross‑functional teams.

Highest‑signal resume keywords
  • Infrastructure Engineering
  • DevOps Practices
  • AWS, GCP, Azure
  • Python or Go Automation
  • Helm and Terraform
Hard Skills
  • Infrastructure Engineering
  • DevOps
  • Automation
  • Containerization
  • Incident Response
  • Root‑Cause Analysis
  • SLO Definition
  • Monitoring and Logging
  • System Design
  • High‑Availability Systems
Soft Skills
  • Excellent Communication
  • Collaboration
  • Problem-Solving
  • Ownership
  • Accountability
Industry Keywords
  • Generative AI
  • High‑Traffic Platforms
  • Production Systems
  • Reliability Engineering
  • High‑Growth Product Company
Tools & Technologies
  • Kubernetes
  • Helm
  • Terraform
  • Pulumi
  • Cloud Tooling
  • Agentic Tooling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Public Cloud Assistant Infrastructure Engineer
Public Cloud Assistant Infrastructure Engineer

Jobtailor • Leeds

On-site
GBP 42,000 - 68,000
Senior Platform Engineer
Senior Platform Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 120,000
Infrastructure engineer (UK)
Infrastructure engineer (UK)

Writer • Greater London

Hybrid
GBP 100,000 - 140,000
Generous PTO
Medical and dental insurance
Parental leave
+6
Principal Platform Engineer
Principal Platform Engineer

Jobtailor • Greater London

On-site
GBP 120,000 - 160,000
Platform Reliability Engineer — AI Infra & Cloud
Platform Reliability Engineer — AI Infra & Cloud

Jobtailor • Greater London

Hybrid
GBP 90,000 - 130,000
Infrastructure engineer (UK)
Infrastructure engineer (UK)

Neura Market • Greater London

On-site
GBP 89,160 - 133,740
Generous PTO
Health & dental insurance
Parental leave
+7
Senior Software Engineer – Agentic Development Enablement
Senior Software Engineer – Agentic Development Enablement

Jobtailor • Greater London

On-site
GBP 90,000 - 110,000
Corporate IT Engineer
Corporate IT Engineer

Jobtailor • Greater London

On-site
GBP 70,000 - 110,000
Senior Software Engineer, GoLang
Senior Software Engineer, GoLang

Jobtailor • Greater London

On-site
GBP 90,000 - 120,000
Infrastructure Engineer
Infrastructure Engineer

Atarus • Greater London

On-site
GBP 70,000 - 120,000