AI Infrastructure Engineer (Storage)

CommonAI CIC

Cambridge

On-site

GBP 65,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary package
Pension

Job summary

CommonAI CIC is seeking an experienced Infrastructure Engineer to design, deploy and maintain high‑performance storage for AI and data workloads. You will combine Linux expertise with automation to deliver scalable, secure storage across on‑premises and cloud environments.

You will manage distributed storage systems such as Ceph, Lustre, or BeeGFS, optimise tiered storage, and ensure data integrity and availability while collaborating with AI platform and infrastructure teams.

Qualifications

  • Proven experience installing, configuring, and maintaining Ceph clusters in production.
  • Familiarity with Lustre or BeeGFS and cloud storage services.
  • Scripting and automation with Bash and Python; IaC tools like Terraform or Ansible.
  • Understanding of data security and compliance best practices.
  • Experience with containerized workloads and image registries (Docker/Kubernetes).

Responsibilities

  • Design and maintain storage platforms for AI/data workloads.
  • Manage distributed storage systems (Ceph, Lustre, BeeGFS).
  • Oversee tiered storage and data movement across storage tiers.
  • Ensure data integrity, availability, security across on‑prem and cloud.
  • Develop automation/monitoring scripts; integrate with cloud services.
  • Collaborate with AI platform and infra teams on performance and capacity.

Skills

Linux system administration
Analytical problem solving
Documentation
Communication

Tools

Ceph
Lustre
BeeGFS
Docker
Kubernetes
Terraform/OpenTofu
Ansible
AWS
Azure
GCP
S3 / MinIO / RADOS Gateway

Job description

CommonAI CIC is a non-profit membership organisation, founded on a belief in collaborative engineering for the safe and responsible development of foundational AI technologies. A place where AI startups, enterprises large and small, public sector bodies and academia can share resources and knowledge, to codevelop and grow businesses, fast.

We support technology-focused start ups, each with unique data management challenges, and are seeking an experienced Infrastructure Engineer to help them design, deploy and maintain high-performance storage systems for their AI and data-driven workloads. The successful candidate will combine deep experience architecting and managing distributed, cloud, and tiered storage solutions with strong Linux and automation skills.

In this role you will:

  • Design, implement, and maintain storage platforms that support large-scale AI and data pipelines
  • Manage distributed storage systems such as Ceph, Lustre, or BeeGFS
  • Oversee tiered storage architectures, optimising data movement across high-performance, object, and archival tiers
  • Ensure data integrity, availability, and security across on-premises and cloud environments
  • Develop automation and monitoring tools using Bash, Python, or similar scripting languages
  • Manage and secure container images and related storage used for AI and ML workloads
  • Integrate storage systems with public cloud services (AWS, Azure, GCP) and hybrid environments
  • Troubleshoot complex storage and data flow issues, collaborating closely with AI platform and infrastructure teams
  • Contribute to ongoing architecture improvements, performance tuning, and capacity planning
Requirements
  • Strong Linux system administration background
  • Proven experience installing, configuring, and maintaining Ceph clusters or similar technologies in a production environment
  • Familiarity with distributed filesystems (e.g., Lustre, BeeGFS) and cloud-based storage services (e.g. EC2)
  • Experience with tiered storage management and lifecycle data policies
  • Scripting and automation proficiency (e.g. Bash, Python, Terraform/OpenTofu, Ansible)
  • Understanding of data security best practices and compliance considerations
  • Experience working with container technologies (e.g. Docker, Kubernetes) and image storage registries
  • Strong analytical, troubleshooting, communication and documentation skills
We also value:
  • Knowledge of GPU compute environments or AI training infrastructure
  • Experience with monitoring and observability tools (Prometheus, Grafana, etc.)
  • Contributions to open-source storage, data management, or infrastructure projects
  • Familiarity with object storage systems (S3, RADOS Gateway, MinIO, etc.)
Benefits
  • A collaborative and supportive work environment
  • The opportunity to have a high impact in a growing organisation
  • Competitive salary package and pension
  • Professional development opportunities
  • Networking opportunities with influential people from across the tech sector and academia
  • A vibrant office environment located a few minutes walk away from Cambridge train station
CommonAI CIC is an equal opportunity employer and is committed to creating an inclusive and diverse workplace.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Storage Architect
Senior Storage Architect

CommonAI CIC • Cambridge

On-site
GBP 60,000 - 75,000
Competitive salary package
Pension
Professional development opportunities
+2
AI Storage Architect: High-Performance, Distributed Systems
AI Storage Architect: High-Performance, Distributed Systems

CommonAI CIC • Cambridge

On-site
GBP 65,000 - 90,000
Competitive salary package
Pension
Software Engineer - AI-Native Cloud Infrastructure
Software Engineer - AI-Native Cloud Infrastructure

CommonAI C.I.C. • Cambridge

On-site
GBP 45,000 - 65,000
Office near railway station
Free snacks
On-site gym
Senior Software Engineer - AI-Native Cloud Infrastructure
Senior Software Engineer - AI-Native Cloud Infrastructure

CommonAI C.I.C. • Cambridge

On-site
GBP 60,000 - 90,000
Stock options
On-site gym
Free snacks
Infrastructure Engineer (Storage, WEKA/CEPH): £200k + Bonus
Infrastructure Engineer (Storage, WEKA/CEPH): £200k + Bonus

Hunter Bond • Greater London

On-site
GBP 70,000 - 90,000
AI Infrastructure Lead Architect
AI Infrastructure Lead Architect

WeAreTechWomen • Greater London

On-site
GBP 85,000 - 120,000
Senior Software Engineer - AI-Native Cloud Infrastructure
Senior Software Engineer - AI-Native Cloud Infrastructure

CommonAI CIC • Cambridge

On-site
GBP 75,000 - 110,000
Stock options
Competitive salary
Professional development
+2
Project Technical Lead - AI Systems Simulation
Project Technical Lead - AI Systems Simulation

CommonAI Holdings Ltd • Cambridge

On-site
GBP 110,000 - 170,000
High‑impact role
Collaborative environment
Competitive salary and benefits
+2
Project Technical Lead - AI Systems Simulation
Project Technical Lead - AI Systems Simulation

CommonAI CIC • Cambridge

On-site
GBP 90,000 - 150,000
Competitive salary
Pension
Healthcare
+1
AI Infrastructure Architect
AI Infrastructure Architect

Accenture UK & Ireland • Greater London

On-site
GBP 120,000 - 170,000