HPC Engineer

Mercedes AMG High Performance Powertrains

Brixworth

On-site

GBP 65,000 - 90,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercedes AMG High Performance Powertrains is seeking an experienced HPC administrator to own reliability and performance of Linux-based HPC platforms used for simulation and engineering workloads. You will optimize compute, storage, networking and scheduling services to deliver scalable workload capacity and engage in technical escalation for incidents.

You will administer HPC clusters, troubleshoot across system layers, and automate tasks with Bash, Python and Ansible, translating user

Qualifications

  • Bachelor's degree or equivalent in a STEM field or IT.
  • Strong Linux system administration experience in HPC or scientific computing.
  • Experience with scheduling systems and queue management.
  • Automation scripting using Bash, Python, or similar.

Responsibilities

  • Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads.
  • Improve compute, storage, networking and scheduling services for efficient workload delivery.
  • Provide technical escalation for HPC incidents, capacity issues and bottlenecks.
  • Administer Linux-based HPC clusters including compute nodes, schedulers and shared services.
  • Troubleshoot across hardware, OS, network, storage, applications and user workflows.
  • Automate operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools.
  • Translate technical user requirements into practical service improvements.

Skills

Linux administration
HPC
Scripting (Bash/Python)
Scheduling (Slurm)
Performance tuning
Networking

Education

Bachelor's degree in a STEM field

Tools

Slurm
PBS
LSF
Ansible

Job description

  • Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads
  • Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery
  • Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems
  • Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services
  • Troubleshooting issues across hardware, OS, network, storage, applications and user workflows
  • Managing capacity, performance and availability for engineering and simulation workloads
  • Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools
  • Translating technical user requirements into practical service improvements
Purpose of the role is to…
  • Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads
  • Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery
  • Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems
  • Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services
  • Troubleshooting issues across hardware, OS, network, storage, applications and user workflows
  • Managing capacity, performance and availability for engineering and simulation workloads
  • Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools
  • Translating technical user requirements into practical service improvements
THEREFORE WE NEED YOU TO…
Be skilled at…
  • Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services
  • Troubleshooting issues across hardware, OS, network, storage, applications and user workflows
  • Managing capacity, performance and availability for engineering and simulation workloads
  • Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools
  • Translating technical user requirements into practical service improvements
Have Experience Of…
  • Supporting Linux-based HPC, scientific computing, simulation or high-throughput compute environments
  • Diagnosing workload, queue, licence, performance, data movement and application issues
  • Operating at a senior technical level in an enterprise or engineering-led environment
  • Delivering maintenance, upgrades, patching and change activity with minimal service impact
  • Working with suppliers and internal teams to resolve platform issues and improve service maturity
Demonstrate knowledge of…
  • HPC architecture, parallel workloads, scheduling, queues and resource allocation
  • Linux administration, scripting, patching and secure configuration
  • Schedulers such as Slurm, PBS, LSF or equivalent
  • Scale-out storage, file systems, backup, archive and data lifecycle management
  • Networking, interconnects, latency, bandwidth and data locality considerations
  • Monitoring, performance tuning, benchmarking and capacity forecasting
  • Security, vulnerability management, access control and compliance for shared platforms
  • Desirable: motorsport, automotive, CFD, simulation or data science experience
Hold These Qualifications…
  • Relevant degree, apprenticeship, professional qualification or equivalent technical experience
  • Relevant technical certifications, or equivalent experience, in Linux, HPC, storage, networking, automation or ITIL
Be…
  • Analytical, curious and comfortable solving complex technical problems
  • Proactive in improving resilience, reducing risk and removing operational friction
  • Structured, communicative and effective across hands-on delivery and change control
  • Collaborative, customer-focused and willing to share knowledge
Success in this role will be if you… (deliverables)
  • Maintain stable, secure and performant HPC services for critical engineering workloads
  • Improve compute and storage utilisation through effective monitoring, queue management and capacity planning
  • Resolve incidents quickly and reduce repeat issues through automation, documentation and service improvement
  • Deliver upgrades, maintenance and project work safely with clear communication and change control
  • Improve simulation throughput, data availability and user productivity
  • Define and guide strategic direction on HPC related topics
Additional Information…
  • The role combines operational support and project delivery, including planned maintenance, capacity improvement, lifecycle management and occasional out-of-hours activity
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Linux Engineer
HPC Linux Engineer

The ONE Group Ltd • Cambridge

On-site
GBP 65,000 - 110,000
HPC Engineer (Linux) - Sussex, Onsite
HPC Engineer (Linux) - Sussex, Onsite

Arden Resourcing • Pease Pottage

On-site
GBP 75,000 - 85,000
Annual bonus scheme
Enhanced pension contribution
Private medical and dental options
+2
HPC Engineer
HPC Engineer

Monash University • Brixworth

On-site
GBP 55,000 - 75,000
HPC Senior Hardware Engineer
HPC Senior Hardware Engineer

CGG Services (UK) Limited • Bolney

On-site
GBP 90,000 - 120,000
Competitive salary
Bonus scheme
Relocation sponsorship
+1
Senior Linux / HPC Engineer
Senior Linux / HPC Engineer

Randstad Technologies Recruitment • Greater London

Hybrid
GBP 65,000 - 95,000
Platform Engineer
Platform Engineer

Red Bull Racing & Red Bull Technology • Milton Keynes

On-site
GBP 70,000 - 110,000
HPC Engineer
HPC Engineer

LinuxRecruit • Greater London

On-site
GBP 50,000 - 70,000
HPC Engineer
HPC Engineer

CGG Services (UK) Limited • Crawley

On-site
GBP 65,000 - 95,000
Bonus scheme
22 days leave
Company pension
+5
Senior HPC Engineer
Senior HPC Engineer

Amentum • West of England

On-site
GBP 55,000 - 75,000
Free medical cover
Digital GP service
Enhanced parental leave pay
+1
HPC Operations Lead
HPC Operations Lead

LinuxRecruit • Greater London

On-site
GBP 100,000 - 120,000
Competitive salary
Flexible compensation
Excellent benefits