HPC Linux Operations & Infrastructure Engineer

Westbury Partners

Sydney

On-site

AUD 90,000 - 120,000

Full time

36 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Westbury Partners in Sydney NSW is seeking a hands‑on Linux systems engineer to provide operational support for large‑scale HPC environments, troubleshoot issues, automate recurring tasks, and maintain reliability across compute, storage, networking, and research workloads.

With 2+ years of Linux experience, you will develop tooling, participate in on‑call rotations, and collaborate with engineering teams to keep production systems healthy in a fast‑paced research setting.

Qualifications

  • 2+ years of professional Linux systems experience.
  • Strong programming or scripting skills in Go, Python, C, or similar.
  • Genuine interest in hands-on operational engineering.

Responsibilities

  • Provide front-line operational support across 24/7 Linux HPC environments.
  • Troubleshoot compute, storage, networking, and interconnect issues.
  • Respond rapidly to infrastructure alerts and problem reports.
  • Develop automation for diagnostics, troubleshooting, and recurring tasks.
  • Monitor system performance, reliability, and infrastructure health.
  • Support global infrastructure projects and production environments.

Skills

Linux
Python
Go
C

Job description

Westbury Partners – Sydney NSW

Full time

Provide hands‑on operational support for large‑scale Linux HPC environments, troubleshooting complex infrastructure, automating recurring tasks, maintaining reliability, and supporting compute, storage, networking, and research workloads.

What You’ll Do:
  • Provide front‑line operational support across 24/7 Linux HPC environments.
  • Troubleshoot compute, storage, networking, and interconnect issues.
  • Respond rapidly to infrastructure alerts and problem reports.
  • Develop automation for diagnostics, troubleshooting, and recurring tasks.
  • Monitor system performance, reliability, and infrastructure health.
  • Support global infrastructure projects and production environments.
Your responsibilities will include:
  • Support HPC compute, storage, and high‑performance network infrastructure.
  • Troubleshoot parallel filesystems, batch schedulers, and RDMA fabrics.
  • Diagnose complex production issues and perform root cause analysis.
  • Develop tools for operational automation and system monitoring.
  • Manage the complete lifecycle of technical incidents and problems.
  • Participate in maintenance operations, including scheduled evening and weekend windows.
  • Collaborate with engineering teams and external technology vendors.
  • Maintain comprehensive systems and user documentation.
  • Participate in an operational on‑call rotation.
Why Join Us:
  • Work hands‑on with sophisticated, large‑scale HPC infrastructure.
  • Solve challenging and unpredictable operational problems every day.
  • Gain exposure to compute, storage, networking, and automation technologies.
  • Contribute to global infrastructure projects and production environments.
  • Work closely with technical teams supporting demanding research workloads.
  • Build expertise across a diverse and constantly evolving technology landscape.
About You:
  • 2+ years of professional Linux systems experience.
  • Strong programming or scripting skills in Go, Python, C, or similar.
  • Genuine interest in hands‑on operational engineering.
  • HPC experience is advantageous but not essential.
  • Strong troubleshooting and root cause analysis capabilities.
  • Excellent verbal and written communication skills.
  • Comfortable managing multiple projects and competing priorities.
  • Strong collaboration skills and a sense of urgency.
  • Willing to participate in evening, weekend, and on‑call support.
  • Reliable, adaptable, and comfortable working in a fast‑paced environment.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Linux Operations Engineer | 24/7 Infra & Automation
HPC Linux Operations Engineer | 24/7 Infra & Automation

Westbury Partners • Sydney

On-site
AUD 90,000 - 120,000
Linux HPC Operations Engineer - 24/7 High-Performance Infra
Linux HPC Operations Engineer - 24/7 High-Performance Infra

Westbury Partners • Sydney

On-site
AUD 120,000 - 170,000
Linux HPC Operations Engineer
Linux HPC Operations Engineer

Westbury Partners • Sydney

On-site
AUD 120,000 - 170,000
HPC Systems Engineer
HPC Systems Engineer

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
HPC Operations Engineer
HPC Operations Engineer

Jump Trading • Sydney

On-site
AUD 80,000 - 120,000
Senior HPC Systems Engineer - Linux, Slurm & Networking
Senior HPC Systems Engineer - Linux, Slurm & Networking

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
HPC Operations Engineer — 24/7 Linux & RDMA Systems
HPC Operations Engineer — 24/7 Linux & RDMA Systems

Jump Trading • Sydney

On-site
AUD 80,000 - 120,000
Lead HPC Ops Engineer for Global Low-Latency Trading
Lead HPC Ops Engineer for Global Low-Latency Trading

Mode Talent Pty Ltd • Sydney

On-site
AUD 360,000 - 440,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Experis • City of Melbourne

On-site
AUD 140,000 - 200,000
Weekly pay
Hybrid work arrangements
Cutting-edge technology access
HPC Operations Engineer
HPC Operations Engineer

Mode Talent Pty Ltd • Sydney

On-site
AUD 360,000 - 440,000