Linux HPC Operations Engineer

Westbury Partners

Sydney

On-site

AUD 120,000 - 170,000

Full time

42 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Westbury Partners is seeking a hands-on Linux engineer to operate and optimize large-scale HPC environments in Sydney. You will resolve production issues, automate tasks, and maintain compute, storage, and networking systems around the clock.

You will collaborate with researchers and engineers, develop tooling, participate in on-call rotations, and contribute to global infrastructure projects, while following security policies.

Qualifications

  • Minimum two years of professional Linux experience.
  • Strong scripting in Python, Go, or C.
  • Ability to diagnose complex HPC infrastructure issues.
  • Willingness to participate in on-call rotation and maintenance windows.

Responsibilities

  • Provide front-line operational support for 24/7 Linux HPC compute, storage, and network infrastructure.
  • Troubleshoot complex issues across RDMA fabrics, parallel filesystems, batch schedulers, FUSE filesystems, hardware, and software.
  • Manage incidents and problem reports through their entire lifecycle, escalating when required.
  • Respond quickly to infrastructure alerts.
  • Participate in coordinated maintenance activities, including scheduled evening and weekend maintenance windows.
  • Support global infrastructure projects across a broad range of technologies.
  • Develop code and tooling to diagnose, troubleshoot, triage, and resolve difficult operational problems.
  • Automate frequently performed operational tasks.
  • Contribute to testing infrastructure and codebases across multiple programming languages.
  • Develop and maintain performance and fault‑monitoring systems.
  • Create and improve technical and user documentation.
  • Work with external technology vendors and support vendor relationships.
  • Participate in an on‑call rotation and provide operational support as a core responsibility.
  • Follow cybersecurity and IT policies when managing infrastructure and production systems.

Skills

Linux
Scripting
Python
Go
C
HPC
Networking
RDMA
Storage
FUSE

Job description

Support and optimise large-scale Linux HPC environments, resolving complex infrastructure issues, automating operations, and maintaining reliable compute, storage, and networking systems around the clock.

What You’ll Do:
  • Join a highly technical infrastructure team responsible for keeping large-scale Linux HPC environments running reliably and efficiently. You’ll provide hands‑on operational support across compute, storage, networking, and interconnected infrastructure while responding to challenging and unpredictable production issues.
  • You’ll work closely with researchers, engineers, vendors, and infrastructure teams to investigate problems, implement solutions, automate repetitive tasks, and continuously improve the reliability and performance of the environment.
Your responsibilities will include:
  • Provide front-line operational support for 24/7 Linux HPC compute, storage, and network infrastructure.
  • Troubleshoot complex issues across RDMA fabrics, parallel filesystems, batch schedulers, FUSE filesystems, hardware, and software.
  • Manage incidents and problem reports through their entire lifecycle, escalating when required.
  • Respond quickly and effectively to infrastructure alerts.
  • Participate in coordinated maintenance activities, including scheduled evening and weekend maintenance windows.
  • Support global infrastructure projects across a broad range of technologies.
  • Develop code and tooling to diagnose, troubleshoot, triage, and resolve difficult operational problems.
  • Automate frequently performed operational tasks.
  • Contribute to testing infrastructure and codebases across multiple programming languages.
  • Develop and maintain performance and fault‑monitoring systems.
  • Create and improve technical and user documentation.
  • Work with external technology vendors and support vendor relationships.
  • Participate in an on‑call rotation and provide operational support as a core responsibility.
  • Follow cybersecurity and IT policies when managing infrastructure and production systems.
Why Join Us:
  • This is an opportunity to work directly with large-scale, high‑performance computing infrastructure where every day brings new technical challenges.
  • You’ll gain hands‑on exposure to Linux, HPC, distributed storage, high-speed networking, automation, monitoring, hardware, and production operations while working alongside highly technical infrastructure and research teams.
  • The role is ideal for engineers who genuinely enjoy operational work and want to develop their expertise by solving difficult problems in a fast-moving environment. You’ll also have opportunities to contribute to global projects, work with leading technology vendors, and build automation that makes complex infrastructure easier to operate.
About You:
  • You’re a hands‑on Linux engineer who enjoys operational work and thrives when solving unpredictable technical problems.
  • You have at least two years of professional Linux experience and strong programming or scripting skills in a language such as Python, Go, or C. HPC experience is valuable but not essential — what matters is your ability to learn quickly and tackle unfamiliar technologies.
  • You’re comfortable performing root cause analysis, managing multiple workstreams, and communicating clearly with both technical colleagues and external vendors.
  • You bring a strong sense of urgency, excellent collaboration skills, reliable availability, and the flexibility to support scheduled maintenance and on‑call responsibilities when required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hpc Operations Engineer - 24/7 Linux & Rdma Systems
Hpc Operations Engineer - 24/7 Linux & Rdma Systems

Jump Trading • New South Wales

On-site
AUD 120,000 - 180,000
HPC Operations Engineer
HPC Operations Engineer

Jump Trading • Sydney

On-site
AUD 80,000 - 120,000
HPC Systems Engineer
HPC Systems Engineer

P2P • Sydney

On-site
AUD 140,000 - 190,000
HPC Operations Engineer — 24/7 Linux & RDMA Systems
HPC Operations Engineer — 24/7 Linux & RDMA Systems

Jump Trading • Sydney

On-site
AUD 80,000 - 120,000
HPC Systems Engineer
HPC Systems Engineer

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
Senior HPC Systems Engineer - Linux & Scale
Senior HPC Systems Engineer - Linux & Scale

P2P • Sydney

On-site
AUD 140,000 - 190,000
24/7 Linux HPC Operations Engineer (RDMA & Storage)
24/7 Linux HPC Operations Engineer (RDMA & Storage)

Jump Trading • Sydney

On-site
AUD 120,000 - 180,000
Senior HPC Systems Engineer - Linux, Slurm & Networking
Senior HPC Systems Engineer - Linux, Slurm & Networking

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
Linux Engineer
Linux Engineer

Cipher Recruitment • City of Melbourne

Hybrid
AUD 110,000 - 170,000
System Engineer
System Engineer

Precision Sourcing • Sydney

On-site
AUD 120,000 - 180,000