HPC Operations Engineer: High-Performance Cluster Lead

Evergreen Statistical Trading

Bellevue (KY)

On-site

USD 180,000 - 220,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Evergreen is a proprietary trading firm based in Bellevue, Washington. As an HPC Operations Engineer, you will own the day-to-day operation of our research clusters across multiple datacenters and the scheduled workloads our research and trading rely on.

You will work on a modern stack without legacy infrastructure and take on major projects across software and hardware as our HPC footprint expands, collaborating with a hands-on team.

Qualifications

  • Linux systems administration on Red Hat-family distros (RHEL, Rocky, AlmaLinux)
  • Python for automation and tooling
  • Git and versioned configuration management experience
  • Prometheus/Grafana monitoring and alerting experience
  • Slurm batch scheduler experience including queue config and resource management
  • Cluster provisioning tools such as Warewulf or xCAT
  • ZFS and parallel filesystems like GPFS or Lustre
  • InfiniBand networking experience
  • Ability to own critical processes and resolve issues outside business hours
  • Hands-on hardware coordination in datacenters
  • Ability to diagnose complex production problems under time pressure
  • Growth-oriented and collaborative mindset
  • Passion for high-performance infrastructure with minimal drama

Responsibilities

  • Own day-to-day operation of research clusters across multiple datacenters and scheduled workloads
  • Work on a modern stack unencumbered by legacy infrastructure and tackle major projects across software and hardware as HPC footprint grows
  • Collaborate with teams to scale and improve HPC performance and reliability

Skills

Linux systems administration
Python scripting
Ownership of critical processes
Team collaboration

Tools

Git
Configuration management
Prometheus
Grafana
Slurm
Warewulf
xCAT
ZFS
GPFS
Lustre
InfiniBand

Job description

Evergreen is a proprietary trading firm based in Bellevue, Washington. As an HPC Operations Engineer, you will own the day-to-day operation of our research clusters across multiple datacenters and the scheduled workloads our research and trading rely on.

You will work on a modern stack without legacy infrastructure and take on major projects across software and hardware as our HPC footprint expands, collaborating with a hands-on team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Operations Engineer – High-Performance Infra & Automation
HPC Operations Engineer – High-Performance Infra & Automation

Evergreen Statistical Trading • Bellevue (WA)

On-site
USD 170,000 - 230,000
Signing bonus
Performance bonus
Company-paid medical benefits
HPC Operations Engineer
HPC Operations Engineer

Evergreen Statistical Trading • Bellevue (WA)

On-site
USD 170,000 - 230,000
Signing bonus
Performance bonus
Company-paid medical benefits
HPC Operations Engineer
HPC Operations Engineer

Evergreen Statistical Trading • Bellevue (KY)

On-site
USD 180,000 - 220,000
HPC Systems Engineer — Research Compute & Automation
HPC Systems Engineer — Research Compute & Automation

University of Washington • Seattle (WA)

Hybrid
USD 97,000 - 158,000
GPU HPC Cluster Ops Engineer — Equity & Growth
GPU HPC Cluster Ops Engineer — Equity & Growth

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Life Insurance
Flexible Spending Account
+8
HPC Operations Engineer - Optimize Compute Systems (Equity)
HPC Operations Engineer - Optimize Compute Systems (Equity)

NVIDIA • United States

On-site
USD 124,000 - 242,000
Equity
Comprehensive benefits
HPC & Compute Engineering Lead
HPC & Compute Engineering Lead

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 250,000
Remote HPC Systems Engineer — Research Compute
Remote HPC Systems Engineer — Research Compute

Socket.dev • Seattle (WA)

On-site
USD 97,000 - 158,000
Hybrid HPC Systems Engineer
Hybrid HPC Systems Engineer

FHLB Des Moines • Seattle (WA)

On-site
USD 97,000 - 158,000
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2