Systems Developer: HPC Scheduling & Platform Reliability

engineeringjobs.net, Inc.

New York (NY)

On-site

USD 110,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

engineeringjobs.net, Inc. is seeking engineers to design and improve workload scheduling, fleet management, and clustered file systems for large-scale compute environments.

You will investigate kernel and network performance, build metrics tooling, and collaborate with a global team supporting research workloads. The role requires strong programming ability in Python, Go, or Rust and a technical degree or equivalent experience, along with curiosity, organization, and time-management skills;

Qualifications

  • Candidates should have strong programming ability in Python, Go, Rust, or a comparable language, along with a computer science or related technical degree or comparable practical experience.
  • The role also calls for curiosity, communication, organization, and time-management skills; Linux, networking, HPC, storage, GPU infrastructure, or data-center hardware experience is desirable.

Responsibilities

  • Design and improve workload scheduling, fleet management, clustered file systems, software lifecycle tooling, and platform reliability for large-scale compute environments.
  • Investigate and tune kernel and network performance, build infrastructure metrics tools, and collaborate with a global team supporting research and trading workloads.

Skills

Python
Go
Rust
Linux Systems
Networking
Workload Scheduling
Fleet Management
Clustered File Systems
Kernel Performance
Network Performance Tuning
Metrics Collection
High-Performance Computing
Storage Infrastructure
GPU Infrastructure
Data-Center Hardware
Platform Reliability

Education

BS in Computer Science or related field

Job description

engineeringjobs.net, Inc. is seeking engineers to design and improve workload scheduling, fleet management, and clustered file systems for large-scale compute environments.

You will investigate kernel and network performance, build metrics tooling, and collaborate with a global team supporting research workloads. The role requires strong programming ability in Python, Go, or Rust and a technical degree or equivalent experience, along with curiosity, organization, and time-management skills;

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Systems Developer: HPC Scheduling & Clustered File Systems
Systems Developer: HPC Scheduling & Clustered File Systems

Referment • New York (NY)

On-site
USD 120,000 - 160,000
Systems Developer: HPC Workload & Fleet Management
Systems Developer: HPC Workload & Fleet Management

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 120,000 - 160,000
HPC Systems Engineer - Scheduling & Performance
HPC Systems Engineer - Scheduling & Performance

Referment • New York (NY)

On-site
USD 90,000 - 130,000
Quant Systems: Systems Developer (New York) (E00E524)
Quant Systems: Systems Developer (New York) (E00E524)

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 110,000 - 170,000
Quant Systems: Systems Developer (New York) (7B419E5)
Quant Systems: Systems Developer (New York) (7B419E5)

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Quant Systems: Systems Developer (New York) (4295F7B)
Quant Systems: Systems Developer (New York) (4295F7B)

Referment • New York (NY)

On-site
USD 90,000 - 130,000
Quant Systems: Systems Developer (New York) (2D66E00)
Quant Systems: Systems Developer (New York) (2D66E00)

Referment • New York (NY)

On-site
USD 120,000 - 160,000
Quant Systems: Systems Developer (New York) (B419E59)
Quant Systems: Systems Developer (New York) (B419E59)

Referment • New York (NY)

On-site
USD 120,000 - 180,000
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 241,500