Senior HPC Systems Engineer – Research Computing Backbone

SLAC

Palo Alto (CA)

On-site

USD 168,000 - 200,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

PTO
Health benefits
Tuition assistance
403(b) plan
Mental health programs
Commuter benefits

Job summary

Stanford Research Computing seeks a talented systems engineer to strengthen the backbone of our multi-petabyte research computing environment. You will steward bastion and jump hosts, manage FLEXlm license services, and lead the central observability stack while supporting a co-location initiative and replication across sites.

You’ll also lead container hosts, backbone services, and deployment projects, ensuring security, resiliency, and proper documentation.

Qualifications

  • Bachelor's degree and ten years of relevant experience or equivalent.
  • Experience with complex, multi-system platforms and coordinating across vendors.
  • Ability to design and implement disaster recovery and business continuation plans.
  • Proficiency in multiple programming languages and operating systems.
  • Experience leading large deployment projects and cross-team collaboration.
  • Knowledge of security trends and best practices.

Responsibilities

  • Lead bastion hosts administration, hardening and high availability across data centers.
  • Operate and improve FLEXlm license servers and related services.
  • Oversee central observability services (Prometheus, Grafana, Splunk, XDMoD).
  • Support co-location design, production hand-off and remote replication.
  • Manage container hosts, backbone services, and Ansible/Git-driven config.
  • Establish and enforce secure build standards and access lifecycle.
  • Document core service requirements and recovery procedures.
  • Plan hardware lifecycle and capital needs with platform leads.
  • Engage with vendors to triage issues, manage RMAs and warranties.

Skills

Multi-system platforms
Disaster recovery
System deployment
Programming languages
Security best practices
Vendor coordination

Education

Bachelor's degree

Tools

Ansible
Git

Job description

Stanford Research Computing seeks a talented systems engineer to strengthen the backbone of our multi-petabyte research computing environment. You will steward bastion and jump hosts, manage FLEXlm license services, and lead the central observability stack while supporting a co-location initiative and replication across sites.

You’ll also lead container hosts, backbone services, and deployment projects, ensuring security, resiliency, and proper documentation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC & Observability Infrastructure Engineer
Senior HPC & Observability Infrastructure Engineer

Stanfordlivetickets • Palo Alto (CA)

On-site
USD 168,000 - 200,000
Health benefits
PTO 18+ days
Tuition assistance
+2
Lead Research Computing Systems Engineer
Lead Research Computing Systems Engineer

Stanford University • Palo Alto (CA)

On-site
USD 168,000 - 200,000
Health benefits
Flexible work options
Tuition assistance
+1
Senior Research Computing Systems Architect
Senior Research Computing Systems Architect

Stanford University • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Research Computing Systems Engineer
Research Computing Systems Engineer

Stanford University • Palo Alto (CA)

On-site
USD 140,000 - 190,000
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions, LLC. • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
paid time off
401k match
health care benefits
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions • Berkeley (CA)

Remote
USD 140,000 - 190,000
Paid time off
401k match
Health care benefits
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine • Berkeley (CA)

Remote
USD 140,000 - 190,000
Paid time off
401k match
Health care benefits
Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Guardant Health • Palo Alto (CA)

Hybrid
USD 173,000 - 238,000
Hybrid work model
Senior Backend Engineer: Distributed Systems & HPC
Senior Backend Engineer: Distributed Systems & HPC

Colossus Technologies Group • Watertown (MA)

On-site
USD 120,000 - 180,000
Hybrid Cloud & Data Center Engineer
Hybrid Cloud & Data Center Engineer

SLAC • Palo Alto (CA)

Hybrid
USD 137,000 - 157,000
PTO 18+ days
Health benefits
Flexible work options
+4