Research Computing Systems Engineer

Stanford University

Palo Alto (CA)

On-site

USD 168,000 - 200,000

Full time

8 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
Flexible work options
Tuition assistance
PTO days

Job summary

Stanford Research Computing seeks a talented systems engineer to steward bastion and backbone services across data centers, while contributing to co-location and replication initiatives. You will drive automated provisioning, security hardening, and observability, collaborating with researchers and admins across Stanford's schools.

The role emphasizes lifecycle management of infrastructure, planning for hardware refresh, and ensuring reliable, compliant operation of a multi-petabyte research

Qualifications

  • Extensive experience with complex, multi-system platforms and vendors.
  • Notable experience coordinating multi-system environments in independent facilities.
  • Experience developing business continuation and disaster recovery plans.
  • Ability to lead large deployment projects and manage cross-team communication.

Responsibilities

  • Lead design, development, installation and maintenance of operating systems and applications.
  • Anticipate risks and de-escalate issues to prevent disruptions.
  • Develop and enforce system security measures and policies.
  • Plan capacity, configure services, and optimize interdependencies.
  • Manage vendor relationships and oversee hardware/software maintenance agreements.

Skills

Multi-system platforms
Security practices
Disaster recovery

Education

Bachelor's degree

Tools

Ansible

Job description

Stanford Research Computing is looking for a talented systems engineer to join our team of collaborative and innovative professionals helping Stanford's faculty and students use advanced computing and data tools to explore new frontiers in knowledge and solve some of humanity's most urgent problems. Our staff work directly with some of the world's top researchers in a broad range of disciplines, across all of Stanford's seven schools — while also supporting and learning from each other in cross-project endeavors. We maintain and steadily improve an advanced research computing facility, and we support a variety of environments for Stanford research. In Stanford Research Computing, you'll have a rare opportunity to contribute to discoveries and inventions that have global reach and positive impact, and to share in the curiosity and commitment of the scholars and scientists who lead these projects.

About the Role

Stanford Research Computing operates a portfolio of large-scale HPC clusters and petabyte-scale storage systems serving thousands of researchers. Every one of those platforms depends on a layer of shared infrastructure that researchers never see: the bastion hosts our administrators pass through to reach management networks, the network license servers that let research software start, the container hosts carrying our internal services, and the observability stack that tells us something is wrong before a researcher has to.

In this role you will be the primary steward of that layer, serving as lead for the "backbone" services the rest of our environment depends on. You will also contribute to a co-location initiative, standing up management infrastructure and geographically distinct replication in a remote data center. The work spans a wide range of technologies, with substantial engineering ahead in hardware lifecycle, redundancy, and configuration management.

Why Stanford? You won't be maintaining back-office IT. You will be building the connective tissue that a multi-petabyte, multi-platform research computing environment runs on — the systems that decide whether a downtime is a scheduled inconvenience or a lost quarter of someone's research.

Responsibilities

Bastion hosts: Lead the administration, hardening, and modernization of bastion and jump hosts across multiple data centers, and design the next generation with redundancy and high availability. Support partner unit administrators moving from VDI to bastion-based access.

Network license services: Operate and improve the primary and secondary FLEXlm license servers behind commercial research and statistical software. Right-size firewall policy, retire orphaned license hosts, and bring runbooks and user documentation current.

Central observability service: Lead the central observability service and the servers behind it — Prometheus, Grafana, Splunk, XDMoD, and facility telemetry. Consolidate legacy monitoring and build alerting that reaches the right person with enough context to act.

Co-location project: Support a funded co-location initiative from design through production hand-off: network architecture, management infrastructure and a bastion host in the remote facility, and a geographically distinct replication target for independent backups.

Container hosts and backbone services: Lead the container hosts and the internal services they carry, along with cluster head and service nodes, bare-metal provisioning, and the VMs behind other Research Computing backbone services. Bring these under Ansible and Git-based configuration management.

Security and compliance: Establish and maintain secure build standards across the systems you own, in line with university, regulatory, and contractual requirements. Deploy endpoint and logging agents, remediate scan findings, and manage credential and access lifecycle.

Resilience and documentation: Reduce single points of failure, and document core service requirements and details so that any member of the team can operate, restore, or rebuild a service.

Planning and hardware lifecycle: Meet regularly with platform leads and service owners to understand upcoming needs, own hardware lifecycle and capital planning for the infrastructure you lead, and help develop business cases for replacement and expansion.

Vendor engagement: Liaise with hardware and software vendors and support partners to triage and resolve issues, manage RMAs, and keep systems under appropriate warranty and support coverage.

Responsibilities

Core Duties:

Lead the design, development, installation and maintenance ofoperating systems, utilities, and applications software on computingsystems.

Anticipate risks, de-escalate issues, and prevent emergencies tolimit disruptions to system operations and protect the integrity of userdata and systems.

Safeguard the university’s data and system assets – formulatesystem security strategies and develop viable policies and proceduresthat will enable the design and implementation of system securitymeasures at the university.

Establish and enforce systems policies and procedures andvalidate that university software/hardware standards are aligned withexternal best practice.

Partner with other information technology specialty areas toconfirm information technology strategies, devise and deploy plans toensure information technology objectives are met, and advise ontechnical feasibility of information technology initiatives,particularly regarding system compatibility within the university’scurrent, or proposed technical or structural framework(s).

Review and conduct capacity planning for system configuration,software services, network services, load distribution, and serviceinterrelationships among computer systems.

Act as technical expert or lead for university-wide computersystem administration. May manage system administration staff.

Provide project management for large and complex university-widecomputing projects.

Manage vendor relationships and negotiate cost effective hardwareand software maintenance agreements with vendors.

Minimum Education and Experience:

Bachelor's degree and ten years of relevant experience, or acombination of education and relevant experience.

Knowledge, Skills and Abilities:

Extensive experience with complex, multi-system platforms andvendors.

Notable experience coordinating multi-system and computingenvironments in independent computing facilities.

Extensive experience developing/implementing a businesscontinuation and disaster recovery plan.

Exceptional ability to develop appropriate plans to meetcomputing needs.

Expert ability to program in multiple programming languages inmultiple operating systems.

Superior ability to lead and work on large/complex systemdeployment projects in a team environment.

Expert knowledge of security trends and best practice.

  • Pay Range $167,539-$200,383
  • Pay Frequency Annually
  • Fixed Term N
  • Stanford Schools and Units Office of Vice President for Business Affairs and Chief Financial Officer
  • Full time or Part time Full Time
  • Regular or Temporary Regular
  • Job Category Information Technology Srvcs
  • Posting Date 10/08/2026, 08:52 PM
  • Grade L
  • Job Identification 201341
Get job alerts

Be the first to know about job openings and updates.

Benefits that support you, and your future

18+ PTO days per year

Health benefits start when you start

Flexible work options

$6000+ annual tuition and training assistance

403(b) plan

Mental health and wellness programs

Free and discounted commuter transportation

Applies to regular full-time positions.

Notice to Applicants The job duties listed are typical examples of work performed by positions in this job classifications and are not designed to contain or be interpreted as a comprehensive inventory of all duties, tasks and responsibilities. Specific duties and responsibilities may vary depending on department or program needs without changing the general nature and scope of the job or level of responsibility. Employees may also perform other duties as assigned. Consistent with its obligations under the law, the University will provide reasonable accommodation to any employee with a disability who requires accommodation to perform the essential functions of their job. Stanford is an equal employment opportunity and affirmative action employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by law. Stanford University provides job pay ranges representing its good faith estimate of what the university reasonably expects to pay for a particular job. The specific pay offered to a selected candidate will be determined based on a wide range of factors that are unique to each candidate, including but not limited to geographic work location, relevant knowledge, skills and abilities, relevant education, years of relevant experience, depth and breadth of relevant experience, and performance; further including but not limited to other business and organization needs such as the scope and responsibilities of the position, the minimum qualifications, departmental budget availability, and market and internal equity across the university, school/ unit, department, as well as job reporting relationships.

Why work at Stanford

World-changing work in a culture of continuous learning

An inspiring community that values collaboration

High-quality health and wellness benefits that put you and your family first

Access to world-class education and career development, with tuition reimbursement, training assistance, and more

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Hosting Services Operations Engineer
Hosting Services Operations Engineer

SLAC • Palo Alto (CA)

Hybrid
USD 137,000 - 157,000
PTO 18+ days
Health benefits
Flexible work options
+4
Hosting Services Operations Engineer
Hosting Services Operations Engineer

Stanfordlivetickets • Palo Alto (CA)

Hybrid
USD 137,000 - 157,000
Health benefits
Flexible work options
Tuition assistance
+2
Hosting Services Operations Engineer
Hosting Services Operations Engineer

Stanford University • Palo Alto (CA)

Hybrid
USD 137,000 - 157,000
PTO 18+ days per year
Tuition assistance
Health benefits
+2
Program Manager
Program Manager

Stanford University • Palo Alto (CA)

Hybrid
USD 65,000 - 73,000
Health benefits
Flexible work options
Tuition assistance
+3
Program Manager
Program Manager

SLAC • Palo Alto (CA)

Hybrid
USD 135,005,000 - 151,536,000
18+ PTO days per year
Health benefits start when you start
Flexible work options
+3
Program Manager
Program Manager

Stanfordlivetickets • Palo Alto (CA)

Hybrid
USD 65,000 - 73,000
PTO 18+ days
Health benefits
Flexible work options
+4
AI Engineer
AI Engineer

Stanfordlivetickets • Redwood City (CA)

On-site
USD 170,000 - 195,000
Health benefits
Tuition assistance
Flexible work options
+4
Administrative Services Administrator 1
Administrative Services Administrator 1

Stanford University • Palo Alto (CA)

Hybrid
USD 90,000 - 169,000
PTO 18+ days/year
Health benefits
Tuition assistance $6,000+ per year
+4
Administrative Associate 3
Administrative Associate 3

Stanfordlivetickets • Palo Alto (CA)

On-site
USD 45,000 - 84,000
Health benefits start when you start
Tuition assistance
PTO days
Research Administrator 2 (Remote Eligible)
Research Administrator 2 (Remote Eligible)

SLAC • Palo Alto (CA)

Hybrid
USD 60,000 - 117,000
Health benefits
Tuition assistance
Flexible work options
+1