Senior Systems Administrator

Clear Destination Inc.

Mountain View (CA)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ASRC Federal is seeking a Staff HPC Engineer to support InuTeq LLC out of NASA Ames in Mountain View, CA. The role centers on Linux system administration and HPC operations, maintaining clusters with 2000+ nodes and Lustre storage, and contributing to the PBS scheduler and InfiniBand networking.

You will design and implement automation tools, monitor performance, patch systems, deploy new hardware, and help users run jobs efficiently.

Qualifications

  • Bachelor's degree in computer science or related field.
  • 5+ years of Linux administration experience.
  • Strong scripting skills: Python, Perl or Bash.
  • Experience interacting with customers to understand needs and relay feedback.
  • Familiarity with HPC components: Lustre, PBS schedulers, InfiniBand networking.

Responsibilities

  • Maintain HPC clusters with 2000+ nodes and large storage systems.
  • Develop and deploy automation scripts for administration and monitoring.
  • Patch OS, deploy new systems and perform hardware upgrades.
  • Troubleshoot hardware, software, and network issues affecting jobs.
  • Document processes and provide user support and after-hours assistance.

Skills

Linux administration
Scripting (Python/Perl/Bash)
Customer interaction
Performance tuning
OpenMP/MPI knowledge

Education

Bachelor's degree in computer science or related field

Tools

Lustre
PBS
InfiniBand

Job description

ASRC Federal is searching for a Staff HPC Engineer to support InuTeq LLC out of NASA AMES, CA

ASRC Federal, InuTeq proudly supports NASA's High-Performance Computing Services program with our site in Mountain View, CA at the Ames Research Center. Make a DIFFERENCE on a program that supports 4 On-site Supercomputers totaling ~9,000 nodes and 25+ combined petaflops. This program provides High-Performance Computing services throughout the HPC lifecycle for computational requirements, architecture, acquisition, to our NASA customer. Our employees embrace innovation and are committed to a culture of continuous, standards-driven process improvement, and assimilation of industry best practices.

Summary

The successful candidate will be an active supporting member of the ASRC Federal team reporting directly to the Manager of the HPC Computer Systems and Storage (CSS) group. An individual at this skill level should have demonstrated Linux administration experience as well as working knowledge of HPC operations while contributing to the support of users of HPC resources on the various issues they might have getting jobs to run efficiently. This individual will be expected to participate in all aspects of HPC System Administration that includes such activities as installing, maintaining, and upgrading HPC systems. The individual, along with the entire HPC team, will be engaged in the day-to-day operations and support of the HPC resources. Activities may include system patching, OS upgrades, deploying new systems, writing scripts, and troubleshooting system issues on the HPC system. The ability to interact with users to determine symptoms and then reproduce their issues to isolate the causes is critical skills for this work. There may also be activities supporting testing, benchmarking, tool scripting, and analyzing trouble tickets to find patterns indicating system or user education issues.

Duties and Responsibilities
  • Maintains HPC clusters with over 2000+ nodes with InfiniBand, 100+ petabytes of data storage in production.
  • Designs and develops scripts for system administration, monitoring and usage reporting.
  • Modify existing software to correct errors and/or improve performance
  • Designs and develops scripts for system regression test and performance (file systems (Lustre), scheduler (PBS), interconnect (HDR/NDR, Slingshot, ), high availability, etc.).
  • Troubleshoots, isolates and resolves application, system and other technical problems (hardware, software, and network).
  • Understands research use cases, researches and deploys new technologies, defining cost, performance and other trade-offs.
  • Manages and maintains tools for configuration management (HPCM, Ansible & GIT), resource management, scheduling and all necessary aspects of HPC in accordance with best practices.
  • Researches, deploys and manages networking and security infrastructure, including development of policies and procedures.
  • Assists in developing and writing proposals and publications.
  • Creates and provides clear documentation.
  • After hours/weekend support as required
  • Moderate Supercomputing System Administration that contributes to:
    • Day-to-day operations of the Linux HPC clusters and storage systems
    • Proactive monitoring, analyzing, and correct system issues
    • Development of scripts to automate repetitive tasks or tools to enhance support of the HPC systems
    • System performance analysis and tuning
    • Building, installing, and supporting user-requested software
    • Supporting evaluation and assessment of new HPC technology
    • Resolving user report issues and managing support tickets requests in Remedy
Requirements
  • Bachelor’s degree in computer science or related field
  • Strong computer science background with in-depth systems-level knowledge in operating systems and networking
  • A minimum of 5 years’ experience of administration of Linux systems
  • Strong ability to analyze, debug and maintain the integrity of an existing code base
  • Demonstrated equivalence of 5 years of Linux/UNIX user support experienceand hands-on experience with administration of Linux systems
  • Superior scripting skills and excellent attention to detail; proficiency in at least Python, Perl, or Bash
  • Strong ability to interact with customers to understand needs, elicit requirements, and get feedback on prototype solutions
  • Excellent communication and people skills; excellent time management and organizational skills
  • Experience with system configuration management tools e.g. , ansible, puppet, chef
  • Experience with revision control software e.g. CVS, SVN, Git
  • Proficiency at technical writing
Preferred Skills (Requesting Manager Defines)
  • Proficiency with analysis and problem-solving skills for debugging and optimization of applications
  • Familiarity/proficiency with OpenMP and Message Passing Interface (MPI) programming
  • Experience with Lustre, and InfiniBand
  • Experience with cloud technologies (AWS, Azure, GCP), OpenStack or Kubernetes is a plus
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Administrator
Senior Systems Administrator

ASRC Federal Holding Company • Mountain View (CA)

On-site
USD 150,000 - 190,000
Senior Systems Administrator
Senior Systems Administrator

ASRC Federal • Mountain View (CA)

Hybrid
USD 140,000 - 190,000
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Clear Destination Inc. • Mountain View (CA)

Hybrid
USD 150,000 - 210,000
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions, LLC. • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
paid time off
401k match
health care benefits
Senior HPC Systems Engineer for Large-Scale Clusters
Senior HPC Systems Engineer for Large-Scale Clusters

ASRC Federal Holding Company • Mountain View (CA)

On-site
USD 150,000 - 190,000
Senior Systems Administrator (HPC Lab Integration & Development)
Senior Systems Administrator (HPC Lab Integration & Development)

Fuse Engineering • Maryland

On-site
USD 70,000 - 90,000
Linux Systems Administrator – Top Secret HPC
Linux Systems Administrator – Top Secret HPC

GIGATEC Engineering • Maryland

On-site
Sr. HPC Systems Engineer (High Performance Computing)
Sr. HPC Systems Engineer (High Performance Computing)

United States Digital Space LLC • El Segundo (CA)

On-site
USD 165,000 - 265,000
Medical, Vision & Dental
401(k) plan
Disability insurance
+4
HPC System Administrator (Top Secret)
HPC System Administrator (Top Secret)

RedLine Performance Solutions, LLC • Dayton (OH)

On-site
USD 110,000 - 160,000
Full benefits package
401(k) match
Paid time off
HPC System Administrator (Top Secret)
HPC System Administrator (Top Secret)

RedLine Performance Solutions, LLC. • Dayton (OH)

Hybrid
USD 90,000 - 130,000
PTO
401k match
Health insurance