HPC Systems Administrator

100 Eli Lilly and Company

South San Francisco (CA)

Hybrid

USD 141,000 - 231,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401(k)
Pension
Vacation benefits
Medical benefits

Job summary

Lilly is seeking an HPC Systems Administrator to build and operate scalable AI and high-performance computing platforms in the Silicon Valley hub. You will partner with AI scientists, engineers, and domain experts to enable efficient model training, inference, and experimentation across GPU, cloud, and on‑prem environments.

The role emphasizes automation, reliability, and cost efficiency while adapting to evolving AI/HPC technologies.

Qualifications

  • Bachelor’s degree in a technical field and 5+ years deploying or supporting large-scale HPC/GPU environments.
  • Experience operating AI/HPC infrastructure in cloud environments (AWS/Azure/GCP).
  • Strong scripting and automation capabilities across Linux-based systems.

Responsibilities

  • Build and operate scalable AI and HPC platforms for model training and inference.
  • Collaborate with AI scientists, engineers, and domain experts to enable pipelines across GPU, cloud, and on‑prem environments.
  • Enhance productivity and reliability through automation and standardization.
  • Maintain security, performance, and cost efficiency across environments.
  • Support distributed training workloads across multi-GPU, multi-node setups.
  • Troubleshoot complex infrastructure challenges and implement scalable solutions.

Skills

Linux systems administration
Automation
Scripting
Job scheduling platforms
Ansible
Kubernetes
Docker
Distributed computing
GPU infrastructure management
Cloud environments
Infra monitoring & observability
Problem solving

Education

Bachelor’s Computer Science, Electrical/Computer Engineering, Systems Engineering, or related field

Tools

Slurm/Grid Engine
Ansible
Kubernetes
Docker
AWS/Azure/GCP

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley! Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us!

About the Lilly and NVIDIA Partnership

Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.

What You’ll Be Doing

The HPC Systems Administrator will build and operate scalable AI and high-performance computing platforms that power advanced machine learning and scientific workloads. You will partner with AI scientists, engineers, and domain experts to enable efficient model training, inference, and experimentation across GPU, cloud, and on-premises environments. Through platform engineering and automation, you will improve productivity, performance, and access to advanced computing resources while ensuring the reliability, availability, and efficiency of Lilly's AI and HPC infrastructure.

How You’ll Succeed
  • Deliver highly available, secure, and performant AI and HPC platforms that meet the needs of research and engineering teams.
  • Drive operational excellence through automation, standardization, monitoring, and continuous improvement.
  • Balance infrastructure reliability, scalability, and cost efficiency across environments.
  • Enable efficient ML workflows through automation for orchestration, resource scheduling, data access, and reproducibility.
  • Collaborate effectively across scientific, engineering, and infrastructure teams to solve complex technical challenges.
  • Adapt quickly to evolving AI, GPU, and HPC technologies and translate new capabilities into business value.
What You Should Bring
  • Deep expertise in Linux systems administration, automation, and infrastructure management, with strong scripting skills in Python, Bash, and/or Ansible.
  • Experience building, administering, and optimizing large-scale HPC, GPU, or AI/ML computing environments, including job scheduling and resource management platforms such as Slurm or Grid Engine.
  • Proficiency with automation, configuration management, and container technologies, including Ansible, Kubernetes, Docker, and related tooling.
  • Solid understanding of distributed computing, high-performance networking, storage architectures, and cluster infrastructure.
  • Experience supporting large-scale distributed training and inference workloads across multi-GPU and multi-node environments.
  • Knowledge of GPU infrastructure, hardware lifecycle management, monitoring, and observability practices.
  • Demonstrated ability to solve complex infrastructure challenges, identify root causes, and implement scalable automated solutions.
  • Good communication and collaboration skills, with the ability to work effectively across researchers, engineers, and infrastructure teams.
  • Experience running NVIDIA GPU infrastructure and hardware lifecycle operations, including GPU monitoring, partitioning, diagnostics, and out-of-band server management using industry-standard tools and protocols.
  • Experience supporting regulated or critical environments is a plus.
  • Experience operating AI/HPC infrastructure in cloud environments such as AWS, Azure, or GCP.
Your Basic Qualifications

Bachelor’s Computer Science, Electrical/Computer Engineering, Systems Engineering, or a related technical field. 5 years’ experience deploying, administering, or supporting large-scale HPC, GPU, or distributed computing environments in an enterprise or research setting.

Location & Work Flexibility

This role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with three days onsite and two days working remotely each week, supporting both collaboration and work‑life balance.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include:

  • Africa
  • Middle East
  • Central Asia (AMECA)
  • Black Employees at Lilly (BE@Lilly)
  • Chinese Culture Network (CCN)
  • EnAble
  • Evolve
  • Lilly Indian Network (LIN)
  • Organization of Latinx at Lilly (OLA)
  • Pride (LGBTQ+ Allies)
  • Veterans Leadership Network (VLN)
  • Women’s Initiative for Leading at Lilly (WILL)

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is $141,000 - $231,000 Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance).

Benefits
  • Eligibility to participate in a company-sponsored 401(k)
  • Pension
  • Vacation benefits
  • Eligibility for medical, dental, vision and prescription drug benefits
  • Flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts)
  • Life insurance and death benefits
  • Certain time off and leave of absence benefits
  • Well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities)

Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly At Lilly we strive to ensure our employees are part of a team that cares about them and our shared purpose of making life better for those around the world. How do we do this? We continue to look for ways to include, innovate, accelerate and deliver while maintaining integrity, excellence and respect for people. We hope that you seek to join us on our journey as we create medicine and deliver improved outcomes for patients across the globe! #WeAreLilly

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 253,000
401(k)
Pension
Vacation benefits
+1
Data Engineer
Data Engineer

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 158,000 - 231,000
401(k)
Pension
Vacation benefits
+4
Data Engineer
Data Engineer

Eli Lilly and Company • San Francisco (CA)

Hybrid
USD 158,000 - 231,000
Company bonus
401(k) and pension
Health benefits
+2
AI PhD ML Engineering Intern
AI PhD ML Engineering Intern

100 Eli Lilly and Company • Indianapolis (IN)

On-site
USD 120,000 - 136,000
Free parking
Lilly LIFE fitness center
Bike garage
+1
Sr. AI Science Lead
Sr. AI Science Lead

Eli Lilly and Company • San Francisco (CA)

Hybrid
USD 260,000 - 381,000
Company bonus
Health benefits (medical, dental, vis.
401(k) and pension
+2
AI Scientist (Model Building & Training)
AI Scientist (Model Building & Training)

Initial Therapeutics, Inc. • San Francisco (CA)

Hybrid
USD 168,000 - 268,000
Company bonus
401(k) plan
Health, dental, vision benefits
+4
Postdoctoral Fellow - Agentic AI Solutions and SciML
Postdoctoral Fellow - Agentic AI Solutions and SciML

100 Eli Lilly and Company • United States

On-site
USD 58,000 - 123,000
401(k) plan
Pension
Vacation benefits
+1
Senior AI Science Lead
Senior AI Science Lead

Eli Lilly and Company • San Francisco (CA)

Hybrid
USD 260,000 - 381,000
401(k)
Pension
Medical benefits
+1
R-95809 Engineer - HPC Platform
R-95809 Engineer - HPC Platform

Eli Lilly and Company • Indianapolis (IN)

On-site
USD 64,000 - 185,000
Company bonus based on performance
Eligibility to participate in a company-sponsored 401(k)
Comprehensive health benefits including medical, dental, and vision
AI Scientist (Model Building & Training)
AI Scientist (Model Building & Training)

BioSpace • San Francisco (CA)

Hybrid
USD 168,000 - 268,000
401(k)
Pension
Medical benefits
+1