Senior Engineering Manager, AI Infrastructure

Allen Institute for Artificial Intelligence

Seattle (WA)

On-site

USD 146,880 - 220,320

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision
401k plan
Commuting stipend
Wellbeing stipend
Paid time off
Bonuses

Job summary

Ai2, the Allen Institute for Artificial Intelligence, is seeking a Senior Manager, AI Infrastructure in Seattle to lead day‑to‑day HPC operations, focusing on on‑prem GPU clusters and a hybrid cloud orchestration layer. You will own reliability, performance and utilization while delivering the roadmap with your PM counterpart.

The role demands deep expertise in Linux systems, distributed architectures, and orchestration tools, plus leadership of a growing engineering team.

Qualifications

  • 12+ years in infrastructure, systems engineering, or HPC (or 8+ with advanced degree).
  • 2+ years supervising a small engineering team (5+).
  • Experience operating large-scale NVIDIA GPU clusters and high-performance networking.
  • Strong orchestration background in Kubernetes, Slurm, or similar in hybrid-cloud configs.
  • Experience with distributed file systems and cloud storage integration.

Responsibilities

  • Manage availability, performance and health of dense on‑prem GPU clusters.
  • Operate and improve Beaker orchestration platform for efficient resource use.
  • Oversee storage environment for high-throughput training and durability.
  • Manage GPU compute allocation and budget, decide cloud bursting vs on‑prem capacity.
  • Provide technical bridge to research teams to accelerate work.
  • Lead and grow a team of systems engineers, SREs, and developers.

Skills

Linux kernel
Container runtimes
Distributed systems
InfiniBand networking
NCCL optimizations
Kubernetes
Slurm
Hybrid-cloud orchestration
Go or Python

Education

Bachelor's degree in related field

Tools

Beaker
NVIDIA GPU clusters management

Job description

Persons in these roles are expected to work from our offices in Seattle. On-site requirements vary based on position and team. If you have questions about on-site work arrangements for this role, please ask your recruiter.
Our base salary range is $146,880 - $220,320, and in addition we have generous bonus plans to provide a competitive compensation package.

We are seeking a Senior Manager, AI Infrastructure to run the day-to-day operation of the systems that power our research. Reporting to the VP of Engineering, you will own the execution and reliability of our high-performance computing (HPC) environment which includes on-prem GPU clusters and the software orchestration layer that schedules workloads across a hybrid cloud environment. This is a hands‑on operational leadership role: your mandate is to keep the platform fast, reliable, and well-utilized, and to deliver against the roadmap set with your PM counterpart.

Our ideal candidate is a:

  • Systems Expert: You have a deep, hands‑on understanding of the Linux kernel, container runtimes, and distributed systems. You understand the performance implications of InfiniBand topologies and NCCL optimizations.
  • Execution‑Focused Leader: You plan and deliver against near‑term operational goals, keep reliability and researcher velocity high, and turn priorities set with leadership into shipped, dependable systems.
  • Pragmatic Operator: You are comfortable making trade‑offs between technical elegance and operational necessity. You triage and mitigate immediate risks, and know when to handle something yourself versus escalated.

Ai2 is a non‑profit research institute at the forefront of open‑source AI development. Unlike industry peers, our goal is to share our findings, data, code, and models with the global scientific community.

Why Ai2:
  • Open Science: Your work directly enables the release of open models like OLMo, providing the broader research community with tools they can't get elsewhere.
  • Mission‑Driven: We prioritize scientific impact over profit margins. This allows us to focus on building the "right" infrastructure for long‑term research goals.
  • Complexity at Scale: You will manage some of the most dense and high‑performance compute environments currently in operation.

Your Next Challenge:

  • Cluster Operations: Manage the availability, performance, and health of our dense on‑prem GPU clusters. Coordinate with hardware vendors and internal teams to keep physical infrastructure meeting the demands of frontier model training.
  • Orchestration & Scheduling: Operate and improve Beaker, our internal orchestration platform by optimizing resource allocation and driving high utilization across on‑prem assets and elastic cloud resources (AWS/GCP).
  • Storage Operations: Execute and continuously improve our storage environment, balancing high‑throughput performance for active training against cost‑effective durability for petascale research data. Contribute to the longer‑term storage roadmap.
  • Resource Management: Manage GPU compute allocation against budget. Track utilization, surface the data, and recommend when to burst to the cloud versus investing in on‑prem capacity, escalating larger trade‑offs as needed.
  • User Support & Velocity: Serve as the technical bridge to our research teams. Ensure infrastructure is an accelerator, not a bottleneck, for a diverse set of research objectives.
  • Team Leadership: Manage and grow a team of systems engineers, SREs, and software developers. Set the bar for operational rigor, engineering quality, and a collaborative culture, and keep the team unblocked and delivering.

What You’ll Need:

  • Experience: 12+ years in infrastructure, systems engineering, or HPC (or an advanced degree with 8+ years), including 2+ years supervising a small engineering team (5+).
  • Bachelor's degree in a related field: a relevant advanced degree may substitute for equivalent years of technical work experience.
  • GPU/HPC Stack: Direct experience operating large‑scale NVIDIA GPU clusters and high‑performance networking (InfiniBand/RoCE).
  • Orchestration: Strong background in Kubernetes, Slurm, or similar orchestration frameworks, particularly in hybrid‑cloud configurations.
  • Storage: Hands‑on experience with distributed filesystems (e.g., WEKA, Ceph, Lustre) and cloud storage integration at scale.
  • Software Development: Proficient in designing and managing SDLC processes including sprint planning and technical design reviews. Proficient in Go or Python.

Physical Demands and Work Environment:

The physical demands described here are representative of those that must be met by a team member to successfully perform the essential functions of this position. Reasonable accommodations may be made to enable individuals with disabilities to perform the functions.

  • Must be able to remain in a stationary position for long periods of time.
  • The ability to communicate information and ideas so others will understand. Must be able to exchange accurate information in these situations.
  • The ability to observe details at close range.
  • Can work under deadlines.

A Little More About Ai2:

Ai2 is a Seattle based non‑profit AI research institute founded in 2014 by the late Paul Allen. Our mission is building breakthrough AI to solve the world’s biggest problems. We develop foundational AI research and innovation to deliver real‑world impact through large‑scale open models, data, robotics, conservation, and beyond.

Ai2 is proud to be an Equal Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. You may view the related Know Your Rights compliance poster and the Pay Transparency Nondiscrimination Provision by any appropriate means.

This employer participates in E‑Verify and will provide the federal government with your Form I‑9 information to confirm that you are authorized to work in the U.S. If E‑Verify cannot confirm that you are authorized to work, this employer is required to give you written instructions and an opportunity to contact the Department of Homeland Security (DHS) or Social Security Administration (SSA) so you can begin to resolve the issue before the employer can take any action against you, including terminating your employment. Employers can only use E‑Verify once you have accepted a job offer and completed the Form I‑9.

We are committed to providing reasonable accommodations to employees and applicants with disabilities to the full extent required by the Americans with Disabilities Act (ADA). If you feel you need a reasonable accommodation pursuant to the ADA, you are encouraged to contact us at recruiting@allenai.org.

Benefits:

  • Team members and their families are covered by medical, dental, vision, and an employee assistance program.
  • Team members are able to enroll in our health savings account plan, our healthcare reimbursement arrangement plan, and our health care and dependent care flexible spending account plans.
  • Team members are able to enroll in our company’s 401k plan.
  • Team members will receive $125 per month to assist with commuting or internet expenses and will also receive $200 per month for fitness and wellbeing expenses.
  • Team members will also receive up to ten sick days per year, up to seven personal days per year, up to 20 vacation days per year and twelve paid holidays throughout the calendar year.
  • Team members will be able to receive annual bonuses and can participate in the long-term incentive plan.

Note: This job description in no way states or implies that these are the only duties to be performed by the team members(s) of this position. Team members will be required to follow any other job‑related instructions and to perform any other job‑related duties requested by any person authorized to give instructions or assignments. All duties and responsibilities are essential functions and requirements and are subject to possible modification to reasonably accommodate individuals with disabilities. To perform this job successfully, the team member(s) will possess the skills, aptitudes, and abilities to perform each duty proficiently. Some requirements may exclude individuals who pose a direct threat or significant risk to the health or safety of themselves or others. The requirements listed in this document are the minimum levels of knowledge, skills, or abilities. This document does not create an employment contract, implied or otherwise, other than an at‑will relationship.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI Infrastructure
Senior Software Engineer, AI Infrastructure

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401k plan
Paid vacation and sick leave
+2
Senior Software Engineer, AI Infrastructure
Senior Software Engineer, AI Infrastructure

The Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision coverage
401(k) plan
Generous paid vacation and sick leave
+2
Senior Software Engineer, AI Infrastructure
Senior Software Engineer, AI Infrastructure

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401(k) plan enrollment
Monthly stipends for commuting and fitness
+1
Infrastructure Engineer
Infrastructure Engineer

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 100,000 - 151,000
Catered lunches five times a week
Fitness and wellbeing expense assistance
Annual bonuses
Infrastructure Engineer
Infrastructure Engineer

The Allen Institute for AI • Seattle (WA)

On-site
USD 128,000 - 189,000
Catered lunches
Kayaks
Competitive pay
+1
Senior People Operations Business Partner
Senior People Operations Business Partner

The Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 119,000 - 180,000
Medical, dental, vision coverage
401k plan
Generous paid vacation and personal days
+1
Senior People Operations Business Partner
Senior People Operations Business Partner

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 119,000 - 180,000
Health, dental, and vision insurance
401k plan
Monthly stipends for commuting and fitness
Senior Software Engineer, Agent Frameworks
Senior Software Engineer, Agent Frameworks

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision coverage
Health savings account options
401k plan
+3
Senior Software Engineer, Agent Frameworks
Senior Software Engineer, Agent Frameworks

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision coverage
401k plan
Annual bonuses
Senior Software Engineer, Agent Frameworks
Senior Software Engineer, Agent Frameworks

The Allen Institute for AI • Seattle (WA)

On-site
USD 126,000 - 189,000
Annual bonus
401k plan
Commuting stipend $125/mo
+2