Senior Software Engineer, AI Infrastructure

Ai2

Seattle (WA)

On-site

USD 126,000 - 189,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) plan enrollment
Monthly stipends for commuting and fitness
Generous paid leave

Job summary

Ai2 in Seattle is seeking an expert Systems Engineer to build infrastructure enabling transparent AI breakthroughs. You will independently deliver critical systems, optimize performance, and bridge gaps between researchers and GPU clusters.

The ideal candidate has 8+ years of experience, proficiency in Go/Python, and solid Linux expertise. The position offers a base salary between $126,000 and $189,000, along with a comprehensive benefits package including health, 401(k), and generous paid leave.

Qualifications

  • 8+ years of professional experience developing business-critical software and operating large-scale compute infrastructure.
  • Bachelor’s degree in a related field; an advanced degree may substitute for equivalent years of technical work experience.
  • Expert knowledge of Linux internals and container runtimes such as Docker.

Responsibilities

  • Independently design and deliver critical systems spanning the entire stack.
  • Build innovative tooling and software-defined infrastructure.
  • Conduct root-cause analysis on complex distributed system failures.

Skills

Proficiency in Go
Proficiency in Python
Linux internals expertise
Communication skills
Distributed Systems design and debugging
Experience with workload schedulers

Education

Bachelor’s degree in a related field

Tools

Docker
Kubernetes
Slurm

Job description

Location: Seattle, Washington – you are expected to work from our Seattle offices. On‑site requirements vary by position and team; if you have specific questions about on‑site work arrangements, ask your recruiter.

Base salary range: $126,000 - $189,000 with generous bonus plans to provide a competitive compensation package.

Who You Are

An expert systems engineer who works at the intersection of high‑level software orchestration and low‑level system performance. You are comfortable designing resource‑allocation algorithms in Go as well as debugging low‑level HPC issues. You lead by example, blending the rigor of a Senior Software Engineer with the pragmatic urgency of an HPC operator. You build systems that thrive under the pressure of training world‑class AI models.

Your Next Challenge

You will build the infrastructure that makes transparent and accessible AI breakthroughs possible. Your role is to bridge the gap between researchers and our GPU clusters, ensuring that when a researcher submits a job, the software schedules it intelligently and the hardware executes it flawlessly.

Your Responsibilities
  • Full‑Stack Ownership: Independently design and deliver critical systems spanning the entire stack—from the Beaker job scheduler to the execution runtime.
  • System Automation: Build innovative tooling and software‑defined infrastructure to accelerate researcher velocity and automate cluster health management.
  • Performance Optimization: Conduct root‑cause analysis on complex distributed system failures and implement optimizations for distributed workloads.
  • Technical Input & Ownership: Provide input into the roadmap for managing large‑scale HPC systems, including deployment of compute, networking, and storage in partnership with leadership.
  • Mentorship & Culture: Review code/design docs, mentor team members, and drive process improvements within the team.
  • Collaboration: Communicate effectively with internal research staff to share system designs, gather feedback, and support engineers on implementation tasks.
What You’ll Need (Qualifications)
  • 8+ years of professional experience developing business‑critical software and operating large‑scale compute infrastructure. Proficiency in Go and/or Python preferred.
  • Bachelor’s degree in a related field; an advanced degree may substitute for equivalent years of technical work experience.
  • Linux Expertise: Expert knowledge of Linux internals and container runtimes such as Docker.
  • Distributed Systems Expertise: Proven track record of designing, debugging, and optimizing high‑scale distributed systems and databases.
  • Communication: Exceptional written skills and the ability to drive consensus across diverse groups of researchers and engineers.
  • Principled approach to engineering: Care about how systems are built and excitement for the unique constraints of a non‑profit research environment.
Bonus Qualifications
  • Experience with workload schedulers (e.g., Kubernetes, Slurm) and high‑performance networking (NCCL, InfiniBand).
  • Prior experience training or fine‑tuning frontier AI models.
  • Deep systems administration or Site Reliability Engineering (SRE) background in an HPC context.
  • Contribution to open‑source infrastructure or orchestration projects.
  • Familiarity with on‑prem storage systems such as WEKA and Ceph.
Physical Demands and Work Environment
  • Ability to remain in a stationary position for long periods.
  • Clear communication of information and ideas to others.
  • Attention to detail and observation at close range.
  • Capacity to work under deadlines.
Benefits
  • Medical, dental, vision, and an employee assistance program for team members and their families.
  • Health savings account and healthcare reimbursement arrangement; flexible spending account plans for medical and dependent care.
  • 401(k) plan enrollment.
  • Monthly stipend of $125 for commuting or internet expenses and $200 for fitness and wellbeing.
  • Up to ten sick days, up to seven personal days, up to 20 vacation days, and twelve paid holidays per year.
  • Annual bonuses and participation in a long‑term incentive plan.
Equal Opportunity and Legal Statements

Ai2 is a proud Equal Opportunity employer. We do not discriminate based on race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity or expression, gender, gender expression, transgender status, sexual stereotypes, age, veteran status, disability, or other protected characteristics. We welcome applicants from outside the United States. This employer participates in E‑Verify and will provide the federal government with your Form I‑9 information to confirm that you are authorized to work in the U.S. We are committed to providing reasonable accommodations to employees and applicants with disabilities in accordance with the Americans with Disabilities Act (ADA).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI Infrastructure
Senior Software Engineer, AI Infrastructure

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401k plan
Paid vacation and sick leave
+2
Senior Engineering Manager, AI Infrastructure
Senior Engineering Manager, AI Infrastructure

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 146,000 - 221,000
Medical, dental, vision
401k plan
Commuting stipend
+3
Senior Software Engineer, AI Infrastructure
Senior Software Engineer, AI Infrastructure

The Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision coverage
401(k) plan
Generous paid vacation and sick leave
+2
Staff AI Engineer
Staff AI Engineer

Socket.dev • Houston (TX)

On-site
USD 177,000 - 210,000
AI Software Engineer
AI Software Engineer

webAI • Austin (TX)

On-site
USD 100,000 - 140,000
Competitive salary
Equity options
Comprehensive health benefits
+8
Staff HPC Engineer
Staff HPC Engineer

Biohub • San Francisco (CA)

Hybrid
USD 214,000 - 268,000
401(k) employer match
Paid volunteer time off
Relocation support
Senior AI Infra Engineer — HPC & Scheduler
Senior AI Infra Engineer — HPC & Scheduler

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision insurance
401(k) plan enrollment
Monthly stipends for commuting and fitness
+1
Senior Software Engineer, Fullstack
Senior Software Engineer, Fullstack

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 126,000 - 189,000
Comprehensive medical, dental, and vision coverage
401k enrollment
Monthly commuting/internet support
+2
Senior Software Engineer, Agent Frameworks
Senior Software Engineer, Agent Frameworks

Ai2 • Seattle (WA)

On-site
USD 126,000 - 189,000
Medical, dental, and vision coverage
Health savings account options
401k plan
+3
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits