AI Infra Systems Engineer — Hybrid, GPU & Scale

Delos Data Inc

Palo Alto (CA)

Hybrid

USD 140,000 - 200,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
401k
Benefits

Job summary

Delos Data Inc. seeks a System Software Engineer to bridge high-level AI frameworks and low-level system software, enabling efficient execution of large-scale AI models across GPUs.

This hybrid role, based in Palo Alto, CA or Charlottesville, VA, focuses on building communication and execution primitives for scalable AI workloads with meaningful equity and benefits.

Qualifications

  • Proficiency in C++ and Python with strong systems programming basics.
  • Proficient in a Linux development environment.
  • BS/MS in Computer Engineering or Computer Science or a related field.

Responsibilities

  • Collaborate across the stack to influence the design of foundational AI infrastructure.
  • Identify and resolve performance bottlenecks in distributed training and inference.
  • Conduct rigorous benchmarking on multi-node clusters.

Skills

C++
Python
Linux
Systems programming

Education

CS/CE degree

Tools

CUDA
PyTorch
JAX
DeepSpeed

Job description

System Software Engineer - AI

About us:

We are a stealth-mode startup building foundational technology to address performance, scalability, and resiliency challenges in large-scale AI data center clusters. We are backed by top-tier VC firms and notable angel investors.

The company is led by experienced builders and operators who have founded companies, taken them to scale, and exited successfully. We work with a strong sense of unity and shared responsibility, and we expect trust, integrity, and respect in how we collaborate and make decisions. We hold ourselves accountable to one another and to the quality of the work we deliver.

Headquartered in Silicon Valley, we operate across a mix of remote and on-site locations in the U.S. and Canada. We aim to create an environment where people are treated fairly, supported in their growth, and are empowered to do meaningful work alongside others who take the craft seriously.

Some Recent Press:

https://www.eetimes.com/startup-boosts-scale-up-to-1000-gpus-in-a-single-domain/

We are looking for:

We arelooking for a talented System Software Engineer to help us redefine the infrastructure layer of AI. In this role, you will bridge the gap between high-level AI frameworks and low-level system software. You will be responsible for designing and implementing the communication and execution primitives that allow large-scale AI models to run efficiently across thousands of GPUs. We are looking for a "builder" who thrives in the early stages of a product’s lifecycle and is passionate about solving the "hard" systems problems of the generative AI era.

Key Responsibilities:

  • Collaborate across the stack to influence the design of our foundational technology, ensuring it meets the needs of next-generation AI models.

  • Identify and resolve performance bottlenecks in distributed training and inference workloads through deep-dive analysis of the software-hardware interface.

  • Conduct rigorous performance benchmarking and characterization on multi-node clusters.

Required Skills and Qualifications:

  • Strong proficiency in C++ and Python, with a deep understanding of systems programming fundamentals (memory management, concurrency, OS internals).

  • Proficient in a Linux development environment.

Desired Skills:

  • Experience with GPU programming (CUDA) and performance optimization for parallel architectures.

  • Familiarity with distributed AI frameworks (PyTorch, JAX, or DeepSpeed) and/or inference engines (vLLM, SGLang, Dynamo/TRT-LLM).

  • Hands-on experience with large-scale cluster orchestration and telemetry tools.

Education:

  • Bachelor's or Master's degree in Computer Engineering, Computer Science, or a related field.

Location:

This is a hybrid role based in Palo Alto or Charlottesville, VA.

Compensation:

Target base salary for this role is $140,000 - $200,000 per year + meaningful equity + benefits + 401k. Our salary ranges are determined by role, level, experience, and location.

We are an equal opportunity employer. We value a range of perspectives and experiences and make employment decisions based on merit and business needs. We do not discriminate on the basis of legally protected characteristics.

Agency Note:

We do not accept resumes from agencies or search firms. Please do not forward candidate profiles through our careers page, email, LinkedIn messages, or directly to company employees. Any resumes submitted will be deemed the property of the company, and no fees will be paid in the event the candidate is hired.

#LI-EW1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
System Software Engineer - AI
System Software Engineer - AI

Delos Data Inc • Palo Alto (CA)

On-site
USD 140,000 - 200,000
Equity
401k
Benefits
Software Development Engineer in Test - AI
Software Development Engineer in Test - AI

Delos Data • Charlottesville (VA)

On-site
USD 140,000 - 200,000
Equity
401k
Meaningful benefits
Software Development Engineer in Test - AI
Software Development Engineer in Test - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 220,000
Equity
Benefits package
401k
Software Development Engineer in Test - AI
Software Development Engineer in Test - AI

Delos Data Inc • Palo Alto (CA)

Hybrid
USD 140,000 - 220,000
Equity
Benefits
401(k)
Distinguished Engineer, Scaled Out Inferencing
Distinguished Engineer, Scaled Out Inferencing

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Infrastructure Software Engineer - DGX Cloud
Senior AI Infrastructure Software Engineer - DGX Cloud

Socket.dev • Santa Clara (UT)

On-site
USD 184,000 - 357,000
Equity
Health benefits
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Principal System Software Engineer, AI Inference Execution
Principal System Software Engineer, AI Inference Execution

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 260,000