Senior/Staff AI Engineer

Data Direct Networks

California (MO)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Vacation plans
Paid holidays
Bonus programs
Tuition reimbursement
Employee referral program
Excellent medical, dental and vision
Paid leave programs
Anniversary and recognition awards

Job summary

Data Direct Networks is seeking an engineer to build and optimize LLM serving and inference systems for production environments. You will work on performance across GPU and CPU pathways and tackle bottlenecks in memory, storage, and throughput.

You will design and scale systems to support RAG and retrieval-heavy AI workloads, contributing to infrastructure where storage architecture and system efficiency directly affect AI performance.

Qualifications

  • Experience building production AI systems and serving workloads.
  • Understanding of compute/memory/storage interactions on inference.
  • Hands-on work near the systems layer with GPU/CPU tuning.
  • Ownership across model serving, retrieval, caching, or distributed performance.

Responsibilities

  • Build and optimize LLM serving/inference systems for production.
  • Improve performance across GPU and CPU pathways.
  • Scale systems for RAG and retrieval-heavy workloads.
  • Address bottlenecks in memory, storage, and throughput.

Skills

AI systems
Model serving
Inference performance
System optimization
Distributed systems

Job description

We’re Looking for the Best and Brightest

We are the world’s leading data intelligence platform that reliably accelerates massive datasets for actionable real-time insights. Join our team to help the best and brightest minds tackle the world’s biggest challenges in business, science, medicine, academia and government.

What you’ll do

Build and optimize LLM serving and inference systems for production environments.

Improve performance across GPU and CPU pathways.

Work on KV cache, memory, storage, and throughput bottlenecks.

Design and scale systems that support RAG and retrieval-heavy AI workloads.

Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance.

Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure.

What we’re looking for

An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models.

Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture.

Deep hands‑on experience working close to the systems layer— for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency.

Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work.

The ability to move comfortably between architecture decisions and hands‑on implementation, especially in environments where efficiency and scale matter.

A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work.

PhD preferred, but far less important than having built serious systems in the real world.

Why this role is compelling

This is not a “prompt engineering” job.

This is not an “AI wrapper” job.

This is not a generic backend role with AI sprinkled on top.

This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.

If you want to work on the real mechanics of AI performance—serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale—this is where that work happens.

Who will love this role

Engineers who enjoy deep systems problems.

Builders who care about performance, scale, and architecture.

People who want to work where AI meets infrastructure.

Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features.

Who should not apply

This role is not for:

Purely academic researchers without meaningful production ownership.

Generic software engineers without clear AI systems or inference depth.

Candidates focused mainly on prompt engineering or lightweight application integrations.

MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems.

Employment Type

Full time

Location Type

On-site

Benefits
  • Highly Competitive Vacation Plans
  • Paid Holidays
  • Bonus Programs
  • Tuition Reimbursement
  • Employee Referral Program
  • Excellent Medical, Dental and Vision Benefits
  • Paid Leave Programs
  • Anniversary and Recognition Awards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Inference Engineer
Staff Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 170,000 - 230,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Northern (KY)

Hybrid
USD 150,000 - 210,000
Performance bonus
Equity grant
Health, dental, vision fully paid
+5
AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Santa Clara (CA), Boston (MA)

On-site
USD 150,000 - 210,000
Health, Dental, and Vision coverage
401(k) match
Life, Disability and AD&D insurance
+4
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Senior AI Engineer
Senior AI Engineer

People In AI • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior/Staff AI Engineer
Senior/Staff AI Engineer

DDN • San Francisco (CA)

On-site
USD 150,000 - 250,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000