LLM Inference Systems Architect

Netpreme

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance
Visa sponsorship
Lunch stipend
401k match

Job summary

Netpreme is seeking an LLM Systems Engineer to prototype and optimize advanced inference systems on cutting-edge hardware. The role blends engineering with research, guiding hardware teams on product definitions and pushing the frontiers of ML inference software.

The candidate should have a strong track record in ML systems research and deep familiarity with industry-standard LLM inference systems, accelerator programming, and performance engineering.

Qualifications

  • MS or PhD in computer systems, ideally with a focus on LLM inference and/or distributed systems.
  • Prior experience contributing to the core LLM inference infrastructures (vLLM, SGLang, TensorRT, etc.).
  • Prior experience in accelerator programming (e.g. CUDA, JAX/Pallas, ROCm).
  • Advanced computer architectures and performance engineering skills.

Responsibilities

  • Prototype and optimize emerging ML inference systems.
  • Develop novel memory models for expandable vRAM.
  • Write efficient GPU kernels for data movement.
  • Perform design-space exploration, implementation, and benchmarking of inference engines, both in simulations and on real hardware.

Skills

LLM inference
GPU performance
Distributed systems
Performance engineering

Education

MS or PhD in computer systems

Tools

CUDA
JAX/Pallas
ROCm
TensorRT

Job description

Netpreme is seeking an LLM Systems Engineer to prototype and optimize advanced inference systems on cutting-edge hardware. The role blends engineering with research, guiding hardware teams on product definitions and pushing the frontiers of ML inference software.

The candidate should have a strong track record in ML systems research and deep familiarity with industry-standard LLM inference systems, accelerator programming, and performance engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Systems Engineer: Inference Hardware & Research
LLM Systems Engineer: Inference Hardware & Research

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2
LLM Inference Systems Performance Architect
LLM Inference Systems Performance Architect

Doist • San Jose (CA)

On-site
USD 245,000 - 325,000
Health Insurance
Dental Insurance
Vision Insurance
+8
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental