LLM Systems Engineer: Inference Hardware & Research

Netpreme

Cambridge (MA)

On-site

USD 190,000 - 230,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance
Visa sponsorship
Daily lunch stipend
401k match
Equity grant

Job summary

Netpreme is seeking a motivated LLM Systems Engineer to explore and prototype inference systems for next‑gen hardware. You’ll guide the hardware team on product definitions and push ML systems research forward in a hands‑on role.

The role blends engineering and research, requiring a track record in ML systems work and strong familiarity with modern LLM inference stacks. On-site roles available in Santa Clara, CA or Boston, MA.

Qualifications

  • MS or PhD in computer systems with focus on ML inference or distributed systems.
  • Experience contributing to core LLM inference infrastructures (vLLM, SGLang, TensorRT).
  • Accelerator programming experience (CUDA, JAX/Pallas, ROCm).

Responsibilities

  • Prototype and optimize emerging ML inference systems.
  • Develop memory models for expandable vRAM.
  • Write efficient GPU kernels for data movement.
  • Benchmark inference engines in simulations and real hardware.

Skills

LLM inference systems
GPU kernel optimization
Memory modeling
Performance benchmarking
Design-space exploration

Education

MS/PhD in computer systems

Tools

CUDA
JAX/Pallas
ROCm
TensorRT
vLLM
SGLang

Job description

Netpreme is seeking a motivated LLM Systems Engineer to explore and prototype inference systems for next‑gen hardware. You’ll guide the hardware team on product definitions and push ML systems research forward in a hands‑on role.

The role blends engineering and research, requiring a track record in ML systems work and strong familiarity with modern LLM inference stacks. On-site roles available in Santa Clara, CA or Boston, MA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Systems Architect
LLM Inference Systems Architect

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
Staff Research Engineer - LLM Inference & Serving
Staff Research Engineer - LLM Inference & Serving

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Staff Research Scientist - LLM Inference & Systems
Staff Research Scientist - LLM Inference & Systems

Modal • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Systems Engineer for RL & Inference Infrastructure
ML Systems Engineer for RL & Inference Infrastructure

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
Senior ML Systems Scientist: LLM/VLM Inference Expert
Senior ML Systems Scientist: LLM/VLM Inference Expert

Nebius B.V. • Palo Alto (CA)

On-site
USD 210,000 - 320,000