ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources

San Francisco (CA)

On-site

USD 200,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k)
Unlimited PTO
Modern engineering workspace
Equity participation

Job summary

IC Resources seeks an engineer to accelerate production AI systems, focusing on speed and efficiency of large language model inference. You will optimize GPU-heavy pipelines and scale distributed GPU environments in a fast-moving, innovative AI company.

You will work with research and infrastructure teams to translate cutting-edge models into reliable production systems, improving latency and throughput while leveraging modern hardware.

Qualifications

  • Degree in Computer Science, Electrical Engineering, or related discipline.
  • Strong Python and C++ (CUDA preferred) programming skills.
  • Knowledge of modern LLM serving technologies (vLLM, SGLang, PyTorch, etc.).
  • Understanding of GPU architecture and parallel computing.
  • Experience with model serving, distributed inference, quantization, or batching strategies.
  • Strong profiling, debugging, and systems optimization skills.

Responsibilities

  • Improve speed and efficiency of production AI systems.
  • Evaluate performance bottlenecks and build benchmarking tools.
  • Optimize inference pipelines and scale distributed GPU environments.
  • Collaborate with research and infrastructure teams to deploy new modeling techniques.
  • Continuously improve latency, throughput, and hardware utilization.

Skills

Python
C++
CUDA
vLLM
SGLang
PyTorch
GPU architecture
Distributed inference
Profiling
Debugging

Education

Degree in Computer Science or Electrical Engineering

Tools

CUDA toolkit

Job description

IC Resources seeks an engineer to accelerate production AI systems, focusing on speed and efficiency of large language model inference. You will optimize GPU-heavy pipelines and scale distributed GPU environments in a fast-moving, innovative AI company.

You will work with research and infrastructure teams to translate cutting-edge models into reliable production systems, improving latency and throughput while leveraging modern hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Infra Engineer — Peak LLM on GPUs (Hybrid)
AI Infra Engineer — Peak LLM on GPUs (Hybrid)

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
GPU-Optimized LLM Inference Engineer
GPU-Optimized LLM Inference Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000