LLM Inference Frameworks & Optimizations Engineer

Together AI

San Francisco (CA)

On-site

USD 160,000 - 230,000

Full time

44 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Benefits package

Job summary

Together AI is building state-of-the-art LLM inference infrastructure to enable scalable, low-latency deployment for multimodal models. You will design and optimize distributed inference engines, collaborate with hardware and research teams, and push performance, scalability and cost-efficiency.

We seek an engineer with 3+ years in deep learning inference or HPC, proficient in Python and C++/CUDA, and hands-on experience with TensorRT, MoE and GPU optimization.

Qualifications

  • Must have 3+ years in deep learning inference frameworks or high-performance computing.
  • Proficient in Python and C++/CUDA for high-performance inference.
  • Familiar with LLM inference frameworks (e.g., TensorRT‑LLM, vLLM, SGLang, TGI).
  • Background in GPU programming, compilers, quantization or GPU cluster scheduling.

Responsibilities

  • Design and develop fault‑tolerant, high‑concurrency distributed inference engine for text, image, and multimodal models.
  • Implement and optimize distributed inference strategies, including MoE parallelism, tensor parallelism, and pipeline parallelism.
  • Apply CUDA graph optimizations, TensorRT/TRT‑LLM graph optimizations, and PyTorch‑based compilation to enhance efficiency and scalability.

Skills

Python
C++/CUDA
DL inference
TensorRT
Mooncake KV
GPU/TPU

Tools

Kubernetes

Job description

Together AI is building state-of-the-art LLM inference infrastructure to enable scalable, low-latency deployment for multimodal models. You will design and optimize distributed inference engines, collaborate with hardware and research teams, and push performance, scalability and cost-efficiency.

We seek an engineer with 3+ years in deep learning inference or HPC, proficient in Python and C++/CUDA, and hands-on experience with TensorRT, MoE and GPU optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Equity
Health insurance
Benefits package
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
LLM/VLM Inference Optimization Research Engineer
LLM/VLM Inference Optimization Research Engineer

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1