Distributed LLM Inference & Optimization Engineer

Together AI

San Francisco (CA)

On-site

USD 160,000 - 230,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Startup equity
Health insurance
Competitive benefits

Job summary

Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale.

This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment.

Qualifications

  • 3+ years of experience in deep learning inference frameworks, distributed systems, or high-performance computing.
  • Familiar with at least one LLM inference framework (e.g., TensorRT-LLM, vLLM, SGLang, TGI).
  • Background knowledge in GPU programming (CUDA/Triton/TensorRT), compiler, model quantization, and GPU cluster scheduling.
  • Deep understanding of KV cache systems like Mooncake, PagedAttention, or custom in-house variants.
  • Proficient in Python and C++/CUDA for high-performance deep learning inference.
  • Deep understanding of Transformer architectures and LLM/VLM/Diffusion model optimization.
  • Knowledge of inference optimization, workload scheduling, CUDA graph, compiled, efficient kernels.
  • Strong analytical problem-solving skills with a performance-driven mindset and collaborative.

Responsibilities

  • Design and develop fault-tolerant, high-concurrency distributed inference engine for text, image, and multimodal generation models.
  • Implement and optimize distributed inference strategies, including Mixture of Experts (MoE) parallelism, tensor parallelism, pipeline parallelism for high-performance serving.
  • Apply CUDA graph optimizations, TensorRT/TRT-LLM graph optimizations, and PyTorch-based compilation (torch.compile), and speculative decoding to enhance efficiency and scalability.
  • Collaborate with hardware teams on performance bottleneck analysis, co-optimize inference performance for GPUs, TPUs, or custom accelerators.
  • Work closely with AI researchers and infrastructure engineers to develop efficient model execution plans and optimize E2E model serving pipelines.

Skills

Deep learning inference
LLM inference frameworks
GPU programming
KV cache systems
Python and C++/CUDA
Transformer optimization
Inference optimization
Soft skills

Tools

Kubernetes
RDMA
Ceph

Job description

Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale.

This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
LLM Inference Architect & Systems Optimizer
LLM Inference Architect & Systems Optimizer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits