Senior Inference Engineer

Loft Labs, Inc. dba vCluster Labs

United States

On-site

USD 140,000 - 190,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Loft Labs, Inc. dba vCluster Labs is seeking a Senior Inference Engineer to build and optimize the platform's inference layer. The role involves deploying models to production and managing the entire query-to-response pipeline in a remote environment.

The candidate will deploy models, operate serving infra with vLLM/SGLang/TensorRT-LLM, and influence the engineering direction in collaboration with Product for future developments.

Qualifications

  • Experience deploying and serving LLMs using frameworks such as vLLM, SGLang, or TensorRT-LLM.
  • Hands-on knowledge of inference optimization techniques like quantization, batching, caching, and routing.
  • Strong programming skills in Python or Golang with real production code experience.
  • Ability to communicate technical concepts clearly to both technical and non-technical stakeholders.
  • Familiarity with containerized environments (Docker, Kubernetes) is a plus.

Responsibilities

  • Deploying models to production and owning the full pipeline from query to response.
  • Building and operating serving infrastructure using frameworks like vLLM, SGLang, or TensorRT-LLM.
  • Leading the engineering direction for inference, collaborating with Product to shape future developments.

Skills

LLM deployment
Inference optimization
Python
Golang
Communication

Tools

Docker
Kubernetes

Job description

Partnering directly with the CTO, the full-time Senior Inference Engineer will build and optimize the platform's inference layer, deploying models to production and managing the entire query-to-response pipeline in a remote environment.

Key responsibilities
  • Deploying models to production and owning the full pipeline from query to response
  • Building and operating serving infrastructure using frameworks like vLLM, SGLang, or TensorRT-LLM
  • Leading the engineering direction for inference, collaborating with Product to shape future developments
Required qualifications
  • Experience deploying and serving LLMs using frameworks such as vLLM, SGLang, or TensorRT-LLM
  • Hands‑on knowledge of inference optimization techniques like quantization, batching, caching, and routing
  • Strong programming skills in Python or Golang with real production code experience
  • Ability to communicate technical concepts clearly to both technical and non-technical stakeholders
  • Familiarity with containerized environments (Docker, Kubernetes) is a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior Inference Engineer - Production LLM Pipelines
Remote Senior Inference Engineer - Production LLM Pipelines

Loft Labs, Inc. dba vCluster Labs • United States

On-site
USD 140,000 - 190,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior Inference Platform Engineer - Data Center
Senior Inference Platform Engineer - Data Center

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000