Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual performance bonus
Equity

Job summary

DevHub is building scalable LLM infrastructure to power large-scale inference workloads. You will work with cross-functional teams to improve reliability, latency, and efficiency of distributed AI systems in a fast-growing environment.

We seek a senior backend/infrastructure engineer with 8+ years' experience in distributed systems, scalable APIs, and cloud-native infrastructure. Expertise in ML infrastructure, GPU orchestration, and SOA is essential; PyTorch and vLLM experience is a plus.

Qualifications

  • Requires 8+ years of experience in backend or infrastructure engineering with expertise in distributed systems and scalable APIs.
  • Experience with ML infrastructure, GPU orchestration, and service-oriented architecture is essential.

Skills

Backend Engineering
Infrastructure Engineering
Distributed Systems
Scalable APIs
Cloud-native Infrastructure
Real-time Serving
ML Infrastructure
GPU Orchestration
Service-oriented Architecture
Deployment Pipelines
System Observability
LLM Infrastructure
Model Inference

Tools

PyTorch
vLLM

Job description

DevHub is building scalable LLM infrastructure to power large-scale inference workloads. You will work with cross-functional teams to improve reliability, latency, and efficiency of distributed AI systems in a fast-growing environment.

We seek a senior backend/infrastructure engineer with 8+ years' experience in distributed systems, scalable APIs, and cloud-native infrastructure. Expertise in ML infrastructure, GPU orchestration, and SOA is essential; PyTorch and vLLM experience is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Member of the Technical Staff- LLMs
Member of the Technical Staff- LLMs

Amadeus Search • San Francisco (CA)

Hybrid
USD 170,000 - 220,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6