Inference Engineer

techire.®

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance (including dental &视
Dental insurance
Vision insurance
401(k)
Relocation support
Immigration support
Meals in office

Job summary

techire.® seeks engineers to build the infrastructure that runs next-generation multimodal foundation models at scale. You will design and implement real-time inference pipelines and distributed systems to support low latency, high reliability production workloads.

You’ll collaborate with researchers to productionise new model architectures and drive ownership from day one, focusing on end-to-end reliability and observability across the stack.

Qualifications

  • Strong software engineering fundamentals for large-scale systems.
  • Experience with ML inference pipelines or serving generative models in production.
  • Ability to work through ambiguous technical challenges and deliver zero-to-one systems.
  • Experience implementing modern machine learning research into production environments.

Responsibilities

  • Build low-latency inference and serving infrastructure for foundation models across Transformers, SSMs and hybrid architectures.
  • Design scalable, reliable distributed systems that support production AI workloads.
  • Develop monitoring and observability across the inference stack.
  • Work closely with researchers to productionise new model architectures.
  • Help shape technical direction with significant ownership from day one.

Skills

Distributed systems
ML inference
Production engineering
Ambiguity tolerance
Research-to-prod translation

Tools

CUDA
Triton
vLLM
SGLang
Continuous Batching

Job description

Help build the inference stack behind the next generation of multimodal foundation models.

Most inference roles are about making existing models run faster.

This one is about helping define how entirely new model architectures are served at scale.

You’ll be building real-time multimodal AI capable of processing enormous streams of text, audio and video. The research is pushing beyond today’s transformer limitations, and your work will make those models usable in production.

You’ll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models to run with low latency, high reliability and at scale. If you enjoy solving systems problems where every millisecond matters, you’ll feel at home here.

Your focus
  • Build low-latency inference and serving infrastructure for foundation models across Transformers, SSMs and hybrid architectures.
  • Design scalable, reliable distributed systems that support production AI workloads.
  • Develop monitoring and observability across the inference stack.
  • Work closely with researchers to productionise new model architectures.
  • Help shape technical direction with significant ownership from day one.
You’ll bring
  • Strong software engineering fundamentals and experience building large-scale distributed systems.
  • Experience with ML inference pipelines or serving generative models in production.
  • The ability to work through ambiguous technical challenges and deliver zero-to-one systems.
  • Experience implementing modern machine learning research into production environments.

Experience with vLLM, SGLang, Continuous Batching, CUDA or Triton would be particularly valuable but isn’t essential.

Alongside compensation, you’ll receive fully covered medical, dental and vision insurance, 401(k), relocation and immigration support, plus daily meals in the office.

If you’re interested in building the infrastructure that enables the next generation of AI models to run in real time, we’d love to tell you more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Engineer
Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • Redwood City (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

Adaption Labs • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1