Realtime Multimodal Inference Architect

techire.®

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance (including dental &视
Dental insurance
Vision insurance
401(k)
Relocation support
Immigration support
Meals in office

Job summary

techire.® seeks engineers to build the infrastructure that runs next-generation multimodal foundation models at scale. You will design and implement real-time inference pipelines and distributed systems to support low latency, high reliability production workloads.

You’ll collaborate with researchers to productionise new model architectures and drive ownership from day one, focusing on end-to-end reliability and observability across the stack.

Qualifications

  • Strong software engineering fundamentals for large-scale systems.
  • Experience with ML inference pipelines or serving generative models in production.
  • Ability to work through ambiguous technical challenges and deliver zero-to-one systems.
  • Experience implementing modern machine learning research into production environments.

Responsibilities

  • Build low-latency inference and serving infrastructure for foundation models across Transformers, SSMs and hybrid architectures.
  • Design scalable, reliable distributed systems that support production AI workloads.
  • Develop monitoring and observability across the inference stack.
  • Work closely with researchers to productionise new model architectures.
  • Help shape technical direction with significant ownership from day one.

Skills

Distributed systems
ML inference
Production engineering
Ambiguity tolerance
Research-to-prod translation

Tools

CUDA
Triton
vLLM
SGLang
Continuous Batching

Job description

techire.® seeks engineers to build the infrastructure that runs next-generation multimodal foundation models at scale. You will design and implement real-time inference pipelines and distributed systems to support low latency, high reliability production workloads.

You’ll collaborate with researchers to productionise new model architectures and drive ownership from day one, focusing on end-to-end reliability and observability across the stack.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA)

On-site
USD 168,000 - 227,000
Senior AI Infra Engineer — Real-Time Multimodal Inference
Senior AI Infra Engineer — Real-Time Multimodal Inference

Ambient AI, Inc. • Redwood City (CA)

Hybrid
USD 190,000 - 270,000
Stock options
Health, dental, vision
401(k)
+1
Inference Engineer
Inference Engineer

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Multimodal Inference Engineer — Scale GPU AI Models
Multimodal Inference Engineer — Scale GPU AI Models

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Staff AI Inference & Systems Engineer
Staff AI Inference & Systems Engineer

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Inference Engineer - Multimodal Large Models
Senior AI Inference Engineer - Multimodal Large Models

TikTok • San Jose (CA)

On-site
USD 212,000 - 450,000
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000