AI Inference Infrastructure Engineer

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. in San Francisco, California is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models.

Your work will make inference faster, more cost-effective, more reliable, and more reproducible to enable our teams to focus on advancing model capabilities rather than managing bottlenecks. Our focus is on performant and efficient model inference both to power real-world applications and to accelerate research.

Qualifications

  • Bachelor’s degree or equivalent in CS or engineering.
  • Understanding of deep learning frameworks and system architectures.
  • Experience with inference serving for throughput and latency.
  • Collaborative across cross-functional teams.
  • Bias for action and ability to ship across stacks where opportunity is spotted.
  • Strong coding skills and debugging in large codebases.

Responsibilities

  • Collaborate with researchers to bring models into production.
  • Enable high-performance inference for novel architectures.
  • Design tools and architectures improving performance and efficiency.
  • Optimize codebase and compute fleets for GPU usage.
  • Extend orchestration frameworks for distributed inference and large batches.
  • Establish reliability and observability standards across the inference stack.
  • Publish learnings via internal docs, open-source libraries, or technical reports.

Skills

Bachelor level CS/Engineering
Deep learning frameworks
Inference serving systems
Cross-functional collaboration
Bias for action
Strong coding and debugging

Education

Bachelor’s degree or equivalent

Tools

PyTorch
JAX
SGLang
vLLM
Kubernetes
Ray
SLURM

Job description

Thinking Machines Lab Inc. in San Francisco, California is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models.

Your work will make inference faster, more cost-effective, more reliable, and more reproducible to enable our teams to focus on advancing model capabilities rather than managing bottlenecks. Our focus is on performant and efficient model inference both to power real-world applications and to accelerate research.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

Menlo Ventures • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Inference Infrastructure Engineer for Large-Scale AI
Inference Infrastructure Engineer for Large-Scale AI

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure & Platform Engineer
ML Infrastructure & Platform Engineer

Odyssey • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 445,000
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Inference Engineer: Real-Time ML, Hybrid, Equity
AI Inference Engineer: Real-Time ML, Hybrid, Equity

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000