Global Inference Library Engineer

Jobot

San Francisco (CA)

On-site

USD 175,000 - 250,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity (startup)
Competitive compensation
Healthcare, vision, dental

Job summary

Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from complex hardware.

We value experience in AI/ML infrastructure, CUDA/Rocm, and model-serving frameworks, with strong Python and C++/Rust skills. Equity, competitive pay, and excellent benefits are offered.

Qualifications

  • Experience in AI/ML infrastructure or HPC software.
  • Programming in Python and C++, Rust, or similar systems languages.
  • Experience with LLM inference frameworks and model-serving infrastructure.
  • Hands-on CUDA, ROCm, Triton or similar GPU programming.
  • Understanding of transformer/LLM architectures and memory management.
  • Experience benchmarking AI workloads across hardware environments.
  • Familiarity with batching, attention, KV caching, quantization, and memory management.

Responsibilities

  • Build and maintain a high-performance inference library for AI models.
  • Optimize performance across diverse compute architectures.
  • Work with CUDA/ROCm and related frameworks to accelerate workloads.
  • Analyze and improve batching, KV caching, and memory usage.

Skills

AI/ML infrastructure
Python
C++
Rust
LLM inference
GPU programming
Performance benchmarking
Transformer/LLM architectures
Memory management
Kernel optimization

Tools

CUDA
ROCm
Triton
vLLM
TensorRT-LLM
SGLang

Job description

Job Details

This Jobot Job is hosted by: Grant Greenhalgh.

Salary: $175,000 - $250,000 per year.

A bit about us

We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.

Why join us?
  • Well-funded by leading tech investors
  • Cutting edge technical problems with complex solutions
  • Lucrative Equity in a seed stage startup
  • Competitive compensation
  • Excellent benefits (healthcare, vision, dental)
Job Details

We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.

What We’re Looking For
  • Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software
  • Strong programming experience with Python and C++, Rust, or similar systems languages
  • Experience with LLM inference frameworks and model-serving infrastructure
  • Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies
  • Experience developing, integrating, or optimizing performance-critical compute kernels
  • Understanding of modern transformer and LLM architectures
  • Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management
  • Experience benchmarking and profiling AI workloads across different hardware environments
  • Strong understanding of GPU or accelerator architecture and performance characteristics
  • Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable

Jobot is an Equal Opportunity Employer. We provide an inclusive work environment that celebrates diversity and all qualified candidates receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, age (40 and over), disability, military status, genetic information or any other basis protected by applicable federal, state, or local laws. Jobot also prohibits harassment of applicants or employees based on any of these protected categories. It is Jobot's policy to comply with all applicable federal, state and local laws respecting consideration of unemployment status in making hiring decisions.

Sometimes Jobot is required to perform background checks with your authorization. Jobot will consider qualified candidates with criminal histories in a manner consistent with any applicable federal, state, or local law regarding criminal backgrounds, including but not limited to the Los Angeles Fair Chance Initiative for Hiring and the San Francisco Fair Chance Ordinance.

Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal.

By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Jobot, and/or its agents and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: jobot.com/privacy-policy

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global Inference Library Engineer
Global Inference Library Engineer

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1
LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1