Member of Technical Staff (AI Inference Engineer)

Kindredventures

Palo Alto (CA)

On-site

USD 190,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive health insurance
Dental insurance
Vision insurance
401(k) plan

Job summary

A leading financial technology firm in California seeks an AI Inference engineer to join its team. The role involves developing APIs for AI inference, improving system reliability, and optimizing LLM performance. Required qualifications include experience with ML systems, deep learning frameworks, and GPU programming. The position offers a hybrid work model and competitive salary ranging from $190,000 to $250,000, alongside health benefits and a 401(k) plan.

Qualifications

  • Experience with ML systems and deep learning frameworks required.
  • Familiarity with LLM architectures and optimization techniques necessary.
  • Understanding of GPU architectures and CUDA programming.

Responsibilities

  • Develop APIs for AI inference for internal and external users.
  • Benchmark and resolve bottlenecks in inference stack.
  • Enhance system reliability and respond to outages.
  • Research and implement LLM inference optimizations.

Skills

Experience with ML systems
Deep learning frameworks (e.g. PyTorch, TensorFlow)
Common LLM architectures knowledge
Inference optimization techniques
Understanding of GPU architectures
GPU kernel programming using CUDA

Tools

Python
Rust
C++
PyTorch
Triton
CUDA
Kubernetes

Job description

Location

San Francisco

Employment Type

Full time

Location Type

Hybrid

Department

AI

We are looking for an AI Inference engineer to join our growing team. Our current stack is Python, Rust, C++, PyTorch, Triton, CUDA, Kubernetes. You will have the opportunity to work on large-scale deployment of machine learning models for real-time inference.

Responsibilities
  • Develop APIs for AI inference that will be used by both internal and external customers

  • Benchmark and address bottlenecks throughout our inference stack

  • Improve the reliability and observability of our systems and respond to system outages

  • Explore novel research and implement LLM inference optimizations

Qualifications
  • Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow, ONNX)

  • Familiarity with common LLM architectures and inference optimization techniques (e.g. continuous batching, quantization, etc.)

  • Understanding of GPU architectures or experience with GPU kernel programming using CUDA

The cash compensation range for this role is $190,000 - $250,000.

Final offer amounts are determined by multiple factors, including, experience and expertise, and may vary from the amounts listed above.

Equity: In addition to the base salary, equity may be part of the total compensation package. Benefits: Comprehensive health, dental, and vision insurance for you and your dependents. Includes a 401(k) plan.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid AI • Boston (MA)

Hybrid
USD 180,000 - 240,000
Equity
Health insurance
401(k) matching
+2
Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1