AI Inference Engineer

Socket.dev

Palo Alto (CA)

On-site

USD 230,000 - 350,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity opportunity (0.5%)
Professional growth
High-impact work

Job summary

Premier Global Links LLC in Palo Alto, CA is seeking a Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.

The role focuses on LLM inference, model serving, distributed systems, and inference runtime performance. Strong systems engineering skills and hands-on production experience are required; Rust experience is highly valued.

Qualifications

  • 2–10 years of experience in backend, distributed systems, or systems engineering.
  • Hands-on experience building, operating, or optimizing LLM inference or serving systems.
  • Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling.
  • Experience with production inference engines such as vLLM, SGLang, or TensorRT-LLM.

Responsibilities

  • Build and optimize production LLM inference and model-serving systems.
  • Develop inference runtime components using Rust and other systems-level technologies.
  • Design and implement batching, scheduling, request routing, and serving infrastructure.
  • Build and optimize KV cache and prefix caching systems.
  • Scale inference workloads across multi-GPU and multi-node environments.
  • Profile, benchmark, and optimize latency, throughput, reliability, and cost.
  • Work with inference engines such as vLLM, SGLang, or TensorRT-LLM.
  • Investigate performance bottlenecks across the inference stack.
  • Contribute to core architecture and technical decisions for the platform.
  • Collaborate with a small, hands-on engineering team in a fast-paced environment.

Skills

Rust
C++
Go
Python

Education

Bachelor's or Master's in CS/CE

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
NCCL

Job description

About the Role

Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.

This role is focused on LLM inference, model serving, distributed systems, and inference runtime performance. The ideal candidate has hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued.

Key Responsibilities
  • Build and optimize production LLM inference and model-serving systems.
  • Develop inference runtime components using Rust and other systems-level technologies.
  • Design and implement batching, scheduling, request routing, and serving infrastructure.
  • Build and optimize KV cache and prefix caching systems.
  • Scale inference workloads across multi-GPU and multi-node environments.
  • Profile, benchmark, and optimize latency, throughput, reliability, and cost.
  • Work with inference engines such as vLLM, SGLang, or TensorRT-LLM.
  • Investigate performance bottlenecks across the inference stack.
  • Contribute to core architecture and technical decisions for the platform.
  • Collaborate with a small, hands-on engineering team in a fast-paced environment.
Required Qualifications
  • 2–10 years of experience in backend, distributed systems, or systems engineering.
  • Hands-on experience building, operating, or optimizing LLM inference or serving systems.
  • Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling.
  • Experience with a production inference engine such as vLLM, SGLang, or TensorRT-LLM.
  • Strong programming experience with Rust, C++, Go, or systems-level Python/PyTorch.
  • Experience building performance-critical systems where latency, throughput, and cost are important.
  • Strong distributed systems and production software engineering fundamentals.
  • Ability to work on-site 5 days per week in Palo Alto, CA.
Preferred Qualifications
  • Production Rust experience.
  • CUDA or Triton kernel development experience.
  • Multi-GPU or multi-node serving experience.
  • Experience with NCCL, NVLink, or RDMA.
  • Experience with prefix caching, speculative decoding, or prefill/decode disaggregation.
  • Contributions to open-source inference projects such as vLLM, SGLang, or Dynamo.
  • Experience working on inference systems at an AI provider, accelerator company, research lab, or similar organization.
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field.
Technology Environment

Rust | C++ | Go | Python | PyTorch | vLLM | SGLang | TensorRT-LLM | CUDA | Triton | NCCL | NVLink | RDMA | LLM Inference | Distributed Systems

Compensation & Benefits
  • $230,000–$350,000 annual salary, based on experience and qualifications.
  • Equity opportunity starting at approximately 0.5%, with flexibility based on experience.
  • Professional growth and development opportunities.
  • High-impact work within a fast-paced AI technology environment.
Work Arrangement

On-Site – Palo Alto, CA

Employees are expected to work from the Palo Alto office 5 days per week.

Equal Opportunity Employer

Premier Global Links LLC is an equal opportunity employer. Qualified applicants are considered based on their skills, experience, education, and qualifications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
+3
Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500