AI Infrastructure Engineer

1600 NIO USA, Inc.

San Jose (CA)

On-site

USD 163,500 - 212,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health plans with $0 employee coverage
401(k) with brokerage link option
Onsite gym and free snacks

Job summary

1600 NIO USA, Inc. is looking for a Senior AI Inference Infrastructure Software Engineer in San Jose, California. The role involves designing and optimizing high-performance inference systems for Large Language and Vision-Language Models across various environments.

Candidates must have over 5 years of experience, strong programming skills, and expertise in performance optimization techniques. Benefits include comprehensive health plans, flexible spending accounts, and a competitive salary range.

Qualifications

  • 5+ years of hands-on software development experience building and optimizing AI inference systems.
  • Strong expertise in performance engineering for inference systems.
  • Proficiency in C/C++ programming for high-performance software.

Responsibilities

  • Design and implement scalable inference systems for LLMs and VLMs.
  • Develop and optimize custom kernels for various hardware.
  • Ensure tight hardware-software integration and optimal performance.

Skills

AI Inference Systems
Performance Engineering
C/C++ Programming
GPU/NPU Programming
Deep Learning Frameworks (PyTorch/TensorFlow)

Education

BS/MS in Computer Science or related field

Tools

CUDA
TensorFlow
PyTorch

Job description

About the Position

We are seeking a senior AI Inference Infrastructure Software Engineer with extensive experience designing, building, and optimizing high‑performance, scalable inference systems for Large Language Models (LLMs) and Vision‑Language Models (VLMs). The role focuses on delivering production‑grade software that powers real‑world applications of LLMs/VLMs across cloud, edge, and hybrid edge‑cloud environments.

Roles and Responsibilities
  • Design and implement scalable inference systems for LLMs and VLMs across cloud, edge, and hybrid platforms.
  • Develop and optimize custom kernels and operators for GPU, NPU, DSP, and other accelerator hardware to improve throughput, latency, and memory efficiency.
  • Integrate advanced optimization techniques (KV‑cache management, tensor/model parallelism, quantization, memory‑efficient execution) into production inference pipelines.
  • Partner with system and hardware teams to ensure tight hardware‑software integration and optimal performance across diverse compute environments.
  • Translate architectural requirements into robust, maintainable, production‑ready software that meets performance, safety, and reliability standards.
  • Define and drive the evolution roadmap for LLM/VLM inference in the AIOS stack, ensuring scalability and adaptability to new workloads.
  • Stay ahead of industry trends and competitor solutions, applying best practices from AI and large‑scale systems engineering.
Qualifications
  • 5+ years of hands‑on software development experience building and optimizing AI inference systems at scale.
  • Direct experience with LLM/VLM model internals, including Transformer‑based architectures, inference bottlenecks, and optimization techniques.
  • Strong expertise in performance engineering: kernel development, parallelism strategies, memory optimization, and distributed inference systems.
  • Proficiency with GPU/NPU programming (CUDA or vendor‑specific SDKs), compiler toolchains, and deep learning frameworks (PyTorch or TensorFlow).
  • Strong programming skills in C/C++, with a track record of delivering high‑performance, production‑grade software.
  • Solid foundation in computer architecture, systems programming, and embedded systems.
  • BS/MS in Computer Science, Computer Engineering, or related field.
  • Excellent communication and collaboration skills, with the ability to work across cross‑functional teams.
Preferred Qualifications
  • Master’s or PhD in Computer Science, Electrical/Computer Engineering, or related field.
  • 5+ years industry experience building inference serving systems for large models (batching, scheduling, caching, load balancing).
  • Expertise in hardware‑aware model optimization (kernel fusion, mixed precision, quantization, pruning).
  • Familiarity with edge and embedded AI, including real‑time constraints and limited‑resource optimization.
  • Contributions to widely used AI frameworks, libraries, or performance‑critical software (open source or proprietary).
Compensation

US base salary range: $163,500.00 – $212,400.00. Individual pay is determined by location and other factors.

Benefits
  • Health: Anthem Blue Cross, HSA, and Kaiser HMO medical plans with $0 employee coverage.
  • Dental (including orthodontic) and vision plans with $0 employee coverage.
  • Company‑paid HSA contribution when enrolled in the high‑deductible Anthem plan.
  • Healthcare and dependent care flexible spending accounts (FSA).
  • 401(k) with brokerage link option.
  • Life insurance, AD&D, short‑term and long‑term disability insurance.
  • Employee assistance program, sick and vacation time, 13 paid holidays per year.
  • Paid parental leave: first 8 weeks at full pay (eligible after 90 days).
  • Paid disability leave: first 6 weeks at full pay (eligible after 90 days).
  • Voluntary benefits: life, AD&D, pet insurance, commuter benefits, mobile cell phone credit.
  • Free lunch and snacks, onsite gym, employee discounts and perks program.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
AI Software Engineer
AI Software Engineer

ElastixAI INC. • Seattle (WA)

Hybrid
USD 150,000 - 210,000
Comprehensive medical, dental, and vision coverage
Life insurance
Flexible Time Off (FTO)
+7
AI Technical Lead
AI Technical Lead

1600 NIO USA, Inc. • San Jose (CA)

On-site
USD 192,100 - 249,600
Health insurance
Dental and vision plans
401(k) with brokerage link
+3
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Sr. AI Software Engineer (GPU/C++)
Sr. AI Software Engineer (GPU/C++)

KLA-Belgium • Milpitas (CA)

On-site
USD 166,000 - 284,000
Medical, dental, vision benefits
401(K) with company matching
Tuition reimbursement program
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave