Senior AI Inference Engineer — GPU & Edge Optimization

Jobtailor

California (MO)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a seasoned AI software engineer to advance AI inference on RTX and DGX platforms, focusing on optimizing performance of AI models and inference runtimes on GPU architectures. You will collaborate with NVIDIA teams and industry partners, driving innovations in open and closed source technologies with emphasis on system-level support.

Required are 5+ years in AI inference pipelines, strong C++ skills, and expertise with ML/DL frameworks such as ONNX RT, PyTorch, TensorRT,

Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or a related field (or equivalent experience).
  • Excellent C++ programming and debugging skills with a strong understanding of data structures and algorithms.
  • 5+ years of experience with proficiency in AI inferencing pipelines and applications using ML/DL frameworks, including ONNX RT, PyTorch, Tensor RT, llama.cpp and vLLM.
  • Strong analytical and problem-solving abilities, with the ability to multitask effectively in a dynamic environment.
  • Outstanding written and oral communication skills enabling effective collaboration with management and engineering teams.

Responsibilities

  • Partner with NVIDIA software, research, architecture, and product teams to align strategies and technical needs for AI ecosystem on RTX and DGX PCs.
  • Collaborate closely with industry partners to advance AI across critical domains by driving innovations in open and closed source technologies with emphasis on system-level support.
  • Improve performance on current and next-generation GPU architectures by analyzing and optimizing AI models, data processing pipelines, and inference runtime features.
  • Identify, evaluate, and implement compute and memory optimization techniques—quantization, distillation, and pruning—for large AI models; fine-tune and compress models for edge devices.

Skills

C++ Programming
AI Inference Pipelines
ML/DL Frameworks
Data Structures
Algorithms
Communication Skills
Analytical Thinking

Education

Bachelor/Master/PhD in CS/SE/Math

Tools

ONNX RT
PyTorch
TensorRT
llama.cpp
vLLM

Job description

Jobtailor is seeking a seasoned AI software engineer to advance AI inference on RTX and DGX platforms, focusing on optimizing performance of AI models and inference runtimes on GPU architectures. You will collaborate with NVIDIA teams and industry partners, driving innovations in open and closed source technologies with emphasis on system-level support.

Required are 5+ years in AI inference pipelines, strong C++ skills, and expertise with ML/DL frameworks such as ONNX RT, PyTorch, TensorRT,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer – Edge GPU & Local Inference
Senior AI Systems Engineer – Edge GPU & Local Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior Systems Software Engineer - Local AI on GPUs
Senior Systems Software Engineer - Local AI on GPUs

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Senior AI Performance Engineer – GPU & DL
Senior AI Performance Engineer – GPU & DL

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
AI Tooling Architect for Efficient Training & Inference
AI Tooling Architect for Efficient Training & Inference

Jobtailor • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits