Senior AI Inference Modeling Architect

Neurophos

San Mateo (CA)

On-site

USD 170,000 - 250,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

100% health plan premiums
Unlimited PTO
401(k) matching and stock options
Voluntary benefits
Personalized Benefits

Job summary

Neurophos is hiring a performance engineer to own benchmarking and energy metrics for their optical inference accelerator. You will manage measurement methodology, run end-to-end tests, and compare against competing GPUs and accelerators.

The role emphasizes reproducible results, harness development, and codifying workloads for consistent evaluation. The ideal candidate has 5+ years in GPU performance/mechanics, strong Python and Linux skills, and hands-on experience with Nsight tools, Hugging

Qualifications

  • Degree in a relevant engineering or CS field is required.
  • 5+ years in GPU performance engineering, benchmarking, or ML system measurement.
  • Experience turning papers or models into runnable benchmarks.
  • Hands-on roofline/limiter analysis or analytical modeling.
  • Proficient with Nsight tools and GPU profiling for HBM vs compute.
  • Familiar with Hugging Face/vLLM/TensorRT-LLM stacks and MoE concepts.
  • Python for harnesses and plots; comfortable in Linux; cloud GPU ops on major clouds.

Responsibilities

  • Own performance and energy metrics across fidelities and hardware.
  • Produce workload benchmarks across multiple fidelity levels and align datasets.
  • Keep workload definitions constant across models, sequences, and batching.
  • Integrate inference workloads from external stacks like Hugging Face and Triton.
  • Measure GPUs end-to-end, manage cloud or lab accounts, and run recipes.
  • Report TTFT, ITL, tokens/second, and energy per token with instrumentation.
  • Document discrepancies with RTL simulations and performance models.
  • Maintain an internal benchmark suite with clear separation for external claims.

Skills

GPU performance engineering
Benchmark harnesses
Roofline analysis
Limiter analysis
Performance modeling
Python scripting
Linux
Cloud GPU operations
HPC benchmarking

Education

BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience

Tools

NVIDIA Nsight Systems
NVIDIA Nsight Compute
TensorRT-LLM
PyTorch
CUDA

Job description

Neurophos is hiring a performance engineer to own benchmarking and energy metrics for their optical inference accelerator. You will manage measurement methodology, run end-to-end tests, and compare against competing GPUs and accelerators.

The role emphasizes reproducible results, harness development, and codifying workloads for consistent evaluation. The ideal candidate has 5+ years in GPU performance/mechanics, strong Python and Linux skills, and hands-on experience with Nsight tools, Hugging

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Benchmark Architect
AI Inference Benchmark Architect

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Lead Inference Performance Architect
Lead Inference Performance Architect

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
Modeling Architect — Optical AI Inference Accelerator
Modeling Architect — Optical AI Inference Accelerator

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 230,000
Health coverage
HSA contributions
Unlimited PTO
+3
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos • San Mateo (CA)

On-site
USD 170,000 - 250,000
100% health plan premiums
Unlimited PTO
401(k) matching and stock options
+2
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
GPU Systems Performance Engineer for AI Inference
GPU Systems Performance Engineer for AI Inference

Yoh Services LLC • California (MO)

On-site
USD 250,000 - 300,000
Medical benefits
Dental & Vision
401K Retirement
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2