Lead Inference Performance Architect

Neurophos, Inc.

Sunnyvale, Northern (TX, KY)

Hybrid

USD 150,000 - 210,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
Unlimited PTO
401(k) matching
Stock options
Voluntary benefits
Personalized benefits

Job summary

Neurophos, Inc. is seeking a performance engineer to own benchmarking and energy metrics for the T100 optical inference accelerator.

You will define measurement methods, generate end-to-end results, and ensure reproducibility across models, fidelities, and hardware generations. The role requires 5+ years in GPU performance, with hands-on experience in roofline and analytical modeling, and the ability to work with cloud GPUs and ML stacks.

Qualifications

  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science or equivalent practical experience
  • 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement
  • Track record of building or operating benchmark harnesses that produced measured results on real GPUs or accelerators
  • Hands-on experience with roofline analysis, limiter analysis, or analytical performance modeling
  • Proficiency in Python for harnesses, parsing, and plots, and comfort working in Linux
  • Cloud GPU operations on AWS, GCP, or Azure, including containers, instance types, drivers, quotas, and cost

Responsibilities

  • Own the performance and energy metrics that architecture, product, and leadership rely on, across modeling fidelities and hardware
  • Produce numbers for the same workloads across fidelities: roofline, limiter models, architecture models, RTL simulation, and measured hardware
  • Keep workload definitions constant across fidelities (model, sequence length, batch, precision, prefill vs decode, tensor, parallelism)
  • Bring up inference workloads from Hugging Face, PyTorch, papers, and vendor stacks (vLLM, SGLang, TensorRT-LLM, Triton)
  • Measure competing GPUs and accelerators end-to-end, managing cloud or lab accounts, images, drivers, and run recipes
  • Report TTFT, ITL, tokens/s, tokens/s per watt, and energy using nv-smI, DCGM, or equivalents
  • Document discrepancies between RTL, model, and measured results, with configs/logs behind each
  • Maintain a benchmark suite separating internal vs external results; route external claims for approval

Skills

Python programming
GPU performance analysis
Benchmark harness development
Linux proficiency

Education

BS or MS in Computer Engineering / Electrical Engineering / Computer Science

Tools

NVIDIA Nsight Systems
Nsight Compute
Hugging Face
TensorRT-LLM

Job description

Neurophos, Inc. is seeking a performance engineer to own benchmarking and energy metrics for the T100 optical inference accelerator.

You will define measurement methods, generate end-to-end results, and ensure reproducibility across models, fidelities, and hardware generations. The role requires 5+ years in GPU performance, with hands-on experience in roofline and analytical modeling, and the ability to work with cloud GPUs and ML stacks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Modeling Architect
Senior AI Inference Modeling Architect

Neurophos • San Mateo (CA)

On-site
USD 170,000 - 250,000
100% health plan premiums
Unlimited PTO
401(k) matching and stock options
+2
AI Inference Benchmark Architect
AI Inference Benchmark Architect

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Modeling Architect — Optical AI Inference Accelerator
Modeling Architect — Optical AI Inference Accelerator

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 230,000
Health coverage
HSA contributions
Unlimited PTO
+3
Technical Lead: Inference Benchmarking & ML Infra
Technical Lead: Inference Benchmarking & ML Infra

NVIDIA • United States

On-site
USD 224,000 - 357,000
Equity and benefits
Comprehensive benefits package
Competitive salaries
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos • San Mateo (CA)

On-site
USD 170,000 - 250,000
100% health plan premiums
Unlimited PTO
401(k) matching and stock options
+2
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
Technical Lead: Inference Benchmarking & ML Infra
Technical Lead: Inference Benchmarking & ML Infra

NVIDIA • California (MO)

On-site
USD 224,000 - 357,000
Equity options
Comprehensive benefits package
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits