AI Inference Benchmark Architect

Neurophos

Austin, Sunnyvale (TX, CA)

On-site

USD 150,000 - 210,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
Unlimited PTO
401k matching
Stock options
Voluntary benefits
Personalized benefits

Job summary

Neurophos is seeking a Performance Engineer to own benchmarking numbers for the T100 optical inference accelerator. You will define measurement methodology and generate end-to-end results across workloads, with full harness, plots, logs, and assumptions for reproducibility.

You will compare RTL models against measured GPUs, manage cloud or lab environments, and report time-to-first-token, latency, and energy metrics while maintaining a benchmark suite and governance for external claims.

Qualifications

  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • 5+ years in GPU performance engineering, accelerator benchmarking, HPC performance, or ML systems measurement.
  • Experience building or operating benchmark harnesses with real GPUs or accelerators.
  • Hands-on with roofline or limiter analysis and analytical performance modeling.
  • GPU performance analysis with Nsight tools for HBM vs compute and batching.
  • Familiar with LLM inference stacks and MoE concepts.
  • Proficient in Python for harnesses, parsing, and plots; Linux comfort.
  • Cloud GPU ops on AWS/GCP/Azure, containers, drivers, and quotas.

Responsibilities

  • Own performance and energy metrics across fidelities and hardware.
  • Produce numbers for workloads across fidelities: roofline, models, RTL sim, real GPUs.
  • Keep workload definitions constant across fidelities (model, sequence length, batch, precision).
  • Bring up inference workloads from Hugging Face, PyTorch, papers, and vendor stacks.
  • Measure competing GPUs end-to-end; manage cloud or lab accounts, images, drivers, runs.
  • Report TTFT, ITL, tokens/sec, tokens/sec per watt, and energy metrics with instrumentation.
  • Document discrepancies between RTL, models, and measured results with configs/logs.

Skills

GPU performance
Benchmark harness
Roofline analysis
Nsight Systems
LLM inference stacks
Python scripting
Linux proficiency
Cloud GPU ops

Education

BS/MS in CE/EE/CS

Tools

Hugging Face
vLLM
SGLang
TensorRT-LLM
CUDA
Triton
PyTorch

Job description

Neurophos is seeking a Performance Engineer to own benchmarking numbers for the T100 optical inference accelerator. You will define measurement methodology and generate end-to-end results across workloads, with full harness, plots, logs, and assumptions for reproducibility.

You will compare RTL models against measured GPUs, manage cloud or lab environments, and report time-to-first-token, latency, and energy metrics while maintaining a benchmark suite and governance for external claims.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Modeling Architect
Senior AI Inference Modeling Architect

Neurophos • San Mateo (CA)

On-site
USD 170,000 - 250,000
100% health plan premiums
Unlimited PTO
401(k) matching and stock options
+2
Lead Inference Performance Architect
Lead Inference Performance Architect

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2
Hardware Benchmarking Engineer – AI Accelerators
Hardware Benchmarking Engineer – AI Accelerators

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Modeling Architect — Optical AI Inference Accelerator
Modeling Architect — Optical AI Inference Accelerator

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 230,000
Health coverage
HSA contributions
Unlimited PTO
+3
GPU Benchmark Engineer for AI Cloud Infra
GPU Benchmark Engineer for AI Cloud Infra

Nebius • United States

Remote
USD 150,000 - 190,000
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance