Edge AI Performance Modeling Engineer — Cycle-Level

Quadric

Burlingame (CA)

On-site

USD 180,000 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid life Insurance
Equity
Paid Parental Leave
401(k) Retirement Plan
Flexible PTO
Catered lunch
Caltrain reimbursement
Downtown Burlingame office

Job summary

Quadric in Burlingame, California, is seeking an AI Performance Modeling Engineer to build cycle-level models of AI inference on the GPNPU architecture. You will work in Python to guide hardware decisions and optimize mapping of workloads from vision nets to LLMs.

This role welcomes all experience levels with mentorship, and offers a competitive base salary, equity, and a range of benefits; you’ll collaborate with a small, high-impact team in a growth-stage semiconductor company.

Qualifications

  • Python skills to write, validate and calibrate numerical models.
  • Solid grasp of memory hierarchies, bandwidth, latency trade-offs, and bottlenecks.
  • Experience writing clear technical studies with defended conclusions.
  • Core depth in NN inference operators or equivalent performance modeling.
  • BS/MS/PhD in CS, Electrical/Computer Engineering, or equivalent experience.

Responsibilities

  • Build analytical, cycle-level Python models of AI inference on next-gen GPNPU hardware.
  • Derive hardware lane bindings and model software pipeline overlaps.
  • Model tensor placement, tiling, memory residency, and data movement across memory tiers.
  • Calibrate models against ISA simulator and profiling traces; meet accuracy targets.
  • Write and defend technical studies that inform architecture and product decisions.
  • Balance single-stream latency and throughput across workloads.

Skills

Python
Quantitative modeling
Computer architecture
Technical writing
NN inference operators

Education

BS/MS/PhD in CS/EE/CE

Tools

CUDA
Triton
gem5
Timeloop
MAESTRO
Accel-Sim

Job description

Quadric in Burlingame, California, is seeking an AI Performance Modeling Engineer to build cycle-level models of AI inference on the GPNPU architecture. You will work in Python to guide hardware decisions and optimize mapping of workloads from vision nets to LLMs.

This role welcomes all experience levels with mentorship, and offers a competitive base salary, equity, and a range of benefits; you’ll collaborate with a small, high-impact team in a growth-stage semiconductor company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Performance Modeling Engineer
AI Performance Modeling Engineer

Quadric • Burlingame (CA)

On-site
USD 180,000 - 225,000
Medical, dental, and vision insurance
Company-paid life Insurance
Equity
+6
AI Inference Engineer
AI Inference Engineer

Quadric • Burlingame (CA)

Hybrid
USD 110,000 - 270,000
Competitive salary and meaningful-equy
Medical dental vision plans from day 1
401(k) retirement plan
+5
Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Quadric • Burlingame (CA)

Hybrid
USD 110,000 - 270,000
Competitive salary and meaningful-equy
Medical dental vision plans from day 1
401(k) retirement plan
+5
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
AI Inference Engineer
AI Inference Engineer

Quadric Inc. • Burlingame (CA)

On-site
USD 180,000 - 260,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
AI Inference Engineer
AI Inference Engineer

quadric.io, Inc • Burlingame (CA)

On-site
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

Hybrid
USD 266,000 - 445,000
Lead AI Inference Performance Architect
Lead AI Inference Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Principal Performance Modeling Engineer
Principal Performance Modeling Engineer

Oho Group • San Francisco (CA)

On-site
USD 190,000 - 280,000