Principal Performance Modeling Architect

Oxmiq Labs

Campbell (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OXMIQ is seeking a Principal Performance Modeling Architect in Campbell, CA to own and advance OxSol, the system-solution performance modeling platform, extending its reach to training workloads and multimodal models. You’ll work hands-on to define direction, push for extensibility, and anchor architecture choices with simulation-backed data.

You will read simulation outputs, ensure code quality, and collaborate with engineering leadership to drive high-leverage decisions before capital

Qualifications

  • 5–8 years of experience in performance modeling, computer architecture, or systems performance engineering with a shipped body of work.
  • Hands-on experience building first-principles/performance models for AI or HPC systems.
  • Solid understanding of both inference and training workloads and how they stress compute, memory, and interconnect.
  • Proficiency in Python for production-quality codebase development with tests and maintainable abstractions.
  • Track record of influencing architecture decisions and a tinkering mindset.

Responsibilities

  • Own and extend the system-solution performance modeling platform end to end (Speed-of-Light) across heterogeneous hardware and cluster topologies.
  • Extend modeling to training workloads and multimodal architectures.
  • Influence silicon and system architecture decisions with defensible, simulation-backed numbers.
  • Make the platform extensible and agent-driven; expose analysis capabilities via interfaces.
  • Validate simulation results against observed workloads and maintain model fidelity.
  • Set the platform's technical direction in line with executive vision and provide engineering guidance.

Skills

Python
Performance modeling
Computer architecture
AI accelerators
Data sheet analysis
Strong communication
Code quality
AI-assisted development

Education

BS/MS/PhD in CS/CE/EE

Tools

vLLM
TensorRT-LLM
Megatron
SGLang
PyTorch

Job description

OXMIQ designs GPU and AI silicon for large-scale model inference and training, and is building the system software and analysis platforms that prove out those designs before they reach customers. We are a startup driving innovation across the full stack — from atoms to agents — and we move fast on the strength of our own tools.

The Role

The Principal Performance Modeling Architect owns OXMIQ's system-solution performance analysis — the modeling platform (OxSol) and its Speed-of-Light and OxFabric analysis capabilities that convert a proposed silicon and cluster architecture into simulation-backed numbers: performance, energy, and cost-per-token across heterogeneous accelerator topologies. These numbers feed our architecture decisions, our pricing, and our customer proposals directly, and they are produced before we commit capital — which makes this one of the highest-leverage technical roles at the company.

This is a hands-on individual-contributor role. You own the platform's technical direction in partnership with engineering leadership, and you write the hard parts yourself. The platform models LLM inference today; your mandate is to extend its reach to new workload classes and make the engine more extensible as it grows.

We are looking for someone who has seen it and done it — a senior engineer with a tinkering mindset who is energized by being the person who can answer "is this architecture actually optimal, and what does a token cost on it?" and would rather prove it with a defensible model than hand-wave it. You read simulation output critically, reason about whether an outcome is physically sensible, hold a high bar for code quality, and treat AI-assisted development as a standard part of how you work.

Key Responsibilities
  • Own and extend OXMIQ's system-solution performance modeling platform end to end — Speed-of-Light modeling of performance, energy, and cost-per-token across heterogeneous accelerator hardware and cluster topologies.
  • Extend the platform beyond LLM inference to training workloads and vision-transformer / multimodal models, defining the modeling approach for each new workload class.
  • Influence silicon and system architecture decisions by turning proposed designs into defensible, simulation-backed numbers ahead of capital commitment.
  • Make the platform more extensible and agent-driven, including exposing analysis capabilities through the OxCapsule interface.
  • Validate and reason about simulation results — correlating model outputs against the observed behavior of real workload frameworks (e.g., vLLM, SGLang, training stacks) and continuously tightening model fidelity.
  • Set and own the platform's technical direction in line with the executive vision, uphold engineering and code-quality standards, and provide technical guidance to interns supporting the work.
  • Serve as a technical point of contact in selected proposal and partner engagements.
Required Qualifications
  • 5–8 years of relevant experience in performance modeling, computer architecture, or systems performance engineering, with a demonstrable, shipped body of work.
  • Hands-on experience building analytical / first-principles / roofline performance models of AI or HPC systems — deriving models, not only running benchmarks.
  • Solid understanding of both inference and training workloads and how they stress compute, memory bandwidth, and interconnect.
  • Strong working knowledge of AI accelerator and cluster hardware: GPUs/XPUs, HBM and memory hierarchies, scale-up/scale-out fabrics, and how parallelism strategies (tensor/pipeline/expert/data) map onto them.
  • Proficiency in Python for building and maintaining a production-quality codebase — clean abstractions, tests, and the discipline to keep a large platform maintainable.
  • A track record of influencing architecture decisions and a tinkering mindset — comfortable getting into datasheets, specs, and code to find out why a number is what it is.
  • An AI-first mindset: fluent use of AI-assisted development workflows (Claude Code or equivalent), and the judgment to review and reason critically about simulation results, outcomes, and code quality.
  • Strong written and verbal communication; able to defend a methodology and explain a surprising result to a non-specialist.
Preferred Qualifications
  • Hands-on experience operating and modifying modern workload frameworks — vLLM, SGLang, TensorRT-LLM, or comparable — at the level of understanding their internals and observed performance.
  • Experience modeling or optimizing training stacks (e.g., FSDP, Megatron-style parallelism) and vision-transformer / multimodal architectures.
  • Familiarity with LLM serving and efficiency techniques: continuous batching, KV-cache management, disaggregated prefill/decode, and quantization.
  • Background in silicon economics: reading vendor datasheets, die-yield and cost modeling, and tech-node tradeoffs.
  • Experience with agent / MCP-based tooling or exposing analysis engines through chat-native or programmatic interfaces.
  • Open-source contributions to performance, modeling, or inference-serving projects.
  • Experience where a model you built directly informed pricing or capital decisions.
Education
  • BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field — or equivalent practical experience.

This is a hands-on role reporting directly to the VP of System Architecture and working closely with OXMIQ's silicon, runtime, and systems teams to execute on the technical vision. You will own the technical direction of the platform; you may be supported by one or two interns, but this is not a people-management role. AI-assisted development tools (Claude Code or equivalent) are a standard part of engineering practice at OXMIQ and are expected in daily work.

OXMIQ offers a competitive compensation package, including base salary, equity participation, comprehensive medical, dental, and vision coverage, and the opportunity to contribute to foundational silicon and software technology.

OXMIQ is an equal opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Architect | AI Platform
Principal Architect | AI Platform

Oxmiq Labs • San Francisco (CA)

On-site
USD 250,000 - 400,000
Equity participation
Medical coverage
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Seattle (WA)

On-site
USD 342,000 - 555,000
Performance Modeling Lead
Performance Modeling Lead

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
Principal Performance Modeling Engineer
Principal Performance Modeling Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI System Performance Architect
Senior AI System Performance Architect

Oxmiq Labs • Campbell (CA)

On-site
USD 180,000 - 240,000
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • Los Angeles (CA)

Hybrid
USD 100,000 - 150,000
Relocation assistance
Hybrid work model
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • San Francisco (CA)

Hybrid
USD 266,000 - 445,000