Systems Performance Modeling Engineer

Tensordyne

Sunnyvale (CA)

On-site

USD 140,000 - 210,000

Full time

13 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Comprehensive benefits
Flexible spending options
Recognition programs

Job summary

Tensordyne in Sunnyvale, CA, seeks a Systems Performance Modeling Engineer to build models and tools predicting AI inference workloads on our systems from single accelerators to rack-scale deployments. This hands-on role involves writing simulator code, running experiments, and analyzing discrepancies between predictions and measurements.

You'll collaborate across architecture, silicon, hardware, and software teams to extend models of silicon, interconnects, and fabric, capture workload

Qualifications

  • Hands-on experience building performance models, simulators, or analytical tools for ML workloads.
  • Strong programming skills in C++ and Python with clean, testable code.
  • Experience comparing model predictions against real measurements and debugging divergences.
  • Understanding of distributed ML execution and parallelism strategies.

Responsibilities

  • Build and extend simulation-based performance models for multimodal AI inference at rack, pod, and cluster scale.
  • Model batching, KV-cache management and disaggregation effects on latency and throughput.
  • Create trace-capture and replay tooling for hardware configurations.
  • Run calibration experiments on Tensordyne hardware and compare to model predictions.
  • Provide design-space analyses and performance projections for decisions.

Skills

C++
Python
Distributed systems
Performance modeling
Debugging
Communication

Education

MS or higher in CS/CE/EE or related field

Tools

NCCL
Python testing
Simulation tools
Profiling ML workloads

Job description

Artificial intelligence (AI) is transforming our world. It can perform cognitive functions that previously only humans could do, such as perceiving interactions across different modalities and environments - with the ability to quickly learn and then solve complex problems. Tensordyne is an AI system solution company that builds very high-performance, low-power generative AI inference systems. Our mission, through the creation of custom silicon, hardware and software, is to enable multimodal Generative AI inference acceleration at scale, with safe, sustainable, high-performance systems for our hyperscaler and neocloud data center customers. We are at the leading edge of advancing the latest research and product improvements for generative Al inference solutions that will make Al even more advantageous for compelling new generative AI applications. Tensordyne is a well funded, fast-paced startup company with headquarters in both Sunnyvale, CA, and Munich, Germany. We also have many talented team members working remotely across North America and Europe. We take care of our people and their families with comprehensive benefits, competitive compensation, flexible spending options, and recognition programs, because building category-defining technology starts with a healthy, supported team. Come join us as we shape the future of multimodal generative artificial intelligence!

About the Role

We are looking for a Systems Performance Modeling Engineer to build the models and tools that predict how generative AI inference workloads perform on Tensordyne systems, from a single accelerator up through rack, pod, and cluster scale. This is a hands-on engineering role for someone who likes writing simulator code, running experiments, and digging into why a prediction and a measurement don't match.

Working closely with our architects and the silicon, hardware, networking, and software teams, you'll capture workload behavior, extend simulation and analytical models of our silicon, interconnect, and multi-hop fabrics, and validate them against real hardware. You'll be comfortable moving across the stack, from the model graph through collectives to the network fabric, to track down where performance is going.

What You'll Do
  • Implement and extend simulation-based performance models for multimodal generative AI inference at rack, pod, and cluster scale, covering compute, memory, collective communication, and network fabric.
  • Model how serving strategies (tensor, pipeline, and expert parallelism, prefill/decode disaggregation, batching, and KV-cache placement) interact with Tensordyne silicon and fabric topology, and measure the effect on latency, throughput, and cost per token.
  • Build trace-capture and replay tooling that records real execution from our inference runtime and replays it under hypothetical silicon, system, and network configurations.
  • Model collective communication on multi-hop scale-out fabrics, including implementing custom collective algorithms designed for our topology.
  • Run calibration experiments on Tensordyne hardware as systems come up, compare them against model predictions, and fix the sources of error.
  • Run design-space sweeps and write up clear analyses that architects and engineering teams use in ASIC, fabric, and system configuration decisions.
  • Produce performance projections that support product and customer discussions.
  • Keep the modeling codebase fast, tested, and reproducible so other engineers can run it themselves.
What We're Looking For
  • Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture.
  • Solid understanding of distributed ML execution, including parallelism strategies, collective communication (All-Reduce, All-Gather, All-to-All, etc.), and how they scale.
  • Working knowledge of system architecture across compute, memory, interconnect, and networking, and the ability to reason about bottlenecks between them.
  • Experience comparing model predictions against real measurements, and debugging where they diverge.
  • Strong programming skills in C++ and Python, with clean, testable, maintainable code.
  • Ability to take a loosely defined performance question, break it into experiments, and deliver results with minimal hand-holding.
  • Clear written and verbal communication, especially when presenting data and trade-offs to other engineers.
  • MS or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
Nice to Have
  • Familiarity with LLM inference serving: batching, KV-cache management, disaggregated prefill/decode, and latency/throughput trade-offs.
  • Experience modeling or benchmarking collective communication libraries (NCCL, RCCL, or similar) on real clusters.
  • Background in data center or HPC networking: topologies, RDMA/RoCE, and congestion behavior.
  • Experience profiling ML workloads on accelerators (GPUs, TPUs, or custom ASICs).
  • Exposure to hardware/software co-design or early-stage architecture evaluation.
  • Publications or open-source contributions in ML systems, architecture, or networking.
Tensordyne's culture was built on the following values
  • Put people first. We only succeed when our people succeed.
  • Ethics and integrity always; Being open, honest, and respectful of everyone.
  • Think Big. Be ambitious and have audacious goals of global scale.
  • Aim for excellence. Quality and excellence count in everything we do.
  • Own it and get it done. Results matter!
  • Make each person better together, than they would be as an individual.
  • Embrace each others’ differences, and embrace that there will be differences.

Tensordyne is an equal opportunity employer. We believe that a diverse team is better at tackling complex problems and coming up with innovative solutions. All qualified applicants will receive consideration for employment without regard to age, color, gender identity or expression, marital status, national origin, disability, protected veteran status, race, religion, pregnancy, sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

System Software Engineer, Networking
System Software Engineer, Networking

Tensordyne • San Jose (CA)

On-site
USD 170,000 - 210,000
System Software Engineer, Networking
System Software Engineer, Networking

Tensordyne • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Comprehensive benefits
Flexible spending options
Recognition program
Sr AI Product Director
Sr AI Product Director

Tensordyne • United States

Hybrid
USD 180,000 - 320,000
Senior Al Product Director
Senior Al Product Director

Tensordyne • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
Sr. Staff ASIC Design Engineer
Sr. Staff ASIC Design Engineer

Tensordyne • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
Meals
Snacks & drinks
Unlimited PTO
+1
Sr. ASIC EDA Workflow Engineer
Sr. ASIC EDA Workflow Engineer

Tensordyne • Sunnyvale (CA)

On-site
USD 110,000 - 140,000
Comprehensive benefits
Flexible spending options
Recognition programs
Mid Level ASIC Design Engineer
Mid Level ASIC Design Engineer

Tensordyne • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Flexible spending
Bonusly awards
Health benefits
+1
Sr. Staff ASIC Verification Engineer
Sr. Staff ASIC Verification Engineer

Tensordyne • United States

Remote
USD 180,000 - 320,000
Flexible spending
Bonusly awards
Meals & snacks
Sr. ASIC Verification Engineer
Sr. ASIC Verification Engineer

Tensordyne • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Meals, snacks, and drinks
Learning and development opportunities
Fun office environment
Principal ASIC Architect
Principal ASIC Architect

Tensordyne • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Competitive pay
Comprehensive benefits
Flexibility
+1