System Performance Modeling Engineer/Architect (NPU)

Bitdeer Group

Singapore

On-site

SGD 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer Group in Singapore is looking for a performance modeling expert to help analyze and optimize chip performance through modeling techniques. You will be responsible for building and maintaining cycle-approximate models that guide architecture decisions concerning tile count, NoC bandwidth, and local memory.

Your expertise in C++ and Python, alongside strong collaboration skills, will be essential. The role offers an opportunity to work closely with various teams and make impactful decisions in the performance of advanced chips.

Qualifications

  • 5+ years in performance modeling or architecture for SoCs, accelerators, GPUs, or CPUs.
  • Strong skills in C++ and Python; experience with SystemC or comparable frameworks is needed.
  • Demonstrated ability to model subsystems and translate results into architecture decisions.

Responsibilities

  • Build and maintain performance models of the tile array, NoC, memory subsystem, and IO.
  • Characterize LLM and CNN inference workloads; produce relevant analyses.
  • Model NoC contention and memory bandwidth based on realistic traffic.

Skills

Performance modeling / architecture
C++
Python
SystemC

Job description

About Bitdeer

Bitdeer is a world‑leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry‑leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About the role

You will build the models that tell us how fast the chip will actually run before any RTL exists — and keep telling us as the design evolves. Working at full‑chip scope, you will model the NPU tile array, the self‑developed NoC, the 3D‑DRAM bandwidth path and end‑to‑end inference of real workloads (LLM and CNN). Your models drive the biggest architecture decisions: how many tiles, how much local memory, what NoC bandwidth and what memory QoS are needed to hit our performance targets within the power and thermal envelope. You will be the quantitative conscience of the architecture team and a close partner to the NPU, memory and compiler groups.

Key responsibilities
  • Full‑chip performance models – build and maintain cycle‑approximate / analytical models of the tile array, NoC, memory subsystem and IO in C++/SystemC and Python.
  • Workload & roofline analysis – characterize LLM and CNN inference workloads; produce roofline and bottleneck analyses that map workloads onto the tile array and memory hierarchy.
  • Traffic & contention modeling – model NoC contention, memory bandwidth and QoS under realistic traffic; generate and parameterize traffic scenarios for architecture studies.
  • Architecture trade‑offs – run sweeps (tile count, local‑memory size, NoC topology/bandwidth, memory channels/QoS) and turn results into clear recommendations that feed the architecture freeze.
  • Compiler co‑design – partner with the AI compiler team on dataflow mapping, tiling and scheduling so modeled performance reflects what the compiler can actually achieve.
  • Correlation & validation – correlate model predictions against RTL, emulation/FPGA and (later) silicon; quantify and close the modeling‑to‑reality gap.
  • Golden‑reference hooks – provide performance references and golden‑model hooks used by the NPU performance and verification teams for sign‑off.
  • Reporting – maintain dashboards and trade‑off decks that make performance risk visible to architecture and program leadership.
Required qualifications
  • 5+ years in performance modeling / architecture for SoCs, accelerators, GPUs or CPUs.
  • Strong C++ and Python; experience with SystemC or a comparable cycle‑approximate modeling framework.
  • Solid grasp of computer architecture fundamentals: memory hierarchy, bandwidth/latency trade‑offs, interconnect/NoC behavior and roofline reasoning.
  • Demonstrated ability to model a subsystem and turn results into architecture decisions, not just numbers.
  • Comfortable working from incomplete specs and collaborating across architecture, design and software teams.
Preferred / differentiators
  • Experience modeling AI/ML inference workloads (LLM, CNN, transformers) and quantization/precision effects on performance.
  • Familiarity with NoC/interconnect modeling, memory‑controller behavior or 3D‑DRAM/HBM bandwidth analysis.
  • Exposure to compiler/dataflow co‑design for accelerators.
  • Experience correlating models against RTL, emulation or silicon.
  • Background with tile‑based or dataflow architectures.
What success looks like (first 6–12 months)
  • A trusted full‑chip performance model that the team uses to make tile‑count, NoC and memory decisions at architecture freeze.
  • Clear roofline/bottleneck analyses for the target LLM and CNN workloads.
  • Quantified, agreed performance targets and a documented modeling‑v‑reality correlation plan.
  • Performance trade‑off reporting that program and architecture leadership rely on.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Performance Modeling Engineer/Architect
System Performance Modeling Engineer/Architect

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Full-Chip Performance Modeling Architect
Full-Chip Performance Modeling Architect

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Full-Chip Performance Modeling Architect (AI/LLM)
Full-Chip Performance Modeling Architect (AI/LLM)

Bitdeer Group • Singapore

On-site
SGD 80,000 - 120,000
Senior LLM Inference Performance & Evaluation Engineer
Senior LLM Inference Performance & Evaluation Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior LLM Inference Performance & Evaluation Engineer
Senior LLM Inference Performance & Evaluation Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Senior Inference Runtime Engineer
Senior Inference Runtime Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
Senior Inference Runtime Engineer
Senior Inference Runtime Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
Senior/Staff Front-End Design Engineer (RISC-V)
Senior/Staff Front-End Design Engineer (RISC-V)

Bitdeer Group • Singapore

On-site
SGD 70,000 - 90,000
Attractive welfare benefits
Training and mentoring opportunities
Senior/Staff Front-End Design Engineer (RTL Design - RiscV)
Senior/Staff Front-End Design Engineer (RTL Design - RiscV)

United States Digital Space LLC • Singapore

On-site
SGD 180,000 - 280,000
Senior Physical Design Engineer
Senior Physical Design Engineer

Bitdeer Group • Singapore

On-site
SGD 80,000 - 120,000
Attractive welfare benefits
Training and mentoring opportunities