Acceleration Kernel Developer Lead

Tenstorrent

Warszawa

Hybrid

PLN 240,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Highly competitive compensation package
Benefits included
Equal opportunity employer

Job summary

A leading AI technology firm located in Poland seeks a Software Engineer specializing in Kernel Development and Optimization. Candidates will engage in performance-critical kernel design, optimizing GPU-style operations while collaborating with various teams for integration. The ideal candidate will have a strong background in C++ systems engineering and a data-driven mindset for performance optimization. This position offers a competitive compensation package and supports a hybrid work model from Warsaw or Gdansk.

Qualifications

  • Strong C++ systems engineer with experience in low-level software.
  • Data-driven with profiling experience.
  • Effective at debugging kernel-level issues.
  • Ability to debug complex runtime/kernel-level issues in large codebases.
  • Structured problem-solving to turn performance problems into measurable experiments.

Responsibilities

  • Design and optimize GPU-style kernels.
  • Identify performance bottlenecks and improve throughput.
  • Collaborate with teams to integrate kernels.
  • Develop micro-benchmarks, regression tests, and tooling to ensure correctness and sustained gains.
  • Collaborate with compiler, runtime, ML, and hardware teams to integrate kernels into production systems.

Skills

C++ systems engineering
Concurrency reasoning
Profiling and benchmarking
Debugging complex issues
Performance optimization
Kernel development

Job description

Software Engineer, Kernel Development and Optimization

Tenstorrent is leading the industry on cutting‑edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC‑V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.

Tenstorrent is building next‑generation AI compute. The Kernel Development and Optimization team develops the performance‑critical kernels that unlock the full capability of our hardware across ML and HPC workloads.

This role is hybrid based out of Warsaw or Gdansk, Poland.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.

Who You Are
  • A strong C++ systems engineer with experience writing performance‑critical or low‑level software.
  • Comfortable reasoning about concurrency, synchronization, latency hiding, and compute versus memory trade‑offs.
  • Data‑driven in your approach, using profiling and benchmarking results to guide optimization decisions.
  • Effective at debugging complex runtime or kernel‑level issues in large codebases.
  • Structured thinker who can break down ambiguous performance problems into measurable experiments.
What We Need
  • Engineers who can design, implement, and optimize GPU‑style kernels such as matrix multiplication, attention primitives, and data‑movement operations.
  • Clear ownership of performance, from identifying bottlenecks to delivering measurable throughput improvements.
  • Contribution to host‑side orchestration code and parallelization strategies.
  • Development of micro‑benchmarks, regression tests, and tooling to ensure correctness and sustained performance gains.
  • Close collaboration with compiler, runtime, ML, and hardware teams to integrate kernels into production systems.
What You Will Learn
  • The execution model, memory architecture, and performance characteristics of Tenstorrent AI hardware.
  • How to write and optimize accelerator kernels outside traditional CUDA‑first ecosystems.
  • Practical AI‑assisted and agentic workflows for kernel generation, debugging, and optimization.
  • How to translate performance intuition into rigorous, reproducible engineering results.
  • How low‑level kernels, compilers, runtime systems, and hardware co‑evolve in modern AI platforms.

Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.

This offer of employment is contingent upon the applicant being eligible to access U.S. export‑controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.

As set forth in Tenstorrent’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

We do not discriminate and are committed to equal opportunity for all candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Lead, Debug
Site Reliability Engineer Lead, Debug

Tenstorrent • Warszawa, Gdańsk

On-site
PLN 320,000 - 520,000
Site Reliability Engineer Lead, Debug
Site Reliability Engineer Lead, Debug

Tenstorrent • Województwo pomorskie

Hybrid
PLN 200,000 - 420,000
Infrastructure and Platform Development Engineer
Infrastructure and Platform Development Engineer

Tenstorrent • Warszawa

Hybrid
PLN 371,000 - 1,859,000
Highly competitive compensation package
Benefits included
C++ Machine Learning Engineer, Models Training
C++ Machine Learning Engineer, Models Training

Tenstorrent • Warszawa

Hybrid
PLN 180,000 - 280,000
Highly competitive compensation
Benefits package
Equal opportunity employer
C++ Machine Learning Engineer, Models Training
C++ Machine Learning Engineer, Models Training

Tenstorrent • Województwo pomorskie

Remote
PLN 180,000 - 280,000
Highly competitive compensation package
Benefits
Equal opportunity employer
Senior C++ Software Engineer with CUDA/GPU/TPU
Senior C++ Software Engineer with CUDA/GPU/TPU

EPAM Systems • Poland

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote work within Poland
Occasional international travel
+3
Senior Software Developer, AI Networking
Senior Software Developer, AI Networking

NVIDIA • Warszawa

On-site
PLN 292,000 - 650,000
Staff Software Engineer - ML Kernels & Runtime Gdańsk, Pomeranian Voivodeship, Poland
Staff Software Engineer - ML Kernels & Runtime Gdańsk, Pomeranian Voivodeship, Poland

graphcore • Województwo pomorskie

On-site
PLN 350,000 - 474,000
Annual leave
Medical and dental health plans
Gym card
+1
Acceleration Kernel Lead — High-Performance ML Optimization
Acceleration Kernel Lead — High-Performance ML Optimization

Tenstorrent Inc. • Gdańsk, Warszawa

Hybrid
PLN 260,000 - 380,000
Kernel Performance Engineer — C++ & GPU-Style Kernels
Kernel Performance Engineer — C++ & GPU-Style Kernels

Tenstorrent • Warszawa

Hybrid
PLN 240,000 - 360,000
Highly competitive compensation package
Benefits included
Equal opportunity employer