GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Anyone AI Inc.

Argentina

A distancia

ARS 125.965.000 - 251.931.000

A tiempo parcial

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote work

Descripción de la vacante

Anyone AI Inc. is seeking an experienced GPU Kernel Engineer for a remote, part-time, project-based engagement focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

You will optimize kernels across CUDA and Triton, assess numerical correctness, memory behavior, and performance, and provide actionable feedback to improve efficiency on target hardware. Candidates should have 3+ years of hands-on kernel work and be comfortable translating kernels

Formación

  • 3+ years of hands-on experience developing, debugging, or optimizing GPU kernels.
  • Strong understanding of GPU architecture, memory hierarchy, and numerical correctness.

Responsabilidades

  • Review GPU/accelerator kernel implementations for correctness and performance.
  • Benchmark and verify numerical tolerance against reference results.
  • Provide actionable feedback to improve kernel efficiency across CUDA/Triton and hardware targets.

Conocimientos

GPU kernel debugging
kernel optimization

Herramientas

CUDA
Triton
Nsight / profilers

Descripción del empleo

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

Kernel implementation and debugging

CUDA and Triton optimization

Hardware migration

Operator fusion

Performance profiling and benchmarking

Numerical correctness verification

Compilation and runtime debugging

Memory hierarchy optimization

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

Strong experience with at least two of the following:

CUDA

Triton

Strong understanding of GPU performance optimization

Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

Understanding of:

Memory bandwidth

Compute throughput

Shared memory

Register pressure

Memory coalescing

Bank conflicts

Strong understanding of floating-point numerical correctness and tolerance thresholds

Experience debugging kernel compilation and runtime issues

Ability to distinguish software defects, environment problems, and genuine optimization challenges

Candidates should have experience with several of the following types of work:

Writing kernels from technical specifications

Translating kernels between CUDA, Triton, or other frameworks

Migrating kernels across hardware platforms

Debugging incorrect kernel implementations

Fusing multiple operations into optimized kernels

Nice to Have

Experience across both NVIDIA GPU and custom accelerator ecosystems

Experience with AWS Trainium, TPU, JAX, or other accelerators

Familiarity with MLIR, XLA, or intermediate representation lowering

Contributions to GPU or ML kernel libraries

Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

Reviewing GPU and accelerator kernel implementations for correctness

Comparing outputs against reference implementations

Evaluating numerical tolerance thresholds

Reviewing kernel benchmarks and determining whether comparisons are fair

Identifying performance bottlenecks and optimization opportunities

Assessing whether performance targets are realistic given hardware limits

Reviewing kernel translations and hardware migrations

Identifying compilation, driver, memory, shape, and runtime issues

Determining whether technical tasks are genuinely difficult or incorrectly configured

Providing clear, actionable technical feedback

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote GPU Kernel Engineer CUDA/Triton Performance Debugging
Remote GPU Kernel Engineer CUDA/Triton Performance Debugging

Anyone AI Inc. • Argentina

A distancia
ARS 125.965.000 - 251.931.000
Remote work
AWS Trainium / NKI Kernel Expert
AWS Trainium / NKI Kernel Expert

Anyone AI Inc. • Argentina

A distancia
ARS 167.954.000 - 314.914.000
Remote Part-Time NKI Kernel Engineer for AWS Trainium
Remote Part-Time NKI Kernel Engineer for AWS Trainium

Anyone AI Inc. • Argentina

A distancia
ARS 167.954.000 - 314.914.000
GPU & ML Infrastructure Engineer
GPU & ML Infrastructure Engineer

Svitla Systems, Inc. • Argentina

Híbrido
ARS 182.877.000 - 274.315.000
Remote or office workspace
Technical webinars and meetups
Bonuses for talks and activities
+1
Senior Compiler Engineer
Senior Compiler Engineer

Werben HR • Buenos Aires

A distancia
ARS 146.337.000 - 313.580.000
Senior Software Engineer – Open Source & SWE-Bench Evaluation
Senior Software Engineer – Open Source & SWE-Bench Evaluation

Anyone AI Inc. • Argentina

A distancia
ARS 9.643.000 - 19.286.000
Remote work
Part-time project-based
Flexible schedule
Machine Learning Engineer – ML Evaluation & Experiment Design
Machine Learning Engineer – ML Evaluation & Experiment Design

Anyone AI Inc. • Argentina

A distancia
ARS 1.200.000 - 2.000.000
Remote work
Part-time project-based consulting
Applied AI/ML Engineer
Applied AI/ML Engineer

FutureProofing • Buenos Aires

A distancia
ARS 150.912.000 - 226.367.000
Remote GPU & ML Infrastructure Engineer
Remote GPU & ML Infrastructure Engineer

Svitla Systems, Inc. • Argentina

Híbrido
ARS 182.877.000 - 274.315.000
Remote or office workspace
Technical webinars and meetups
Bonuses for talks and activities
+1
Graphic designer - Remote
Graphic designer - Remote

YO AI Labs • Buenos Aires

A distancia
ARS 60.959.000 - 106.678.000