High-Performance AI Research Engineer (CUDA/ML)

Metamorphic

Palo Alto (CA)

On-site

USD 200,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Competitive compensation
Mentorship and career development
Equity package

Job summary

Metamorphic in Palo Alto seeks Research Engineers to maximize training and inference performance of our foundation models. You will write and optimize GPU kernels, explore quantization, and contribute to efficient, scalable implementations in close collaboration with researchers.

This role blends deep ML research with practical engineering, offering autonomy on a small, high‑impact team and the chance to shape future AI systems at the frontier of neuroscience‑informed AI.

Qualifications

  • Bachelor's degree or higher in Computer Science, ML, or related field.
  • Strong software engineering skills with track record of building complex systems.
  • Proficiency in CUDA, Triton, or similar; experience writing GPU kernels.
  • Hands-on experience with mixed-precision and low-precision training; numerical stability tradeoffs.
  • Deep knowledge of transformer architectures at implementation level.
  • Experience with MoE architectures: routing, load balancing, expert dispatch across GPUs.
  • Experience with GPU profiling tools (Nsight Compute, Nsight Systems, PyTorch Profiler).
  • Experience integrating high-performance libraries (FlashAttention, cuDNN, Triton, Quack).

Responsibilities

  • Write and optimize GPU kernels for training and inference.
  • Improve performance via quantization, low-precision training, and MoE routing optimizations.
  • Collaborate with researchers to translate theory into scalable implementations.
  • Profile code, identify bottlenecks, and drive incremental performance improvements.

Skills

Software engineering
CUDA
Triton
GPU kernels
MoE routing
Profiling tools
Low-precision training
Collaboration
Pair programming

Education

Bachelor's degree or higher in CS/ML

Tools

CUDA
Triton
Nsight Compute/Systems
FlashAttention
cuDNN
Quack

Job description

Metamorphic in Palo Alto seeks Research Engineers to maximize training and inference performance of our foundation models. You will write and optimize GPU kernels, explore quantization, and contribute to efficient, scalable implementations in close collaboration with researchers.

This role blends deep ML research with practical engineering, offering autonomy on a small, high‑impact team and the chance to shape future AI systems at the frontier of neuroscience‑informed AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer [ Performance Engineering ]
Research Engineer [ Performance Engineering ]

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship and career development
+1
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Research Engineer [ Distributed Training ]
Research Engineer [ Distributed Training ]

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Equity package
+1
Founding ML Inference Performance Engineer
Founding ML Inference Performance Engineer

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
GPU Kernel Engineer — Fast ML Training & Inference
GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
Research Engineer [ Data Engineering ]
Research Engineer [ Data Engineering ]

Metamorphic • Palo Alto (CA)

On-site
USD 175,000 - 250,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed ML Engineer for High-Performance AI Training
Distributed ML Engineer for High-Performance AI Training

Ifm Us • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Medical benefits
Dental benefits
Vision benefits
+7