Research Engineer [ Performance Engineering ]

Metamorphic

Palo Alto (CA)

On-site

USD 200,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Competitive compensation
Mentorship and career development
Equity package

Job summary

Metamorphic in Palo Alto seeks Research Engineers to maximize training and inference performance of our foundation models. You will write and optimize GPU kernels, explore quantization, and contribute to efficient, scalable implementations in close collaboration with researchers.

This role blends deep ML research with practical engineering, offering autonomy on a small, high‑impact team and the chance to shape future AI systems at the frontier of neuroscience‑informed AI.

Qualifications

  • Bachelor's degree or higher in Computer Science, ML, or related field.
  • Strong software engineering skills with track record of building complex systems.
  • Proficiency in CUDA, Triton, or similar; experience writing GPU kernels.
  • Hands-on experience with mixed-precision and low-precision training; numerical stability tradeoffs.
  • Deep knowledge of transformer architectures at implementation level.
  • Experience with MoE architectures: routing, load balancing, expert dispatch across GPUs.
  • Experience with GPU profiling tools (Nsight Compute, Nsight Systems, PyTorch Profiler).
  • Experience integrating high-performance libraries (FlashAttention, cuDNN, Triton, Quack).

Responsibilities

  • Write and optimize GPU kernels for training and inference.
  • Improve performance via quantization, low-precision training, and MoE routing optimizations.
  • Collaborate with researchers to translate theory into scalable implementations.
  • Profile code, identify bottlenecks, and drive incremental performance improvements.

Skills

Software engineering
CUDA
Triton
GPU kernels
MoE routing
Profiling tools
Low-precision training
Collaboration
Pair programming

Education

Bachelor's degree or higher in CS/ML

Tools

CUDA
Triton
Nsight Compute/Systems
FlashAttention
cuDNN
Quack

Job description

About Metamorphic

Metamorphic is developing new approaches to intelligence by combining machine learning with large-scale experimental neuroscience, informed by the principles that make the brain efficient, flexible, and robust. We are building foundation models trained on rich, continuous neural data – a high-resolution model of the brain at a scale never before possible. Our founding team spans machine learning, neuroscience, and neurotechnology, with prior work including the MICrONS project, Neuropixels, and the Enigma project, as well as foundational scientific contributions in learning, neural computation, and generative modeling. Our work sits at the frontier of AI research, and we believe the highest‑impact discoveries will come from researchers and engineers working as a single, tightly collaborative team. The name Metamorphic reflects our belief that the next advances in intelligence will come from a change in form, beyond scale – from artificial to natural intelligence.

About The Role

We are seeking Research Engineers to join our growing AI research team. You will maximize the training and inference performance of Metamorphic's foundation models, from quantization and low‑precision training, to MoE routing optimization, to writing custom CUDA/Triton kernels for our novel architecture. This is a high‑impact, technically deep role at the frontier of ML research and engineering. You will write and optimize GPU kernels, profile and eliminate performance bottlenecks, tune low‑precision training strategies, and work closely with researchers to ensure architectural decisions translate to efficient and scalable implementations. You'll have substantial autonomy to shape foundational technical decisions on a small, high‑impact team.

You'll Thrive In This Role If You
  • Have significant software engineering experience and can move quickly without sacrificing rigor
  • Can balance research goals with practical engineering constraints
  • Translate theory and practice, turning paper ideas into robust and performant implementations
  • Get excited about nitty‑gritty engineering details and incremental performance improvements that others gloss over
  • Are willing to take on tasks outside your job description to support the team
  • Enjoy pair programming and deeply collaborative work
  • Are eager to learn more about machine learning research in a novel scientific domain
  • Are enthusiastic to work at an organization that functions as a single, cohesive team pursuing large‑scale AI research
  • Have ambitious goals for AI progress and are excited to create the best outcomes over the long term
We Offer
  • The chance to work on one of the most scientifically consequential AI projects being pursued today
  • A small, world‑class team where your contributions directly shape the science and the company
  • Competitive compensation and benefits, along with visa sponsorship
  • Strong mentorship and career development
Salary Range

$200,000 - $280,000 USD – Based on experience. We additionally offer a competitive equity package and comprehensive benefits, as well as visa sponsorship for international candidates.

Minimum Qualifications
  • Bachelor's degree or higher in Computer Science, Machine Learning, or a related field
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong proficiency in CUDA, Triton, or similar, with demonstrated experience writing and optimizing GPU kernels
  • Hands‑on experience with mixed‑precision and low‑precision training and a practical understanding of numerical stability tradeoffs
  • Deep knowledge of transformer architectures at the implementation level
  • Experience with MoE architectures: routing algorithms, load balancing, and the systems‑level challenges of expert dispatch across GPUs
  • Hands‑on experience with GPU profiling tools (Nsight Compute, Nsight Systems, PyTorch Profiler)
  • Experience integrating, customizing, and extending third‑party high‑performance libraries (FlashAttention, cuDNN, Triton, Quack, or similar) into production training stacks
Nice to Have
  • Experience with CUTLASS, cuDNN APIs, and NCCL internals
  • Familiarity with inference optimization techniques and serving frameworks
  • Familiarity with diffusion models or multimodal model architectures
  • Experience with inference optimization techniques (KV‑cache management, speculative decoding, post‑training quantization) and serving frameworks (vLLM, TensorRT‑LLM)

We encourage you to apply even if you do not believe you meet every single qualification. If you don't see a role that fits, we encourage you to submit a general application and tell us how you'd like to contribute to our mission.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer [ Data Engineering ]
Research Engineer [ Data Engineering ]

Metamorphic • Palo Alto (CA)

On-site
USD 175,000 - 250,000
Research Engineer [ Distributed Training ]
Research Engineer [ Distributed Training ]

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Equity package
+1
Research Engineer [ Agentic Systems ]
Research Engineer [ Agentic Systems ]

Metamorphic • Palo Alto (CA)

On-site
USD 140,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship & career development
ML Research Scientist (Embodied AI & Reinforcement Learning)
ML Research Scientist (Embodied AI & Reinforcement Learning)

Metamorphic • Palo Alto (CA)

On-site
USD 175,000 - 250,000
Competitive compensation
Strong mentorship
Career development opportunities
+1
Robotics Engineer (Simulation, Hardware & Deployment)
Robotics Engineer (Simulation, Hardware & Deployment)

Metamorphic • Palo Alto (CA)

On-site
USD 175,000 - 250,000
Competitive compensation
Strong mentorship and career development
Visa sponsorship
High-Performance AI Research Engineer (CUDA/ML)
High-Performance AI Research Engineer (CUDA/ML)

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship and career development
+1
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Research Engineer, Infrastructure, Numerics
Research Engineer, Infrastructure, Numerics

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1