Senior AI Kernel & Performance Engineer (MTIA)

Meta

Menlo Park (CA)

On-site

USD 154,000 - 217,000

Full time

3 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking an experienced kernel and performance engineer for MTIA Software to own accelerator workloads, from fused attention kernels to convergence decisions. You will read hardware specs and RTL-adjacent docs, and determine whether the hardware or software is the bottleneck.

This hands-on role covers kernel libraries, C++/Python tooling, and pre-silicon bring-up, with broad impact on PyTorch operators and production workloads.

Qualifications

  • Bachelor's degree in CS/CE or related field.
  • Experience writing and optimizing kernels for parallel architectures (GPU/ASIC).
  • Strong reading of hardware specs and RTL-adjacent docs.
  • Proven ability to measure and close roofline gaps.

Responsibilities

  • Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators.
  • Profile and root-cause performance across stacks to drive fixes.
  • Build and extend kernel authoring frameworks and libraries.
  • Deliver and maintain broad kernel coverage for PyTorch operators across workloads.
  • Partner with silicon teams on hardware/software co-design and roofline validation.
  • Mentor engineers on performance methodology and accelerator programming.

Skills

C++ & Python
Kernel performance engineering
GPU/AI accelerator experience
Performance profiling
Architecture awareness
Debugging performance issues
RTL/architecture reading

Education

Bachelor's degree in CS/CE or related field

Tools

CUTLASS
Triton / CuTe
LLVM / MLIR / XLA
PyTorch internals
NCCL / RCCL

Job description

Meta is seeking an experienced kernel and performance engineer for MTIA Software to own accelerator workloads, from fused attention kernels to convergence decisions. You will read hardware specs and RTL-adjacent docs, and determine whether the hardware or software is the bottleneck.

This hands-on role covers kernel libraries, C++/Python tooling, and pre-silicon bring-up, with broad impact on PyTorch operators and production workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Kernel & Performance Engineer
Senior AI Kernel & Performance Engineer

Meta Careers • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Senior AI Kernel & Performance Engineer - Pre-Silicon
Senior AI Kernel & Performance Engineer - Pre-Silicon

Meta • New York (NY)

On-site
USD 184,000 - 257,000
AI Accelerator Kernel Architect for ASICs
AI Accelerator Kernel Architect for ASICs

Meta • Austin (TX)

On-site
USD 146,000 - 209,000
Bonus
Equity
Benefits
Senior ASIC Kernel Architect for AI Accelerators
Senior ASIC Kernel Architect for AI Accelerators

Meta • Sunnyvale (CA)

On-site
USD 190,000 - 250,000
Senior ASIC Kernel Architect for ML Accelerators
Senior ASIC Kernel Architect for ML Accelerators

Meta Careers • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
ASIC Engineer, Architecture Kernel Development
ASIC Engineer, Architecture Kernel Development

Meta • Sunnyvale (CA)

On-site
USD 190,000 - 250,000
Systems ML Engineer - AI Compiler & Kernel Optimizations
Systems ML Engineer - AI Compiler & Kernel Optimizations

Meta Careers • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer, AI Kernels & Performance Optimization — MTIA Software
Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
AI Accelerator Tooling Systems Engineer
AI Accelerator Tooling Systems Engineer

Meta Careers • Menlo Park (CA)

On-site
USD 150,000 - 190,000
Software Engineer, Systems ML - Compilers / Kernels
Software Engineer, Systems ML - Compilers / Kernels

Meta Careers • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000