Remote AI Systems Research Engineer Intern

Yotta Labs

United States

On-site

USD 34,000 - 62,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive internship compensation
Flexible remote work environment
Fast path to full-time offer

Job summary

Yotta Labs is seeking a motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization. Over a 12–16 week remote internship, you will own a well-scoped project at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, and Large Language Model Infrastructure.

You will implement and optimize kernels for Attention, GEMM, MoE, and quantization; build custom operators with CUDA, Triton, ROCm/HIP, or Neuron; profile and

Qualifications

  • Pursuing BS, MS, or PhD in Computer Science, Computer Engineering, or related field.
  • Strong Python and C++ programming skills.
  • Solid understanding of GPU/accelerator architecture (memory, parallelism, occupancy).
  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels.
  • Familiarity with AI frameworks (PyTorch, Dynamo, LMCache) and profiling tools.
  • Ability to work independently in a collaborative, remote environment.

Responsibilities

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.
  • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
  • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations.
  • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.
  • Ship code upstream to open-source AI infrastructure projects, with tests and documentation.

Skills

Python
C++
GPU/accelerator architecture
PyTorch
Profiling tools

Education

BS/MS/PhD in CS/CE

Tools

CUDA
Triton
ROCm/HIP
Neuron SDK
PyTorch/XLA

Job description

Yotta Labs is seeking a motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization. Over a 12–16 week remote internship, you will own a well-scoped project at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, and Large Language Model Infrastructure.

You will implement and optimize kernels for Attention, GEMM, MoE, and quantization; build custom operators with CUDA, Triton, ROCm/HIP, or Neuron; profile and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Research Intern: Kernel Optimization & AI
GPU Systems Research Intern: Kernel Optimization & AI

Togetherai • San Francisco (CA)

On-site
USD 79,900 - 86,788
Competitive compensation
Housing stipends
Opportunity to work with industry-leading engineers
AI Compiler Research Intern (MLIR, GPU)
AI Compiler Research Intern (MLIR, GPU)

ByteDance • San Jose (CA)

On-site
USD 100,000 - 167,000
AI Operations Intern: Research & Workflow Optimizer
AI Operations Intern: Research & Workflow Optimizer

Ryplaced • Newark (DE)

On-site
Official internship certificate
Letter of recommendation
Mentorship from marketing lead
+3
Systems Research Engineer Intern - GPU Programming (Fall 2026)
Systems Research Engineer Intern - GPU Programming (Fall 2026)

Togetherai • San Francisco (CA)

On-site
USD 79,900 - 86,788
Competitive compensation
Housing stipends
Opportunity to work with industry-leading engineers
ML Research Intern
ML Research Intern

United States Digital Space LLC • New York (NY)

On-site
Remote Generative AI Engineer Intern
Remote Generative AI Engineer Intern

Peraton • United States

On-site
Medical benefits
401(k)
Paid time off
+1
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
ML/AI Intern: Research & Dev for Next-Gen AI
ML/AI Intern: Research & Dev for Next-Gen AI

Advanced Micro Devices, Inc. • Longmont (CO)

Hybrid
USD 27,552,000 - 48,216,000
AMD benefits at a glance
ML Research Intern
ML Research Intern

Modal Labs • New York (NY)

On-site