Lead Research Software Engineer, Portable AI Performance Engineering

Massachusetts Institute of Technology

Cambridge (MA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A prestigious educational institution in Cambridge, MA is seeking a Lead Research Software Engineer to enhance AI performance engineering. This role involves optimizing workloads for AMD GPUs and requires expertise in Python, C++, and AI frameworks like PyTorch. Candidates should have at least five years of experience in technical fields and demonstrate excellent communication skills. Remote work options are available for this two-year term position.

Qualifications

  • Minimum of five years of experience in technical fields or computational research.
  • Deep familiarity with AI/ML frameworks such as PyTorch, TensorFlow, or JAX.
  • Experience with performance profiling and benchmarking on Linux systems.

Responsibilities

  • Lead applied performance engineering for AI workloads.
  • Optimize existing workloads for AMD GPUs.
  • Profile AI workloads to identify performance bottlenecks.

Skills

Proficiency in Python
Proficiency in C++
Familiarity with AI/ML frameworks
Hands-on experience with GPU programming models
Performance profiling skills
Excellent communication skills
Self-motivation
Collaboration skills

Education

Bachelor’s degree or equivalent

Tools

NVIDIA GPUs
AMD MI355X GPUs
AMD ROCm software stack
Linux-based High-Performance Computing systems

Job description

Posting Description

LEAD RESEARCH SOFTWARE ENGINEER, PORTABLE AI PERFORMANCE ENGINEERING, MA Green High Performance Computing Center, to be a hands‑on research software engineering professional and serve as lead for applied performance engineering for AI workloads. Will work closely with research groups and leading computer industry collaborators to evaluate, adapt, and enhance the portable performance of complex AI research workloads on state‑of‑the‑art hardware. The role will have heavy focus on optimizing existing NVIDIA GPU‑based workloads for top‑tier AMD GPUs, such as MI355X and beyond and will analyze and profile existing research AI workloads to identify performance bottlenecks and portability challenges; and port and optimize complex AI models and scientific code to run efficiently on AMD MI355X GPUs using ROCm, HIP, and related translation tools.

Job Requirements

REQUIRED:

  • Bachelor’s degree or equivalent with a minimum of five years of work experience in either deeply technical fields and/or computational research experience;
  • Strong proficiency in Python and C++, with deep familiarity with AI/ML frameworks (PyTorch, TensorFlow, JAX);
  • Hands‑on experience with GPU programming models (e.g., CUDA, HIP, or OpenCL);
  • Experience with performance profiling and benchmarking tools on Linux‑based High‑Performance Computing systems;
  • Excellent communication skills;
  • Ability to collaborate effectively with academic researchers and industry partners;
  • Self‑motivated with the ability to work independently in a remote or hybrid environment.

PREFERRED:

  • Direct experience with the AMD ROCm software stack and translating CUDA code to HIP;
  • Familiarity with AI agentic tools and Large Language Models (LLMs) used for code generation and refactoring;
  • Background in supporting large‑scale, domain‑specific scientific research (e.g., physics, biology, climate science) on institutional clusters;
  • Direct experience with one or more open‑source schedulers and provisioners;
  • Experience with Linux container technologies such as LXC, apptainer and systemd‑nspawn;
  • Or advanced degree in a relevant technical field.

The Lead Software Engineer must comply with all relevant MGHPCC security policies.

This is a two‑year term position.

3/13/2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Performance Engineer, Portable HPC (Remote)
Lead AI Performance Engineer, Portable HPC (Remote)

Massachusetts Institute of Technology • Cambridge (MA)

Hybrid
USD 120,000 - 160,000
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Senior GPU Software Performance Engineer – Post-Training
Senior GPU Software Performance Engineer – Post-Training

AMD • San Jose (CA)

On-site
USD 170,000 - 250,000
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

On-site
USD 170,000 - 250,000
Senior GPU Software Performance Engineer — Post‐Training
Senior GPU Software Performance Engineer — Post‐Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 130,000 - 170,000
Comprehensive health benefits
Inclusive workplace culture
Opportunities for career advancement
Software Engineer- GPU/AI/ML
Software Engineer- GPU/AI/ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

Hybrid
USD 120,000 - 160,000
Principal Software Development Eng. - AI Performance
Principal Software Development Eng. - AI Performance

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
Senior GPU Inference Performance Engineer
Senior GPU Inference Performance Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Opensource Al workload Software Engineer
Opensource Al workload Software Engineer

Socket.dev • San Jose (CA)

On-site
USD 180,000 - 260,000