Inference Optimization Intern – Performance Modeling

Institute of Foundation Models

Sunnyvale (CA)

On-site

USD 34,440 - 55,104

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hands-on experience with advanced AI systems

Job summary

The Institute of Foundation Models in Sunnyvale, California, seeks an intern to help optimize large-scale AI systems on NVIDIA GPUs. You will work alongside researchers in a fast-paced environment, contributing to performance modeling and simulator development.

This internship requires strong programming skills in C++, CUDA, and knowledge of NVIDIA GPU architecture. Ideal candidates are pursuing degrees in Computer Science or related fields and have a passion for AI and performance engineering.

Qualifications

  • Currently pursuing a degree in a relevant field.
  • Strong programming skills in C++, CUDA, and Python are required.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.

Responsibilities

  • Develop analytical performance models for GPU kernels.
  • Build and validate a simulator for NVIDIA GPUs.
  • Document findings and provide actionable recommendations.

Skills

CUDA programming
GPU kernel development
Performance profiling tools
Python
C++

Education

Degree in Computer Science or related field

Tools

NVIDIA Nsight Systems
NVIDIA Nsight Compute

Job description

About the Institute of Foundation Models

The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting‑edge foundation models while pushing the limits of high‑performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real‑world impact.

As part of the team, interns work alongside world‑class researchers and performance engineers to optimize the execution of large-scale foundation models on next‑generation NVIDIA GPU architectures. This internship provides hands‑on experience in low-level GPU performance analysis, kernel optimization, and hardware‑aware inference acceleration.

Key Responsibilities

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVIDIA GPUs.

Responsibilities include:

  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low-level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications

Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.

Preferred Qualifications
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.
Desired Skills
  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware-software co-design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Optimization Intern – Performance Modeling
Inference Optimization Intern – Performance Modeling

Ifm Us • Sunnyvale (CA)

On-site
USD 30,000 - 60,000
GPU Inference Performance Intern & Modeling
GPU Inference Performance Intern & Modeling

Institute of Foundation Models • Sunnyvale (CA)

On-site
Hands-on experience with advanced AI systems
Inference Optimization Intern: GPU Performance Modeling
Inference Optimization Intern: GPU Performance Modeling

Ifm Us • Sunnyvale (CA)

On-site
USD 30,000 - 60,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000