Inference Optimization Intern – Performance Modeling

Ifm Us

Sunnyvale (CA)

On-site

USD 30,000 - 60,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ifm Us is seeking an intern to work on the optimization of large-scale foundation models using NVIDIA GPU architectures. Interns will have the chance to collaborate with world-class researchers and engineers, gaining hands-on experience in GPU performance analysis and kernel optimization.

The ideal candidate should be pursuing a degree in Computer Science or related fields, with strong programming skills in CUDA, C++, and Python. This is a unique opportunity to contribute to cutting-edge developments in AI systems.

Qualifications

  • Currently pursuing a degree in a relevant quantitative discipline.
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.

Responsibilities

  • Develop analytical performance models for GPU kernels.
  • Build and validate a simulator for performance estimation.
  • Identify performance bottlenecks in compute and memory.

Skills

CUDA programming
GPU kernel development
C++
Python
Nsight Systems
Nsight Compute
Deep learning frameworks (PyTorch, TensorFlow)
Performance analysis

Education

Pursuing degree in Computer Science or related field

Job description

About the Institute of Foundation Models

The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting‑edge foundation models while pushing the limits of high‑performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real‑world impact.

As part of the team, interns work alongside world‑class researchers and performance engineers to optimize the execution of large-scale foundation models on next‑generation NVIDIA GPU architectures. This internship provides hands‑on experience in low‑level GPU performance analysis, kernel optimization, and hardware‑aware inference acceleration.

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVIDIA GPUs.

Responsibilities
  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low‑level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer‑based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.
Qualifications
  • Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High‑Performance Computing, or a related quantitative discipline.
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low‑level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.
  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware‑software co‑design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Optimization Intern – Performance Modeling
Inference Optimization Intern – Performance Modeling

Institute of Foundation Models • Sunnyvale (CA)

On-site
Hands-on experience with advanced AI systems
Inference Optimization Intern: GPU Performance Modeling
Inference Optimization Intern: GPU Performance Modeling

Ifm Us • Sunnyvale (CA)

On-site
USD 30,000 - 60,000
GPU Inference Performance Intern & Modeling
GPU Inference Performance Intern & Modeling

Institute of Foundation Models • Sunnyvale (CA)

On-site
Hands-on experience with advanced AI systems
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000