Inference Research Intern: Build Scalable AI Serving

Together AI

San Francisco (CA)

On-site

USD 80,000 - 96,000

Full time

40 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Housing stipend
Competitive compensation

Job summary

Together AI in San Francisco is seeking a research intern to accelerate inference systems for large foundation models. You will explore distributed inference, compiler-aware optimization, and cross-layer strategies across models, systems, and hardware.

Join a collaborative team, publish findings, and contribute to faster serving and scalable deployment, with a 12–14 week program from January through April, housing stipends, and competitive hourly compensation.

Qualifications

  • Final year student in CS/EE or related field
  • Strong ML/DL fundamentals
  • Experience with PyTorch or JAX
  • Strong Python programming
  • Familiarity with Transformer architectures and foundation models

Responsibilities

  • Design and conduct rigorous experiments
  • Communicate plans, progress, and results
  • Document findings in publications and blogs

Skills

ML fundamentals
Deep Learning
PyTorch/JAX
Python
Transformer models

Education

Pursuing final year of Bachelor's, Master's, or Ph.D.

Tools

CUDA

Job description

Together AI in San Francisco is seeking a research intern to accelerate inference systems for large foundation models. You will explore distributed inference, compiler-aware optimization, and cross-layer strategies across models, systems, and hardware.

Join a collaborative team, publish findings, and contribute to faster serving and scalable deployment, with a 12–14 week program from January through April, housing stipends, and competitive hourly compensation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Research Intern – Summer 2027
Inference Systems Research Intern – Summer 2027

Together AI • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Research Intern, Inference (Winter 2027)
Research Intern, Inference (Winter 2027)

Teli Labs • San Francisco (CA)

On-site
USD 120,000 - 146,000
Housing stipend
Competitive compensation
Equal opportunity employer
Research Intern, Inference (Winter 2027)
Research Intern, Inference (Winter 2027)

Together AI • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Research Intern, Inference (Summer 2027)
Research Intern, Inference (Summer 2027)

Togetherai • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive benefits
Research Intern, Inference (Winter 2027)
Research Intern, Inference (Winter 2027)

Together • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipends
Equal opportunity employer
Research Intern, Inference (Winter 2027)
Research Intern, Inference (Winter 2027)

Togetherai • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Paid Inference Architecture Intern (12-Week)
Paid Inference Architecture Intern (12-Week)

The Consensus • San Jose (CA)

On-site
12-week paid internship
Housing support for relocation
Daily lunch and dinner in office
+2
Research Intern, Inference (Summer 2027)
Research Intern, Inference (Summer 2027)

Together AI • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Inference Architecture Intern: AI Accelerator Design
Inference Architecture Intern: AI Accelerator Design

Etched • San Jose (CA), Northern (KY)

Hybrid
USD 13,225,000 - 19,837,000
Housing support
Lunch & dinner
Mentorship
+1
AI Infrastructure Intern — Inference at Scale
AI Infrastructure Intern — Inference at Scale

DeepInfra • Palo Alto (CA)

On-site