AI Infrastructure Intern — Inference at Scale

DeepInfra

Palo Alto (CA)

On-site

USD 78,120 - 89,280

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

DeepInfra seeks talented Software Engineering Interns to join its team in Palo Alto, offering hands-on experience designing, developing, and deploying scalable AI models. You will work closely with engineers to implement inference solutions and optimize AI systems using Python, C++, CUDA, and NCCL.

The role includes monitoring live services and participating in code reviews to ensure high-quality software delivery.

Qualifications

  • Currently pursuing a Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field.
  • Proficiency in Python and knowledge of AI/ML libraries is required.
  • Familiarity with AI models, Transformers and Diffusers is preferred.
  • Experience with Git and agile development methodologies is beneficial.
  • Strong problem-solving and collaboration skills are essential.

Responsibilities

  • Collaborate with the engineering team to design, develop, and test inference solutions for top AI models.
  • Implement and optimize AI models using Python, C++, CUDA, NCCL.
  • Monitor and maintain the live service.
  • Participate in code reviews, feature development, and bug fixes to ensure high-quality software delivery.
  • Engage in daily stand-ups and design discussions to enable seamless collaboration.
  • Stay updated with AI/ML industry trends and advancements.

Skills

Python
Algorithms
Design patterns
Teamwork

Education

Bachelor's or Master's in CS/CE or related field

Tools

Git
NumPy
pandas
SciPy
TensorFlow
PyTorch
Transformers
Diffusers
CUDA
NCCL

Job description

DeepInfra seeks talented Software Engineering Interns to join its team in Palo Alto, offering hands-on experience designing, developing, and deploying scalable AI models. You will work closely with engineers to implement inference solutions and optimize AI systems using Python, C++, CUDA, and NCCL.

The role includes monitoring live services and participating in code reviews to ensure high-quality software delivery.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer Intern — Build at-scale AI models
AI Infrastructure Engineer Intern — Build at-scale AI models

DeepInfra • United States

Remote
USD 40,000 - 60,000
AI Infrastructure Software Engineer
AI Infrastructure Software Engineer

DeepInfra • United States

Remote
USD 120,000 - 170,000
Software Engineer Intern (US)
Software Engineer Intern (US)

DeepInfra • Palo Alto (CA)

On-site
Software Engineer Intern (Europe)
Software Engineer Intern (Europe)

DeepInfra • United States

Remote
USD 40,000 - 60,000
Early-Career AI Infra Engineer - Scale Production
Early-Career AI Infra Engineer - Scale Production

DeepInfra • Palo Alto (CA)

On-site
USD 140,000 - 150,000
Paid Inference Architecture Intern (12-Week)
Paid Inference Architecture Intern (12-Week)

The Consensus • San Jose (CA)

On-site
12-week paid internship
Housing support for relocation
Daily lunch and dinner in office
+2
AI Infrastructure Engineer Intern
AI Infrastructure Engineer Intern

Feedinkoo • United States

Remote
Competitive compensation
Mentorship from experienced engineers
Real-world project ownership
+1
AI Infrastructure & ML Systems Intern
AI Infrastructure & ML Systems Intern

ByteDance • San Jose (CA)

On-site
USD 97,000 - 138,000
Health insurance
Housing allowance
Paid holidays
AI Infrastructure Engineer Intern: Large-Scale Inference
AI Infrastructure Engineer Intern: Large-Scale Inference

TikTok • San Jose (CA)

On-site
USD 70,000 - 95,000
Health insurance
Housing allowance
Paid holidays
AI Infrastructure Intern: Foundation Models & Training
AI Infrastructure Intern: Foundation Models & Training

ByteDance • San Jose (CA)

On-site
USD 63,000 - 89,000