AI Computing Research Intern

Naïve

Mountain View (CA)

On-site

USD 32,000 - 52,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Naïve is an autonomous company builder that deploys AI employees to run entire companies. We’re a $120M company backed by prominent investors.

We’re building cost-effective, scalable AI systems that can operate fleets of agents with lower costs and higher reliability. As an intern, you’ll research and ship the cost/performance frontier for thousands of AI agents, optimize local model inference, and develop routing strategies to choose the cheapest model that can do the job.

Qualifications

  • Strong systems and ML engineering background with Python and PyTorch.
  • Deep understanding of how transformers run: attention, KV cache, memory bandwidth.
  • Built and shipped a real project—OSS, side project, or internship.
  • Experience with LLMs as tooling and objects of study via API calls and prompts.
  • Move fast, take feedback, push back when you're right.

Responsibilities

  • Research and ship cost/performance frontier for AI agents.
  • Optimize local/self-hosted model inference and related techniques.
  • Build model routing to select cheapest capable model.
  • Benchmark and deploy across hardware (GPUs, edge, on-prem, accelerators).
  • Push on agent infrastructure: orchestration, caching, context management, parallelization.
  • Prototype recursive self-improvement loops.
  • Own a research question end-to-end, from framing to production.

Skills

High agency
Systems ML engineering
Python & PyTorch
Profiling & optimization
Ship without hand-holding

Education

CS/EE/Math student

Tools

vLLM
TensorRT-LLM
SGLang
llama.cpp

Job description

About Naïve

Naïve is an autonomous company builder — we let businesses deploy AI employees to create/run entire companies. We’re a $120M company backed by Y Combinator, Liquid2, DEEPCORE (Softbank), and more.

What You'll Do
  • Research and ship the systems that make running thousands of AI agents dramatically cheaper, faster, and more reliable
  • Optimize local / self-hosted model inference — quantization, batching, speculative decoding, KV-cache strategy, tensor & pipeline parallelism
  • Build model routing that sends every request to the cheapest model that can actually do the job — frontier API when it matters, local when it doesn't
  • Benchmark and deploy across hardware — GPUs, edge, on-prem, alternative accelerators — and turn the numbers into real deployment decisions
  • Push on agent infrastructure: orchestration, caching, context management, and parallelization for fleets of concurrent agents
  • Prototype recursive self-improvement loops — agents that improve their own tooling, prompts, and evals
  • Own a research question end-to-end — frame it, run the experiments, ship the result into production

The Role You're not here to write papers nobody reads. You're here to find the cost/performance frontier and ship past it. Every dollar and millisecond you save compounds across an entire fleet of AI employees. Our best interns take a benchmark on Monday and land a production cost win by Friday — and write the changelog entry themselves. This is research with a deploy button.

Must-Haves
  • High agency
  • Strong systems + ML engineering — comfortable in Python and PyTorch, can profile, optimize, and ship without hand-holding
  • Real understanding of how transformers actually run — attention, KV cache, memory bandwidth, throughput vs. latency tradeoffs
  • Built and shipped something real — side project, OSS, hackathon win, research artifact, prior internship
  • Comfortable with LLMs both as tooling and as objects of study — API calls, prompts, tool use, and what's happening under the hood
  • Move fast, take feedback, push back when you're right
Nice-to-Haves
  • Hands-on with inference engines (vLLM, TensorRT-LLM, SGLang, llama.cpp)
  • GPU kernel or low-level perf work (CUDA, Triton)
  • Hardware benchmarking or deployment experience (cloud GPUs, on-prem, edge, alt accelerators)
  • Built with agent frameworks
  • Published, open-sourced, or blogged research/tooling
  • Currently enrolled in a CS / EE / math program — or dropped out of one to build

P.S. If you're serious about this role, send Sean a connection request with a note over LinkedIn.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Research Intern - Optimize, Deploy, Benchmark
AI Systems Research Intern - Optimize, Deploy, Benchmark

Naïve • Mountain View (CA)

On-site
USD 32,000 - 52,000
AI Researcher Intern
AI Researcher Intern

Serverlessvc • Austin (TX), Northern (KY)

Hybrid
USD 34,000 - 55,000
AI Engineer Intern
AI Engineer Intern

Spatium Lab LLC • Bellevue (WA)

On-site
Competitive pay
Flexible schedule
Mentorship by founders
+2
Applied AI Researcher (Brazil)
Applied AI Researcher (Brazil)

Articul8 • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Senior Research Scientist
Senior Research Scientist

Adaption Labs • United States

Hybrid
USD 100,000 - 140,000
Annual travel stipend
Weekly meal allowance
Comprehensive medical benefits
+1
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Senior Research Scientist
Senior Research Scientist

adaption • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Travel stipend
Lunch stipend
+1
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
AI Engineer
AI Engineer

Teserac, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 130,000
Health Care Plan (Medical, Dental & Vision)
Paid Time Off (Vacation, Sick & Public Holidays)
Free Food & Snacks
+2
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits