AI Infrastructure Engineer

Netpreme

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Performance bonus
Equity grant
Health, dental, vision fully paid
401(k) match
Life, disability insurance
Lunch stipend
Office hubs with parking and EV
Visa sponsorship

Job summary

Netpreme is seeking an AI Infrastructure Engineer to build and operate high-performance serving infrastructure on Kubernetes, working hands-on with vLLM and SGLang in a small, high-autonomy team.

You will design parallelism strategies, tune KV cache and batching, profile bottlenecks, and develop benchmarks to balance latency, throughput, and capacity while collaborating with hardware and model engineers.

Qualifications

  • Hands-on experience deploying and tuning vLLM and/or SGLang.
  • Understanding of LLM inference fundamentals: prefill vs decode, batching, KV cache, latency/throughput trade-offs.

Responsibilities

  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency, GPU and memory utilization.
  • Analyze model architecture and serving implications (attention, KV cache, MoE, long context, decoding).
  • Tune vLLM and SGLang configurations (batching, KV-cache, CUDA Graphs, disaggregation).
  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
  • Collaborate with hardware/systems team to translate performance requirements into backend decisions.

Skills

LLM inference
ML systems
GPU systems
Performance engineering
Python
Inference-serving
Kubernetes

Education

BS/MS/PhD in CS/CE or related field

Tools

Nsight Systems
PyTorch Profiler
Kubernetes

Job description

About the Role

We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.

Essential Duties & Responsibilities
  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.
  • Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.
  • Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.
  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
  • Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.
  • Work with model/system engineers to bring newly released models into production efficiently.
  • Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.
  • Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.
Qualifications
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.
  • Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.
  • Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.
  • Strong Python engineering skills.
  • Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.
  • Clear written and verbal communication skills to work effectively with a small, fully distributed team.
Preferred Qualifications (optional)
  • Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc.
  • Experience with MoE / long-context model deployment.
  • Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.
  • Experience with Nsight Systems / PyTorch Profiler.
  • Familiarity with Kubernetes / production GPU serving.
  • Previous startup experience.
Compensation & Benefits
  • Competitive salary with performance-based bonus and early-stage equity grant
  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents
  • 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals
  • 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks
  • Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)
  • Daily lunch stipend
  • Enterprise-level Claude & ChatGPT access with a generous token budget
  • Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office
  • Visa sponsorship and relocation assistance to one of our office hubs
  • A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture
The Opportunity
  • Impact: Humanity stands at the dawn of a new industrial revolution driven by AI - one with the potential to redefine how we live on this planet. We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency. The work we do here compounds across state-of-the-art AI models, systems, and real-world applications.
  • Timing: Breakthrough technology matters most when it meets the right time. Joining now means real ownership of the company and meaningful influence over product direction and execution. In this early-stage environment, your ideas shape the trajectory of the technology - not just its implementation. You’ll work from first principles, move quickly from insight to execution, and see your contributions directly reflected in what we build.
  • Culture: You’ll work alongside a group of people who care deeply about rigor, clarity, and impact. We value thoughtful disagreement, fast learning, and intellectual fearlessness. This is a place where strong ideas shine, curiosity is encouraged, and growth is a daily practice - not a future promise.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Santa Clara (CA), Boston (MA)

On-site
USD 150,000 - 210,000
Health, Dental, and Vision coverage
401(k) match
Life, Disability and AD&D insurance
+4
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff Inference Engineer
Staff Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 170,000 - 230,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
AI/ML Engineer - $84.13 - $120.19 per hour
AI/ML Engineer - $84.13 - $120.19 per hour

7Seventy Recruiting • United States

On-site
USD 140,000 - 210,000
Health Insurance
Conference travel budget
Flexible paid time off
+1