LLM Inference Systems Performance Engineer

3M HEALTHCARE

Austin (TX)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision coverage
Income protection benefits
Paid family leave
Paid time off and holidays

Job summary

Micron Technology is seeking an engineer to work on AI training and inference systems, focusing on LLM execution engines, memory hierarchies, and performance optimization across data-center platforms.

The role involves end‑to‑end profiling, benchmarking, and collaboration with senior engineers and researchers. You will develop tools and methods to improve throughput, latency, and resource utilization for large-scale AI workloads.

Qualifications

  • Bachelor’s or Master’s degree in CS/EE or equivalent experience.
  • Proficiency in C/C++ and Python.
  • Experience in Linux environments, debugging, profiling, and automation.
  • Solid understanding of memory systems, NUMA, GPUs, and server architectures.
  • Strong written and verbal communication skills.

Responsibilities

  • Profile and optimize LLM training and inference workloads.
  • Design and evaluate KV-cache and state-management strategies for LLM serving.
  • Build benchmarking, simulation, and emulation frameworks for AI workloads across memory tiers.
  • Develop data placement, migration, and prefetching algorithms for heterogeneous memory pools.
  • Characterize LLM execution engines and analyze throughput and latency across deployments.
  • Collaborate with engineering, architecture, and research teams to influence future AI systems.

Skills

C/C++ programming
Python
Linux
Memory systems
Performance profiling
System optimization
Communication

Education

Bachelor’s or Master’s in CS/EE

Tools

gem5
Ramulator
Perf tooling

Job description

Micron Technology is seeking an engineer to work on AI training and inference systems, focusing on LLM execution engines, memory hierarchies, and performance optimization across data-center platforms.

The role involves end‑to‑end profiling, benchmarking, and collaboration with senior engineers and researchers. You will develop tools and methods to improve throughput, latency, and resource utilization for large-scale AI workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Systems & Infrastructure Intern for LLMs & GPUs
AI Systems & Infrastructure Intern for LLMs & GPUs

Micron • Austin (TX)

On-site
USD 34,440,000 - 48,216,000
Medical coverage
Paid time off
Paid holidays
+1
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
AI Infrastructure Systems Intern — Optimize LLM Workloads
AI Infrastructure Systems Intern — Optimize LLM Workloads

Micron Technology, Inc • Austin (TX)

On-site
USD 34,000 - 48,000
Medical, dental, and vision plans
Paid time off
Paid holidays
LLM AI Systems & Infrastructure Intern
LLM AI Systems & Infrastructure Intern

1000 Micron Technology, Inc. • Austin (TX)

On-site
USD 41,000 - 55,000
Medical, dental, vision plans
Paid time off
Paid holidays
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
ML Systems Engineer — LLM Inference & Multi-Node Performance
ML Systems Engineer — LLM Inference & Multi-Node Performance

ScOp Venture Capital LLC. • Santa Barbara (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity
Professional growth opportunities
+1
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000