LLM Inference Systems Performance Engineer

3M HEALTHCARE

Austin (TX)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision coverage
Income protection benefits
Paid family leave
Paid time off and holidays

Job summary

Micron Technology is seeking an engineer to work on AI training and inference systems, focusing on LLM execution engines, memory hierarchies, and performance optimization across data-center platforms.

The role involves end‑to‑end profiling, benchmarking, and collaboration with senior engineers and researchers. You will develop tools and methods to improve throughput, latency, and resource utilization for large-scale AI workloads.

Qualifications

  • Bachelor’s or Master’s degree in CS/EE or equivalent experience.
  • Proficiency in C/C++ and Python.
  • Experience in Linux environments, debugging, profiling, and automation.
  • Solid understanding of memory systems, NUMA, GPUs, and server architectures.
  • Strong written and verbal communication skills.

Responsibilities

  • Profile and optimize LLM training and inference workloads.
  • Design and evaluate KV-cache and state-management strategies for LLM serving.
  • Build benchmarking, simulation, and emulation frameworks for AI workloads across memory tiers.
  • Develop data placement, migration, and prefetching algorithms for heterogeneous memory pools.
  • Characterize LLM execution engines and analyze throughput and latency across deployments.
  • Collaborate with engineering, architecture, and research teams to influence future AI systems.

Skills

C/C++ programming
Python
Linux
Memory systems
Performance profiling
System optimization
Communication

Education

Bachelor’s or Master’s in CS/EE

Tools

gem5
Ramulator
Perf tooling

Job description

Micron Technology is seeking an engineer to work on AI training and inference systems, focusing on LLM execution engines, memory hierarchies, and performance optimization across data-center platforms.

The role involves end‑to‑end profiling, benchmarking, and collaboration with senior engineers and researchers. You will develop tools and methods to improve throughput, latency, and resource utilization for large-scale AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Systems Architect
LLM Inference Systems Architect

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
LLM Inference Systems Performance Architect
LLM Inference Systems Performance Architect

Doist • San Jose (CA)

On-site
USD 245,000 - 325,000
Health Insurance
Dental Insurance
Vision Insurance
+8
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Systems Performance Engineer
Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Post-Training LLM Inference Platform Engineer
Post-Training LLM Inference Platform Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 100,000 - 150,000
Comprehensive benefits
Collaborative work environment
Opportunities for career advancement
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1