Inference Optimization Manager

Modular

United States

Remote

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
401k with match
Tax advice
Home office stipend
Unlimited PTO
Flexible hours

Job summary

Modular is seeking a senior leader to head a high‑impact team driving LLM inference on Modular Cloud. You will partner across GTM, Product, and Engineering to redefine inference performance, shaping the product direction and building scalable systems that meet demand.

You’ll own the technical direction to achieve Pareto‑optimal performance across GPUs and ASICs, translating customer workloads into actionable engineering roadmaps while growing and mentoring a top-tier team.

Qualifications

  • 5+ years in distributed systems or performance engineering.
  • Hands-on GPU kernel programming experience.
  • Experience leading or managing engineering teams.
  • Ability to translate customer/product needs into engineering direction.
  • Familiarity with modern LLM architectures and inference optimization techniques.

Responsibilities

  • Lead a high impact team partnering with GTM, Product, and Engineering.
  • Shape product direction of Modular Cloud and build systems for scalable performance.
  • Own technical direction to achieve state-of-the-art LLM performance on Modular Cloud.
  • Collaborate with GTM to tailor inference to customer workloads and drive full-stack optimizations.
  • Translate customer engagements into a clear technical roadmap for engineering teams.
  • Grow and develop a high impact team and foster an environment for engineers to excel.
  • Represent the team externally through blog posts on LLM inference optimization.

Skills

GPU kernel programming
Inference engine internals
Distributed inference architectures
Cloud native ecosystems
LLM architectures
Team leadership

Tools

Kubernetes

Job description

  • In this role you will lead a high impact team that partners closely with GTM, Product, and Engineering to redefine what an inference platform can be
  • You will turn real customer workloads into a continuous optimization loop, shape the product direction of Modular Cloud, and build the systems that let performance scale with demand
  • Lead the team that builds the optimization platform driving LLM inference on Modular Cloud to state of the art performance across the latest GPU and ASIC architectures
  • Own the technical direction of Modular Cloud, delivering LLM performance on the Pareto frontier for agentic use cases and keeping it there as the landscape evolves
  • Partner closely with the GTM team to deliver highly customized LLM inference tuned to specific customer workloads, and collaborate across engineering to drive optimizations spanning the full stack, from GPU kernels to cloud infrastructure. Translate insights from customer engagements into a clear technical roadmap that engineering teams can execute against
  • Grow and develop a high impact team, set priorities, and create the conditions for engineers to do their best work on hard performance problems
  • Champion the team externally through blog posts on innovative approaches to LLM inference optimization that shape industry wide best practices
Benefits
  • A variety of fantastic health benefits (health, dental, vision insurance; life insurance etc) are available
  • A 401k plan with up to 5% match
  • Free tax advice on Carta
  • Generous work-from-home stipend of $1500 to help you improve your home office
  • Unlimited paid time off and flexible work hours

Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skillsA track record of shipping durable, reusable software tools and libraries adopted across teams and functions, and of guiding a team to do the sameThe ability to translate ambiguous customer and product needs into focused engineering directionCreativity and curiosity in solving complex problems, a collaborative and team oriented mindset, and alignment with our culture5+ years in distributed systems or performance engineering, including experience leading or managing engineering teamsHands on background in GPU kernel programming, inference engine internals, or distributed inference architecturesExperience with Kubernetes and cloud native ecosystemsFamiliarity with modern LLM architectures and the latest inference optimization techniques

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Optimization Engineer
Inference Optimization Engineer

Modular • United States

On-site
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

On-site
USD 180,000 - 320,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
Inference Engineering and Product Lead
Inference Engineering and Product Lead

Modal • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

On-site
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Inference Optimization Engineer United States - Remote · Remote
Inference Optimization Engineer United States - Remote · Remote

Modular Inc • United States

Remote
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
LLM Inference Optimization Leader
LLM Inference Optimization Leader

Modular • United States

Remote
USD 180,000 - 240,000
Health benefits
401k with match
Tax advice
+3
Remote LLM Inference Optimization Architect
Remote LLM Inference Optimization Architect

Modular • United States

Hybrid
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1