Senior AI Runtime Engineer

Modular, a Qualcomm company

United States

Hybrid

USD 150,000 - 210,000

Full time

13 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Amazing team
World-class benefits

Job summary

Modular, a Qualcomm company, is building the next-generation AI platform. As an AI Runtime Engineer, you will own a runtime across CPU and GPU hardware, optimizing performance for diverse AI models and deployments.

You will design cross-stack optimizations, port the runtime to new hardware, and work with multiple teams to achieve end-to-end performance. This role offers hybrid work with occasional in-person onboarding in the US.

Qualifications

  • 5+ years of experience in high-performance computing systems.
  • Experience with C++ programming and complex software systems.
  • Experience with CPU/GPU runtime optimizations and performance analysis.
  • Proficiency with one or more profiling tools.

Responsibilities

  • Design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency.
  • Port the Modular runtime stack to new hardware platforms and develop an API.
  • Collaborate with compiler, kernels, serving and models teams to achieve state-of-the-art performance.
  • Collaborate with customer success to understand performance requirements and use cases.
  • Collaborate with tooling and infrastructure teams to design systems for automated performance analysis.

Skills

C++ programming
High-performance computing
Performance analysis
Team collaboration

Education

Advanced degree in CS

Tools

Profiling tools

Job description

About Modular

At Modular, a Qualcomm company, we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.

If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value.

About The Role

ML developers today face significant friction when deploying trained models. They work in a fragmented space with incomplete, patchwork solutions that require extensive performance tuning and model-specific optimizations. At Modular, we are building the next-generation AI platform that will radically improve how developers build and deploy AI models. A core part of this offering is a platform that enables customers to achieve state-of-the-art performance across model families and frameworks. As an AI Runtime Engineer, you will own a runtime that operates on various CPU and GPU hardware platforms, optimizing performance for diverse customer AI models.

LOCATION

Candidates based in the US or Canada are welcome to apply. You can work in our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office.

What You Will Do
  • Design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency, addressing issues such as CPU overhead, caching, and data locality across multiple devices.
  • Port the Modular runtime stack to new hardware platforms and develop an API to streamline this process.
  • Collaborate with the compiler, kernels, serving, and models teams to design core technologies that achieve state-of-the-art end-to-end performance on various CPU and GPU hardware.
  • Collaborate with the customer success team and engage with customers to understand their performance requirements and use cases.
  • Collaborate with tooling and infrastructure teams to design systems for automated performance analysis and benchmarking.
What You Bring To The Table
  • 5+ years of experience working on high-performance computing systems.
  • Experience in C++ programming and complex software systems.
  • Experience with CPU or GPU runtime optimizations and performance analysis on CPUs, GPUs, or AI accelerators.
  • Proficiency with one or more profiling tools (CPU or GPU).
  • Creativity and curiosity for solving complex problems, a team-oriented attitude that enables you to work well with others, and alignment with our culture.
Helpful, But Not Required
  • Experience with ML graph optimizations, parallel / distributed programming, heterogeneous ML computation, and/or code generation.
  • Exposure to MLIR, LLVM, and/or the Mojo programming language.
  • Advanced degree in Computer Science or a related area is a plus.
What Modular Brings To The Table
  • Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders.
  • World-class Benefits. In order to attract the best, we need to offer the best. Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources,
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mojo Libraries Engineer
Mojo Libraries Engineer

Modular • United States

On-site
USD 148,000 - 270,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Mojo Libraries Engineer
Mojo Libraries Engineer

Modular, a Qualcomm company • United States

On-site
USD 148,000 - 270,000
Healthcare coverage
Retirement plans
RSU grants
+5
Senior AI Kernel Engineer
Senior AI Kernel Engineer

Modular • United States

On-site
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
AI Runtime Engineer - Cross-Platform Performance (Remote)
AI Runtime Engineer - Cross-Platform Performance (Remote)

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 210,000
Amazing team
World-class benefits
Staff Mojo Compiler Engineer
Staff Mojo Compiler Engineer

Modular, a Qualcomm company • United States

On-site
USD 216,000 - 372,000
RSU grants
Team onsite events
Excellent benefits
Inference Optimization Engineer United States - Remote · Remote
Inference Optimization Engineer United States - Remote · Remote

Modular Inc • United States

Remote
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
Senior Technical Community Manager
Senior Technical Community Manager

Modular, a Qualcomm company • United States

On-site
USD 150,000 - 225,000
World-class benefits
Competitive compensation
Team building events
Driver Engineer
Driver Engineer

Modular, a Qualcomm company • United States

On-site
USD 148,000 - 270,000
RSU grants
Team onsite events in Los Altos, CA
Travel 2–4 times per year
Senior AI Framework Engineer United States / Canada · Remote
Senior AI Framework Engineer United States / Canada · Remote

Modular Inc • Los Altos (CA), Northern (KY)

Hybrid
USD 180,000 - 324,000
RSU grants
Competitive compensation
Team on-sites in Los Altos
Inference Optimization Engineer
Inference Optimization Engineer

Modular • United States

On-site
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2