Senior AI Runtime Engineer

Socket.dev

United States

Hybrid

USD 216,000 - 324,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

RSU grants
Comprehensive healthcare
Paid time off
Team onsite events

Job summary

Modular is building an AI platform to enable developers to deploy state-of-the-art models across hardware. As an AI Runtime Engineer, you will own a runtime that runs on CPU and GPU platforms, optimizing performance for diverse customer models.

Locations: candidates in the US or Canada may apply. You can work in our Los Altos, CA office or remotely from home; onboarding for new hires is conducted in person in Los Altos.

Qualifications

  • 5+ years of experience working on high-performance computing systems.
  • Experience in C++ programming and complex software systems.
  • Experience with CPU or GPU runtime optimizations and performance analysis.
  • Proficiency with profiling tools (CPU or GPU).
  • Creativity and curiosity for solving complex problems with a team-oriented mindset.

Responsibilities

  • Design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency.
  • Port the Modular runtime stack to new hardware platforms and develop an API.
  • Collaborate with compiler, kernels, serving, and models teams to achieve end-to-end performance.
  • Collaborate with customer success to understand performance requirements and use cases.
  • Collaborate with tooling and infrastructure teams to design automated performance analysis and benchmarking.

Skills

5+ years HPC
C++ programming
CPU/GPU runtime tuning
Profiling tools
Team collaboration

Job description

About the role:

ML developers today face significant friction when deploying trained models. They work in a fragmented space with incomplete, patchwork solutions that require extensive performance tuning and model-specific optimizations. At Modular, we are building the next-generation AI platform that will radically improve how developers build and deploy AI models. A core part of this offering is a platform that enables customers to achieve state-of-the-art performance across model families and frameworks.

As an AI Runtime Engineer, you will own a runtime that operates on various CPU and GPU hardware platforms, optimizing performance for diverse customer AI models. LOCATION: Candidates based in the US or Canada are welcome to apply. You can work in our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office.

What you will do:
  • Design and develop runtime and cross-stack optimizations to improve CPU and GPU efficiency, addressing issues such as CPU overhead, caching, and data locality across multiple devices.
  • Port the Modular runtime stack to new hardware platforms and develop an API to streamline this process.
  • Collaborate with the compiler, kernels, serving, and models teams to design core technologies that achieve state-of-the-art end-to-end performance on various CPU and GPU hardware.
  • Collaborate with the customer success team and engage with customers to understand their performance requirements and use cases.
  • Collaborate with tooling and infrastructure teams to design systems for automated performance analysis and benchmarking.
What you bring to the table:
  • 5+ years of experience working on high-performance computing systems.
  • Experience in C++ programming and complex software systems.
  • Experience with CPU or GPU runtime optimizations and performance analysis on CPUs, GPUs, or AI accelerators.
  • Proficiency with one or more profiling tools (CPU or GPU).
  • Creativity and curiosity for solving complex problems, a team-oriented attitude that enables you to work well with others, and alignment with our culture.
Helpful, but not required:
  • Experience with ML graph optimizations, parallel / distributed programming, heterogeneous ML computation, and/or code generation.
  • Exposure to MLIR, LLVM, and/or the Mojo programming language.
  • Advanced degree in Computer Science or a related area is a plus.
What Modular brings to the table:
  • Amazing Team.We are a progressive and agile team with some of the industry’s best engineering and product leaders.
  • World-class Benefits.In order to attract the best, we need to offer the best. Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities. Please note that specific benefit packages may vary based on your location, you can read more about benefits offered by Qualcomm here.
  • Competitive Compensation.We offer very strong compensation packages, including RSU grants. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce.
  • Team Building Events.We organize regular team onsites and local meetups in Los Altos, CA as well as different cities. Traveling 2-4 times a year is expected for all roles.

Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and a purpose to truly change the world.

The estimated base salary range for this role to be performed in the US, regardless of the state, is $216,000.00 - $324,000.00 USD.

The estimated base salary range for this role to be performed in Canada, regardless of the province, is $207,000.00 - $310,400.00 CAD.

The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of your total compensation. For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply as we may have upcoming openings that are lower/higher level than the ones advertised.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Manager, Cloud United States - Remote · Remote
Product Manager, Cloud United States - Remote · Remote

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 223,000 - 345,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Senior AI Kernel Engineer
Senior AI Kernel Engineer

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Inference Optimization Engineer
Inference Optimization Engineer

Modular • United States

Hybrid
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2
Developer Advocate, MAX Inference & Serving
Developer Advocate, MAX Inference & Serving

Modular • United States

Hybrid
USD 150,000 - 225,000
Premier insurance
401k matching
Flexible PTO
+1
Inference Optimization Engineer United States - Remote · Remote
Inference Optimization Engineer United States - Remote · Remote

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
Senior AI Runtime Engineer - High-Performance Platform
Senior AI Runtime Engineer - High-Performance Platform

Socket.dev • United States

Hybrid
USD 216,000 - 324,000
RSU grants
Comprehensive healthcare
Paid time off
+1
Senior Open Source Community Engineer
Senior Open Source Community Engineer

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premium health insurance
401k matching
Flexible PTO
+1
Developer Advocate, MAX Inference & Serving
Developer Advocate, MAX Inference & Serving

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Amazing team
World-class benefits
Stock options
+1
Software Engineer, Inference Infrastructure
Software Engineer, Inference Infrastructure

Modular • United States

On-site
USD 167,000 - 273,000
Competitive salary
Premier insurance plans
Flexible paid time off
+1
Developer Advocate, Mojo
Developer Advocate, Mojo

Modular, a Qualcomm company • United States

Hybrid
USD 150,000 - 225,000
Premier insurance plans
Stock options
401k matching (up to 5%)
+1