AI/DSP Kernel Optimization Engineer

Zoho

Chennai District

On-site

INR 1,200,000 - 1,800,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

MulticoreWare, a global software solutions company, is seeking an AI/DSP Kernel Optimization Engineer to develop and optimize AI inference kernels and improve overall performance. You will work on low-level C/C++ optimization, SIMD/vectorization, and profiling, collaborating with hardware and software teams to push efficiency across platforms.

The role requires a strong foundation in C/C++, AI/DL basics, and experience with DSPs or hardware accelerators.

Qualifications

  • B.E., B.Tech., M.E., M.Tech. or equivalent degree.
  • Strong knowledge of C/C++ programming.
  • Experience with DSPs or hardware accelerators.
  • Knowledge of SIMD/vectorization, intrinsics, and low-level optimization.
  • Good profiling and debugging skills.

Responsibilities

  • Develop and optimize AI inference kernels using C/C++.
  • Optimize kernels for data types such as INT8, BF16, FP32.
  • Perform SIMD/vectorization and memory optimization.
  • Profile, debug, and improve inference performance.
  • Integrate kernels into AI inference pipelines with hardware teams.

Skills

C/C++ programming
AI/Deep Learning basics
Performance profiling

Education

B.E., B.Tech., M.E., M.Tech., or equivalent

Tools

DSPs
Hardware accelerators
SIMD intrinsics

Job description

MulticoreWare is a global software solutions & products company with its HQ in San Jose, CA, USA. With worldwide offices, it serves its clients and partners in North America, EMEA and APAC regions. Started by a group of researchers, MulticoreWare has grown to serve its clients and partners on HPC & Cloud computing, GPUs, Multicore & Multithread CPUS, DSPs, FPGAs and a variety of AI hardware accelerators.

MulticoreWare was founded by a team of researchers that wanted a better way to program for heterogeneous architectures. With the advent of GPUs and the increasing prevalence of multi-core, multi-architecture platforms, our clients were struggling with the difficulties of using these platforms efficiently.

We started as a boot-strapped services company and have since expanded our portfolio to span products and services related to compilers, machine learning, video codecs, image processing and augmented/virtual reality. Our hardware expertise has also expanded with our team; we now employ experts on HPC and Cloud Computing, GPUs, DSPs, FPGAs, and mobile and embedded platforms. We specialize in accelerating software and algorithms, so if your code targets a multi-core, heterogeneous platform, we can help.

Job Description

We are lookingfor an AI/DSP Kernel Optimization Engineer to develop and optimize AI inferencekernels, analyze application performance, and integrate optimized kernels intoAI inference pipelines. The role involves low-level C/C++ optimization,SIMD/vectorization, profiling, debugging, and close collaboration with hardwareand software teams to improve overall inference performance.

Responsibilities:

  • Develop and optimize AI inference kernels usingC/C++.
  • Optimize kernels for different data types,including INT8, INT16, BF16, and FP32.
  • Perform SIMD/vectorization, intrinsics, memoryoptimization, and performance tuning.
  • Profile and analyze application performance toidentify optimization opportunities.
  • Debug and resolve functional and performanceissues.
  • Integrate optimized kernels into the AI inferencepipeline.
  • Work closely with hardware and software teams to improve overall inference performance.
Requirements

Education:

B.E.,B.Tech., M.E., M.Tech., or equivalent.

Technical Skills (Must haves):

  • Good knowledge of C/C++ programming.
  • Basic understanding of AI/Deep Learning modelsand inference.
  • Experience with DSPs, hardware accelerators, orsimilar compute platforms.
  • Knowledge of SIMD/vectorization, intrinsics, andlow-level software optimization.
  • Good profiling, debugging, and performanceanalysis skills.

Need to have (Can be bridged):

  • Understanding of computer architecture and memoryoptimization.

Good to have (Not essential):

  • Experience in AI/ML inference optimization.
  • Knowledge of different numerical data typessuchas INT8, INT16, BF16, and FP32.
  • Familiarity with performance profiling andbenchmarking tools.

Preferred Qualifications (Optional):

  • Experience with C7x DSP, HWA-MMA, or similarhardware accelerators.
  • Hands-on experience optimizing inferenceworkloads for DSPs or specialized compute platforms.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

MulticoreWare, Inc. • Chennai District

On-site
INR 900,000 - 1,400,000
Software Engineer
Software Engineer

Zoho • Chennai District

On-site
INR 1,200,000 - 1,800,000
Edge AI Engineer
Edge AI Engineer

Vedya Labs • Hyderabad, Bengaluru

On-site
INR 900,000 - 1,500,000
Compute Optimization Engineer (DSP/Kernels)
Compute Optimization Engineer (DSP/Kernels)

Qualcomm • Bengaluru

On-site
INR 800,000 - 1,200,000
AI20P Library Engineer, Machine Learning Acceleration
AI20P Library Engineer, Machine Learning Acceleration

Qpisemi • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Software Engineer – NPU/Hexagon DSP Kernel Optimization
AI Software Engineer – NPU/Hexagon DSP Kernel Optimization

Qualcomm • Bengaluru

On-site
INR 800,000 - 1,500,000
Principal Software Architect- High Performance Computing
Principal Software Architect- High Performance Computing

Applied Materials India • Chennai District

On-site
INR 4,000,000 - 6,000,000
Sr. Software Engineer – AI Operators
Sr. Software Engineer – AI Operators

NXP Semiconductors • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Kernel Engineer
Kernel Engineer

Cerebras Systems, Inc. • India

On-site
INR 1,200,000 - 2,000,000
Open source cutting-edge AI research
Job stability with startup vitality
Inclusive work culture
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • Mumbai

On-site
INR 1,653,000 - 3,306,000