SSE - Optimization Engineer

MulticoreWare, Inc.

Annamayya

On-site

INR 1,500,000 - 2,100,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

MulticoreWare, Inc. is seeking a Senior Software Engineer to develop and optimize deep learning models for efficient inference across CPU, GPU, and edge devices.

You will implement and tune quantization techniques and model compression to boost latency and throughput. The role requires strong Python and C++ skills, hands-on experience with CNNs, Transformers, and LLMs, and deployment pipelines using PyTorch/ONNX.

Qualifications

  • BE/BTech/MS/MTech in Computer Science or related field with 4+ years of experience.
  • Strong programming skills in Python and C++.
  • Proven experience in Quantization algorithms (PTQ, QAT, GPTQ, AWQ).
  • Hands-on experience in pruning, model compression, and inference optimization.
  • Experience implementing quantization or optimization techniques from scratch.
  • Strong understanding of CNNs, Transformers, and LLM architectures.
  • Experience with PyTorch / ONNX and model deployment pipelines.
  • Strong problem-solving and performance optimization skills.

Responsibilities

  • Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, and hardware accelerators / edge devices.
  • Design and implement Quantization algorithms (PTQ, QAT, GPTQ, AWQ) from scratch.
  • Apply model compression techniques such as pruning, decomposition, and distillation.
  • Implement and optimize quantized kernels (INT8, INT4, FP8) using C++ for high performance.
  • Translate research papers into production-ready implementations.
  • Optimize latency, throughput, and memory usage for real-world deployment.
  • Work on transformer optimization including KV-cache, PEFT (LoRA/QLoRA), and MoE models.
  • Profile, benchmark, and debug model performance across different hardware platforms.
  • Collaborate with ML, compiler, and hardware teams to deliver optimized solutions.

Skills

Python
C++
Quantization
Model compression
Inference optimization
CNNs
Transformers
LLMs
PyTorch/ONNX

Education

BE/BTech/MS/MTech in Computer Science

Tools

PyTorch
ONNX
TensorRT
MLIR

Job description

MulticoreWare is a global software solutions & products company with its HQ in San Jose, CA, USA. With worldwide offices, it serves its clients and partners in North America, EMEA and APAC regions. Started by a group of researchers, MulticoreWare has grown to serve its clients and partners on HPC & Cloud computing, GPUs, Multicore & Multithread CPUS, DSPs, FPGAs and a variety of AI hardware accelerators.


MulticoreWare was founded by a team of researchers that wanted a better way to program for heterogeneous architectures. With the advent of GPUs and the increasing prevalence of multi-core, multi-architecture platforms, our clients were struggling with the difficulties of using these platforms efficiently.


We started as a boot-strapped services company and have since expanded our portfolio to span products and services related to compilers, machine learning, video codecs, image processing and augmented/virtual reality. Our hardware expertise has also expanded with our team; we now employ experts on HPC and Cloud Computing, GPUs, DSPs, FPGAs, and mobile and embedded platforms. We specialize in accelerating software and algorithms, so if your code targets a multi-core, heterogeneous platform, we can help.


Job Description

Title - Senior Software Engineer

Job Description

Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, and hardware accelerators / edge devices.



  • Design and implement Quantization algorithms (PTQ, QAT, GPTQ, AWQ) from scratch.

  • Apply model compression techniques such as pruning, decomposition, and distillation.

  • Implement and optimize quantized kernels (INT8, INT4, FP8) using C++ for high performance.

  • Translate research papers into production-ready implementations.

  • Optimize latency, throughput, and memory usage for real-world deployment.

  • Work on transformer optimization including KV-cache, PEFT (LoRA/QLoRA), and MoE models.

  • Profile, benchmark, and debug model performance across different hardware platforms.

  • Collaborate with ML, compiler, and hardware teams to deliver optimized solutions.


Must-Have


  • BE/BTech/MS/MTech in Computer Science or related field with 4+ years of experience.

  • Strong programming skills in Python and C++.

  • Proven experience in Quantization algorithms (PTQ, QAT, GPTQ, AWQ).

  • Hands-on experience in pruning, model compression, and inference optimization.

  • Experience implementing quantization or optimization techniques from scratch.

  • Strong understanding of CNNs, Transformers, and LLM architectures.

  • Experience with PyTorch / ONNX and model deployment pipelines.

  • Strong problem-solving and performance optimization skills.


Nice-to-Have


  • Experience with MoE architectures, and PEFT techniques (LoRA, QLoRA).

  • Knowledge of TensorRT, ONNX Runtime, TVM, MLIR.

  • Familiarity with hardware-aware optimization (GPU, NPU, Edge Devices).

  • Experience in research paper implementation or open-source contributions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Algorithms & optimization Engineer
Algorithms & optimization Engineer

MulticoreWare, Inc. • Chennai District

On-site
INR 1,200,000 - 2,100,000
AI Engineer Model Optimization & Acceleration
AI Engineer Model Optimization & Acceleration

Sunrise Biztech Systems • Bangalore Rural

On-site
INR 1,200,000 - 2,400,000
Tech Lead
Tech Lead

MulticoreWare, Inc. • Chennai District

On-site
INR 4,000,000 - 5,500,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 3,000,000 - 6,000,000
GPU Engineer
GPU Engineer

MulticoreWare, Inc. • Chennai District

On-site
INR 1,200,000 - 2,200,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

LeadSoc Technologies Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Edge AI deployment
Generative AI systems
Distributed inference pipelines
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

Mulya Technologies • Delhi

On-site
INR 4,000,000 - 7,000,000
Senior AI Software Performance Engineer
Senior AI Software Performance Engineer

BigStep Technologies • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)
Senior AI Scientist -Research Engineer, Applied AI ( at par with Director)

Mulya Technologies • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Hyderabad

On-site
INR 2,500,000 - 3,500,000