AI model optimization & acceleration Engineer

L&T Technology Services

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

L&T Technology Services in Bengaluru seeks an AI Engineer to optimize and deploy ML models across CPU, GPU, and NPU, delivering scalable AI systems for robotics, healthcare, and automotive.

The role emphasizes production readiness, latency optimization, and hardware acceleration, with 4-10 years of experience in Python/C++, and a strong foundation in transformers and CNNs.

Qualifications

  • Proficient in Python and C++.
  • Experience with GPU/hardware acceleration (CUDA/ROCm or similar).
  • Solid understanding of deep learning models (transformers, CNNs).
  • Knowledge of optimization, quantization, and performance tuning.

Responsibilities

  • Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech.
  • Deploy on hardware accelerators (GPU/NPU) and optimize performance.
  • Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion).
  • Apply quantization and model compression (FP32 to lower precision).
  • Profile and debug system and model performance.

Skills

Python
C++
Transformers
CNNs
Inference optimization

Tools

CUDA
ROCm

Job description

Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU).

Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive.

Experience : 4-10 Years

Job Responsibilities / Day-to-Day Activities
Qualifications & Experiences:
Key Responsibilities
  • Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech
  • Deploy on hardware accelerators (GPU/NPU) and optimize performance
  • Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion)
  • Apply quantization and model compression (FP32 to lower precision)
  • Profile and debug system and model performance
Required Skills
  • Proficient in Python and C++
  • Experience with GPU/hardware acceleration (CUDA/ROCm or similar)
  • Solid understanding of deep learning models (transformers, CNNs)
  • Knowledge of optimization, quantization, and performance tuning
Good to Have
  • Edge AI or embedded deployment
  • Distributed inference or streaming pipelines
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer Model Optimization & Acceleration
AI Engineer Model Optimization & Acceleration

Sunrise Biztech Systems • Bangalore Rural

On-site
INR 1,200,000 - 2,400,000
AI/ML Engineer-1003
AI/ML Engineer-1003

Sunrise Biztech Systems • Bengaluru

On-site
INR 2,500,000 - 4,500,000
AI Benchmarking & Performance Engineer
AI Benchmarking & Performance Engineer

L&T Technology Services • Bengaluru

On-site
INR 4,200,000 - 5,400,000
AI Engineer – Model Optimization & Acceleration
AI Engineer – Model Optimization & Acceleration

AMD • Bengaluru

On-site
INR 2,500,000 - 4,500,000
AI Engineer – Model Optimization & Acceleration
AI Engineer – Model Optimization & Acceleration

AMD • Bengaluru Urban

On-site
INR 2,000,000 - 2,800,000
AI Engineer – Model Optimization & Acceleration
AI Engineer – Model Optimization & Acceleration

Advanced Micro Devices • Bengaluru

On-site
INR 1,600,000 - 2,400,000
AI Benchmarking & Performance Engineer
AI Benchmarking & Performance Engineer

Sunrise Biztech Systems • Bangalore Rural

On-site
INR 1,400,000 - 2,300,000
AI Engineer
AI Engineer

Edgecore Networks Corporation • Bengaluru

On-site
INR 1,500,000 - 2,800,000
AIML Engineer
AIML Engineer

Tekskills • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Senior AI Engineer
Senior AI Engineer

Epam Systems • Chennai District, Coimbatore District

On-site
INR 3,000,000 - 6,000,000