Edge LLM Engineer

Flairdeck

Bengaluru

On-site

INR 2,000,000 - 3,600,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Flairdeck in Bengaluru, Karnataka invites an engineer to deploy and optimize Large Language Models on edge devices, preferably NVIDIA Jetson platforms. You will design practical LLM-based solutions using prompt engineering, preprocessing, caching, data creation, and occasional model fine-tuning to meet real-world constraints.

The role emphasizes reliable and efficient LLM applications under limited compute, memory, latency, and power on edge devices, with opportunities to work on end-to-end

Qualifications

  • HHands-on experience with LLMs, prompt engineering, and scenario-specific prompt design.
  • Experience running AI/ML models on edge devices with compute and memory constraints.
  • Practical knowledge of preprocessing techniques for text, speech transcripts, and structured inputs.
  • Experience implementing caching, context management, and optimization techniques for LLM applications.
  • Ability to create datasets and fine-tune or adapt models for domain-specific use cases.
  • Strong Python programming skills.
  • Understanding of NLP tasks such as intent handling, entity extraction, and text classification.
  • Experience with model evaluation, latency optimization, and debugging AI behavior.
  • Familiarity with NVIDIA Jetson or similar edge AI platforms.

Responsibilities

  • Design practical LLM-based solutions for real-world scenarios using prompt engineering, input preprocessing, caching strategies, data creation, and model fine-tuning when required.
  • Understand how to make LLM applications reliable, efficient, and context-aware under edge-device constraints such as limited compute, memory, latency, and power.

Skills

LLMs
Prompt design
Edge devices
Text preprocessing
Caching strategies
Dataset creation
Model fine-tuning
Python programming
NLP tasks
Latency optimization
AI debugging
Jetson

Tools

TensorRT
ONNX
PyTorch
Hugging Face

Job description

Job Summary

We are looking for an engineer experienced in deploying and optimizing Large Language Models on edge devices, preferably NVIDIA Jetson platforms. The role is not limited to model inference; the candidate should be able to design practical LLM-based solutions for real-world scenarios using prompt engineering, input preprocessing, caching strategies, data creation, and model fine-tuning when required. The ideal candidate should understand how to make LLM applications reliable, efficient, and context-aware under edge-device constraints such as limited compute, memory, latency, and power.


Responsibilities

Design practical LLM-based solutions for real-world scenarios using prompt engineering, input preprocessing, caching strategies, data creation, and model fine-tuning when required. Understand how to make LLM applications reliable, efficient, and context-aware under edge-device constraints such as limited compute, memory, latency, and power.


Mandatory Skills


  • Hands-on experience with LLMs, prompt engineering, and scenario-specific prompt design.

  • Experience running AI/ML models on edge devices with compute and memory constraints.

  • Practical knowledge of preprocessing techniques for text, speech transcripts, and structured inputs.

  • Experience implementing caching, context management, and optimization techniques for LLM applications.

  • Ability to create datasets and fine-tune or adapt models for domain-specific use cases.

  • Strong Python programming skills.

  • Understanding of NLP tasks such as intent handling, entity extraction, and text classification.

  • Experience with model evaluation, latency optimization, and debugging AI behavior.

  • Familiarity with NVIDIA Jetson or similar edge AI platforms.


Good to Have Skills


  • Experience with speech processing, speech-to-text systems, and audio preprocessing.

  • Knowledge of noise handling, speech enhancement, and robust voice input pipelines.

  • Experience with NER models and entity extraction pipelines.

  • Familiarity with TensorRT, ONNX, PyTorch, Hugging Face, or similar model deployment tools.

  • Experience with quantization, pruning, distillation, or other model compression techniques.

  • Knowledge of retrieval-augmented generation, vector databases, or local knowledge caching.

  • Experience building real-time AI applications on embedded Linux systems.

  • Familiarity with multilingual or domain-specific language processing.

  • Experience integrating LLMs with sensors, robotics, industrial systems, or IoT devices.


Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Principal Local Llm & Generative Ai Platform Engineer
Senior/Principal Local Llm & Generative Ai Platform Engineer

Parallelwireless • Maharashtra

On-site
INR 3,000,000 - 5,500,000
Senior AI Engineer – LLM Systems
Senior AI Engineer – LLM Systems

Evollabs Tech • Anupgarh

On-site
INR 4,000,000 - 6,000,000
Prismforce Pvt Ltd - AI Engineer - LLM/RAG
Prismforce Pvt Ltd - AI Engineer - LLM/RAG

Prismforce • Maharashtra

On-site
INR 1,800,000 - 3,200,000
Machine Learning Engineer
Machine Learning Engineer

Tranzeal • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Lead Engineer - AI/ML
Lead Engineer - AI/ML

Mindfire Solutions • India

On-site
INR 2,000,000 - 3,000,000
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)

Solutions By Text • Bengaluru

On-site
INR 2,500,000 - 4,000,000
LLM Engineer (Large Language Models)
LLM Engineer (Large Language Models)

Fospe UK Ltd • Bengaluru

Hybrid
INR 2,500,000 - 5,200,000
Competitive compensation with bonuses
Hybrid work at Bangalore Innovation Cn
Health, dental, wellness insurance
+3
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Bhavitha Tech, CMMi Level 3 Company • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Inference Server Engineer
Inference Server Engineer

Evollabs Tech • Anupgarh

On-site
INR 3,000,000 - 6,000,000