LLM Engineer

Zorba Consulting

Hyderabad

On-site

INR 2,500,000 - 4,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

On-site role with GPU access
Cross-functional product teams

Job summary

Zorba Consulting in Hyderabad, India, is hiring an experienced LLM Engineer to design, fine‑tune, and deploy production‑ready LLM solutions powering search, summarization, agents, and domain‑specific assistants.

You will build RAG pipelines, optimize inference, and develop backend services with Docker/Kubernetes, collaborating with data engineering and ML teams to ensure reliability and governance.

Qualifications

  • 4+ years of hands-on experience working with LLMs or advanced NLP models in production contexts.
  • Proficiency in Python for ML engineering and model development.
  • Experience with PyTorch and Hugging Face Transformers for training and fine-tuning.
  • Practical experience implementing RAG and vector search using tools such as FAISS or similar vector databases.
  • Familiarity with LangChain (or equivalent orchestration) and integration with LLM APIs (OpenAI, Anthropic, etc.).
  • Experience containerizing and deploying ML services using Docker; familiarity with Kubernetes is a plus.

Responsibilities

  • Design, fine‑tune, and validate LLMs for production use‑cases: Instruction tuning, supervised fine‑tuning, and parameter‑efficient tuning (LoRA/adapters).
  • Implement retrieval‑augmented generation (RAG) pipelines: embeddings, vector search, chunking, and context assembly for high‑recall responses.
  • Optimize inference for latency and cost: quantization, model pruning, batching, and deployment with optimized runtimes (CUDA, Triton, bitsandbytes where applicable).
  • Build backend services and APIs to serve LLM inference and orchestration using containerized deployments (Docker/Kubernetes) and CI/CD pipelines.
  • Collaborate with product, data engineering, and ML teams to integrate LLMs into production flows, monitor model performance, and set up automated retraining/rollbacks.
  • Create reproducible training pipelines, implement evaluation suites, and produce documentation and runbooks for model governance and observability.

Skills

LLMs in production
Python
PyTorch
Transformers
RAG
Vector search
LangChain
Docker
Kubernetes
LLM APIs

Tools

FAISS
bitsandbytes
CUDA
Triton

Job description

A leading consulting firm operating in the Enterprise Generative AI and Large Language Model (LLM) services sector, delivering production‑grade LLM solutions, retrieval‑augmented systems, and custom generative AI products for enterprise clients across domains. The team focuses on building secure, scalable, low‑latency inference services and automating model lifecycle workflows for on‑prem and cloud deployments.

Position

LLM Engineer - On‑site (India). We are hiring an experienced LLM engineer to design, fine‑tune, and deploy LLM‑based solutions that power search, summarization, agents, and domain‑specific assistants.

Role & Responsibilities
  • Design, fine‑tune, and validate LLMs for production use‑cases: Instruction tuning, supervised fine‑tuning, and parameter‑efficient tuning (LoRA/adapters).
  • Implement retrieval‑augmented generation (RAG) pipelines: embeddings, vector search, chunking, and context assembly for high‑recall responses.
  • Optimize inference for latency and cost: quantization, model pruning, batching, and deployment with optimized runtimes (CUDA, Triton, bitsandbytes where applicable).
  • Build backend services and APIs to serve LLM inference and orchestration using containerized deployments (Docker/Kubernetes) and CI/CD pipelines.
  • Collaborate with product, data engineering, and ML teams to integrate LLMs into production flows, monitor model performance, and set up automated retraining/rollbacks.
  • Create reproducible training pipelines, implement evaluation suites, and produce documentation and runbooks for model governance and observability.
Skills & Qualifications

Must‑Have :

  • 4+ years of hands‑on experience working with LLMs or advanced NLP models in production contexts.
  • Proficiency in Python for ML engineering and model development.
  • Experience with PyTorch and Hugging Face Transformers for training and fine‑tuning.
  • Practical experience implementing RAG and vector search using tools such as FAISS or similar vector databases.
  • Familiarity with LangChain (or equivalent orchestration) and integration with LLM APIs (OpenAI, Anthropic, etc.).
  • Experience containerizing and deploying ML services using Docker; familiarity with Kubernetes is a plus.

Preferred :

  • Experience with inference optimizations: quantization (bitsandbytes), Triton, or GPU accelerated serving.
  • Exposure to distributed training frameworks (DeepSpeed) and cloud MLOps platforms (SageMaker, Azure ML, GCP AI Platform).
  • Knowledge of monitoring, logging, and model‑evaluation frameworks for production LLMs (MLflow, Prometheus, Grafana).
Benefits & Culture Highlights
  • Collaborative, engineering‑driven culture with strong focus on ownership and rapid iteration.
  • Opportunity to build end‑to‑end LLM products for enterprise clients and influence architecture decisions.
  • On‑site role with hands‑on access to GPU infrastructure and cross‑functional product teams.

Skills : pytorch, cuda, docker, python, agentic, llm

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Engineer
LLM Engineer

Zorba Consulting • Pune District

On-site
INR 1,500,000 - 2,100,000
GPU access on-site
Cross-functional product teams
Ownership and rapid iteration
Senior AI/ML Engineer
Senior AI/ML Engineer

Keka Technologies Private Limited • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Mentorship from senior architects
Innovation-driven environment
Continuous learning opportunities
LLM Specialist
LLM Specialist

TRDFIN Support Services Pvt Ltd • Gurugram District

On-site
INR 1,000,000 - 1,500,000
LLM Engineer (Large Language Models)
LLM Engineer (Large Language Models)

Fospe UK Ltd • Bengaluru, Ernakulam

On-site
INR 800,000 - 1,200,000
Competitive salary package
Opportunity to work on cutting-edge AI technologies
Career growth in AI Product Engineering
+1
Senior AI / LLM Engineer
Senior AI / LLM Engineer

Thinkscoop Technologies • Bengaluru

Hybrid
INR 3,000,000 - 5,000,000
Remote-friendly culture
Learning budget
Performance bonuses
+1
AI Engineer (LLM / Generative AI)
AI Engineer (LLM / Generative AI)

HCLTech • Chennai District

Hybrid
INR 1,800,000 - 2,400,000
Generative AI Engineer
Generative AI Engineer

HCLTech • Chennai District

On-site
INR 1,800,000 - 3,600,000
AI Engineer (LLM / Generative AI/Ops)
AI Engineer (LLM / Generative AI/Ops)

HCLTech • Dadri

Hybrid
INR 1,500,000 - 4,000,000
Data Science – Gen AI / LLM
Data Science – Gen AI / LLM

Greytip Software Private Limited • Hyderabad

On-site
INR 2,500,000 - 3,800,000
AL/LLM Engineer with Python
AL/LLM Engineer with Python

Avirasoft Digital Gcc • Hyderabad

Hybrid
INR 3,000,000 - 5,000,000