AI/GenAI Engineer (LLM Integration Specialist) _freelancer

SMARTnCODE Technology Private Limited

Hyderabad

On-site

INR 1,200,000 - 2,400,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SMARTnCODE Technology Private Limited is seeking an AI/GenAI Engineer to lead LLM integration, optimization, and domain customization for a high-performance chat application. You will connect multiple LLM APIs, implement robust wrappers, and design streaming real-time responses to support thousands of concurrent users.

You will collaborate with React and Python developers to create an intelligent chat experience, implementing prompt engineering, RAG pipelines, and cost-aware token usage while

Qualifications

  • Experience working with LLMs and GenAI technologies.
  • Proficiency with Python (FastAPI, LangChain, LlamaIndex preferred).
  • Strong prompt engineering techniques and best practices.
  • Experience with vector databases (Pinecone, Weaviate, Qdrant, ChromaDB).
  • Knowledge of RAG implementation.
  • Understanding of transformer architecture and attention mechanisms.
  • Experience with API integration and streaming responses.
  • Experience with LangChain, LlamaIndex, or similar LLM frameworks.
  • Knowledge of fine-tuning techniques (LoRA, QLoRA, PEFT).
  • Experience with embedding models and semantic search.

Responsibilities

  • LLM Integration & Architecture: connect APIs, robust wrappers, streaming responses, rate limiting, token/context management.
  • Domain Customization & Fine-tuning: domain prompts, RAG pipelines, LoRA/Prompt tuning, knowledge bases, few-shot learning.
  • Performance & Optimization: reduce latency, caching, token-cost optimization, A/B testing, model routing pipelines.
  • Safety & Quality: content moderation, guardrails, evaluation frameworks, hallucination handling, user feedback loops.
  • Infrastructure: scalable LLM architecture, queuing, monitoring/logging, DevOps deployment for self-hosted models.

Skills

LLMs & GenAI
OpenAI API
Python
Prompt engineering
Vector databases
RAG implementation
Transformer architecture
API integration
LangChain / LlamaIndex
Fine-tuning (LoRA/QLoRA)
Embedding models
Multi-modal models

Tools

LangChain
LlamaIndex
FastAPI
vLLM
TGI (Text Generation Inference)

Job description

Role & responsibilities
About the Role

We're building a high-performance chat application an AI/GenAI Engineer to lead the integration and optimization of Large Language Models (LLMs). You'll be responsible for connecting LLM APIs, implementing domain-specific fine-tuning strategies, prompt engineering, and ensuring optimal performance for production use.

You'll work closely with our React and Python developers to create a seamless, intelligent chat experience that serves thousands of concurrent users.

Key Responsibilities
LLM Integration & Architecture
  • Integrate multiple LLM APIs (OpenAI, Anthropic Claude, Google Gemini, or open-source models)
  • Design and implement robust API wrapper services with retry logic, fallback mechanisms, and error handling
  • Implement streaming responses for real-time chat experience
  • Build rate limiting and quota management systems
  • Handle token counting, context window management, and cost optimization
Domain Customization & Fine-tuning
  • Develop domain-specific prompt engineering strategies
  • Implement RAG (Retrieval Augmented Generation) pipelines using vector databases
  • Fine-tune or adapt models for specific use cases using techniques like LoRA, prompt tuning
  • Create and maintain knowledge bases for domain-specific responses
  • Implement few-shot learning and in-context learning strategies
Performance & Optimization
  • Optimize API response times and reduce latency
  • Implement caching strategies for common queries
  • Monitor and optimize token usage to control costs
  • A/B test different models and prompts for quality improvements
  • Build fallback chains (primary/secondary model routing)
Safety & Quality
  • Implement content moderation and safety filters
  • Build guardrails to prevent prompt injection and jailbreaking
  • Develop evaluation frameworks to measure response quality
  • Monitor and handle hallucinations and inaccuracies
  • Implement user feedback loops for continuous improvement
Infrastructure
  • Design scalable architecture for handling concurrent LLM requests
  • Implement queue systems for managing high-volume API calls
  • Set up monitoring and logging for LLM interactions
  • Work with DevOps to deploy models (if self-hosted)
Preferred candidate profile
Must Have:
  • Experience working with LLMs and GenAI technologies
  • Strong experience with OpenAI API, Anthropic Claude, or similar LLM APIs
  • Proficiency in Python (FastAPI, LangChain, LlamaIndex preferred)
  • Strong understanding of prompt engineering techniques and best practices
  • Experience with vector databases (Pinecone, Weaviate, Qdrant, ChromaDB)
  • Knowledge of RAG (Retrieval Augmented Generation) implementation
  • Understanding of transformer architecture and attention mechanisms
  • Experience with API integration, webhooks, and streaming responses
  • Strong problem-solving skills and ability to debug complex AI systems
  • Experience with LangChain, LlamaIndex, or similar LLM frameworks
  • Knowledge of fine-tuning techniques (LoRA, QLoRA, PEFT)
  • Experience with embedding models and semantic search
  • Familiarity with x library
  • Experience deploying models using vLLM, TGI (Text Generation Inference)
  • Knowledge of function calling/tool use with LLMs
  • Experience with model evaluation metrics (BLEU, ROUGE, BERTScore)
  • Understanding of token economics and cost optimisation
  • Experience with open-source models (Llama, Mistral, Falcon)
  • Knowledge of model quantization and optimization techniques
  • Experience with multi-modal models (vision, audio)
  • Familiarity with MLOps practices and experiment tracking (Weights & Biases, MLflow)
  • Experience with AWS SageMaker, Google Vertex AI, or Azure ML
  • Understanding of chain-of-thought prompting, ReAct, agents
  • Experience building chatbots or conversational AI systems
  • Publications or contributions to AI/ML community
Technical Stack You'll Work With
  • Languages: Python (primary), JavaScript/TypeScript (basic understanding)
  • LLM APIs: OpenAI, Anthropic, Google Gemini, Cohere
  • Frameworks: LangChain, LlamaIndex, FastAPI
  • Vector DBs: Pinecone, Weaviate, Qdrant, or ChromaDB
  • Infrastructure: Docker, Redis, PostgreSQL, Message Queues
  • Cloud: AWS/GCP/Azure (whatever your team uses)
  • Monitoring: Prometheus, Grafana, custom LLM analytics
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Engineer
Senior AI Engineer

Mallow Technologies • Karur

On-site
INR 1,800,000 - 2,400,000
Competitive salary
Career growth
Cutting-edge AI projects
+1
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
AI Engineer
AI Engineer

Anarock Property Consultants Pvt Ltd • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Lead AI Engineer
Lead AI Engineer

FINVASIA CAREER • Nagar

On-site
INR 1,500,000 - 2,100,000
Data Scientist
Data Scientist

Neurealm • Gurugram District

On-site
INR 1,200,000 - 2,000,000
Lead AI Engineer
Lead AI Engineer

Bridge-it • New Delhi

Remote
INR 11,527,000 - 17,291,000
AI Developer
AI Developer

Bot Jobs • Bengaluru

Remote
INR 1,000,000 - 1,200,000
Analytics Engineer-Senior AI / RAG Platform Engineer
Analytics Engineer-Senior AI / RAG Platform Engineer

Trigyn Technologies Limited. • Delhi

On-site
INR 2,000,000 - 3,500,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

Dailoqa • Dadri

On-site
INR 2,000,000 - 4,000,000
AI - Engineer
AI - Engineer

CorroHealth, Inc. • India

On-site
INR 1,400,000 - 2,100,000