About the Role
As an LLM Engineer at Fospe, you will be at the forefront of designing, fine-tuning, and operationalizing cutting-edge Large Language Models and Retrieval-Augmented Generation (RAG) pipelines. You will collaborate directly with our AI research team to embed deterministic cognitive capabilities and high-throughput semantic reasoning into our vertical enterprise SaaS products.
Key Responsibilities
- Architect, fine-tune, and quantize open-source foundation models (Llama 3, Mistral, Qwen, DeepSeek) for domain-specific enterprise workloads.
- Design high-accuracy Retrieval-Augmented Generation (RAG) architectures with hybrid dense/sparse search, reranking, and semantic chunking.
- Build low-latency streaming inference pipelines utilizing vLLM, TensorRT-LLM, and Triton Inference Server.
- Implement guardrail validation, safety filters, prompt routing, and automated hallucination benchmarking.
- Collaborate with backend engineers to expose scalable gRPC and REST interfaces for seamless frontend consumption.
Requirements & Qualifications
- Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Data Science, or related quantitative field.
- 3+ years of professional experience in machine learning and deep learning, with 1+ years dedicated to LLM engineering.
- Deep hands-on expertise with PyTorch, Hugging Face Transformers, vLLM, and LangChain/LlamaIndex.
- Proven experience with vector databases such as Pinecone, Qdrant, Milvus, or pgvector.
- Strong proficiency in Python, asynchronous programming, Docker containerization, and Linux/CUDA environments.
- Demonstrated understanding of model optimization techniques including LoRA, QLoRA, AWQ, and GPTQ quantization.
Preferred Qualifications
- Experience with multi-agent orchestration frameworks (LangGraph, CrewAI, AutoGen).
- Contributions to open-source AI libraries or published research in NLP/LLMs.
- Experience deploying models on Kubernetes and cloud infrastructure (AWS/GCP/Azure GPU instances).
Primary Technologies
Python, PyTorch, Hugging Face, vLLM, LangChain, LlamaIndex, Qdrant, Docker, Kubernetes, CUDA, FastAPI
What We Offer
- Competitive compensation package with performance bonuses
- Flexible hybrid working schedule (Bangalore Innovation Center)
- Comprehensive health, dental, and wellness insurance coverage
- Generous annual compute and AI research learning stipend
- Latest Apple MacBook Pro M-series workstation & equipment budget
- Stock option and co-ownership opportunities