Get more replies from employers
Send a job-specific resume in minutes.
EXL is seeking an AI/ML Engineer to fine-tune open-source LLMs, build AI agents, and implement hybrid LLM solutions. You will develop Python APIs with FastAPI, containerize with Docker, and deploy on AWS/Azure while collaborating with senior engineers to optimize performance.
The role requires hands-on CUDA toolkits experience, familiarity with LangChain-like frameworks, and strong knowledge of vector databases and cloud deployments.
Fine-tune and adapt open-source LLMs (e.g., LLaMA 4, Mistral and Bert) using NVIDIA GPU tools. Build AI agents using frameworks like LangChain, LangGraph, or AutoGen with structured workflows (memory, tools, retries, etc.). Implement hybrid LLM solutions using OpenAI/Claude APIs and open-source models. Develop APIs using FastAPI and containerize apps with Docker. Deploy, monitor, and scale AI solutions on AWS, Azure, or similar cloud providers. Collaborate with senior engineers to optimize performance and reliability of deployed systems.
Hands-on experience with LLM fine-tuning and NVIDIA GPU toolkits (CUDA). Familiarity with LangChain or similar agent frameworks. Experience developing APIs with FastAPI and deploying via Docker. Proficiency in using OpenAI/Anthropic APIs and building basic RAG pipelines. Solid foundation in Python, cloud deployment (AWS/Azure), and vector databases (e.g., FAISS, Pinecone). Nice to have: Exposure to tools like LangServe and Semantic Kernel Familiarity with CI/CD pipelines and monitoring tools (e.g., GitHub Actions, Prometheus). Contribution to open-source AI/ML projects.
General Shift: 1:30 PM to 11:30 PM IST. Flexibility to extend hours based on critical deployments or support needs.