Get more replies from employers
Send a job-specific resume in minutes.
Infosys is looking for a professional skilled in deploying and managing LLMs like OpenAI and Mistral in production environments. The role involves building scalable inference pipelines, integrating LLMs into applications, and implementing prompt engineering and retrieval-augmented generation (RAG).
The successful candidate will have expertise in managing LLM performance, and applying MLOps for version control and experiment tracking. Strong Python programming and experience with frameworks like LangChain are essential. Join us to operationalize Generative AI applications!
Deploy and manage LLMs (OpenAI, Llama, Mistral, etc.) in production environments.
Build scalable inference pipelines (real‑time & batch).
Integrate LLMs into applications via APIs and microservices.
Design and implement end‑to‑end LLM pipelines: prompt engineering, retrieval‑augmented generation (RAG), fine‑tuning / embeddings, using frameworks such as LangChain, LangGraph, LlamaIndex.
Build and optimize RAG pipelines using vector databases (Pinecone, FAISS, Weaviate, Chroma), handle document ingestion, chunking, indexing, and retrieval.
Monitor LLM performance: latency, accuracy / hallucinations, cost efficiency.
Implement prompt optimization, feedback loops, guardrails & evaluation frameworks.
MLOps for LLMs: build CI/CD pipelines for model updates, prompt/version control, experiment tracking, deployments; ensure reproducibility of LLM workflows.