Role Overview
We are looking for a hands‑on AI Application Engineer to build and ship the GenAI/SLM applications designed by our Technical Program Leads — including for air-gapped and on-prem environments. RAG pipelines, agents, fine‑tuned models, and the APIs/UI that expose them.
Location: Noida
Type: Full-Time, Permanent
Experience: 5+ years
Role Overview
We are looking for a hands‑on AI Application Engineer to build and ship the GenAI/SLM applications designed by our Technical Program Leads — including for air-gapped and on-prem environments. RAG pipelines, agents, fine‑tuned models, and the APIs/UI that expose them.
Key Responsibilities
- Build and productionize RAG pipelines, agentic workflows, and LLM/SLM‑backed features from architecture specs handed off by the Technical Program Lead.
- Fine‑tune, quantize, and package SLMs for constrained/offline environments; benchmark accuracy, latency, and cost against alternatives.
- Implement local/offline inference serving (vLLM, Ollama) and vector store integrations (FAISS, Milvus, Weaviate, Qdrant) for air‑gapped deployments.
- Write clean, testable, well‑documented Python — APIs, data pipelines, and integration layers connecting LLM components to enterprise systems.
- Containerize and deploy applications (Docker/Kubernetes) across cloud (AWS/Azure/GCP) and on‑prem targets.
- Build evaluation harnesses, guardrails, and monitoring/logging for model outputs in line with the governance framework set by the Technical Program Lead.
- Work sprint‑to‑sprint in JIRA — pick up stories, raise blockers early, keep the board current, and demo working software each sprint.
Required Skills & Experience
- 5+ years professional software engineering; 2+ years building GenAI/ML applications in production.
- Strong Python; hands‑on with LangChain, LlamaIndex, or similar frameworks.
- Practical experience with LLMs/SLMs — prompting, RAG, fine‑tuning (LoRA/QLoRA), or model quantization.
- Working knowledge of vector databases and embedding pipelines.
- Comfortable with Docker/Kubernetes and at least one major cloud (AWS/Azure/GCP).
- Expert with Claude‑driven development — uses Claude Code / Claude‑based agents daily as part of the build workflow; comfortable authoring or using custom Skills/MCP tools to speed up delivery.
- Reviewer, not just implementer: most code is agent‑generated first; your core skill is writing tight specs, critically reviewing agent output line‑by‑line, catching bugs/edge cases/security issues, and deciding when to trust vs. override the agent — rather than manually writing everything from scratch.
- Solid understanding of REST/API design, git workflows, and CI/CD basics.
Behavioural Expectations
- Execution-focused: comfortable taking a spec from the Technical Program Lead and running with it with minimal hand‑holding — but "execution"; here means directing and reviewing agentic output, not manual coding for its own sake.
- Fluent in Agile/Scrum — active participant in ceremonies, disciplined about JIRA hygiene and sprint commitments.
- Clear communicator — flags risks/blockers early, documents decisions, and can explain technical trade‑offs to the Technical Program Lead and, when needed, the client.
- Mentors junior AI Application Engineers — reviews their code/PRs, helps them write better specs for AI coding agents, and brings them up to speed on RAG/SLM patterns and air‑gapped deployment practices.
- Self‑driven and self‑governed, per company's high‑ownership hybrid culture.
- Mentor juniors
Preferred: background in an IT/consulting services company.
Good to Have
- Exposure to Big Data tooling (Spark/Hive/Hadoop) or Graph Analytics.
- Experience in a regulated or air‑gapped delivery environment (defense, government, BFSI).
- Familiarity with AI governance/evaluation frameworks (guardrails, red‑teaming, model cards).
- Contribution to open source projects, academic papers published, filled patents.