About The Role
The role focuses on building production LLM systems - not demos. Expect to own RAG architectures, agentic workflows, evaluation pipelines, and inference optimization that serve real users at scale.
About The Role
The role focuses on building production LLM systems - not demos. Expect to own RAG architectures, agentic workflows, evaluation pipelines, and inference optimization that serve real users at scale.
The work sits at the intersection of applied research and software engineering: translating state-of-the-art techniques into reliable, observable, cost-efficient services alongside a team of ML engineers and platform engineers.
Key Responsibilities
- Design, build, and ship LLM-powered features - RAG pipelines, agents, structured extraction - using LangChain, LlamaIndex, or custom orchestration layers
- Integrate and optimize vector stores (Pinecone, Weaviate, pgvector) and embedding pipelines for retrieval quality and latency targets
- Build systematic evaluation harnesses: golden datasets, LLM-as-judge scoring, regression suites, and guardrail testing before every release
- Fine-tune and adapt open-weight models (Llama, Mistral, Qwen) using LoRA/QLoRA and distillation on domain-specific data
- Optimize inference for cost and latency - quantization, batching, caching strategies, and vLLM or TensorRT-LLM deployment
- Instrument LLM applications with tracing, observability, and monitoring (LangSmith, OpenTelemetry, custom dashboards) to catch drift and degradation
- Collaborate with product and design teams to translate ambiguous requirements into scoped, testable AI capabilities
What We Are Looking For
- 3–6 years of software engineering experience, with at least 1–2 years building and shipping LLM/GenAI systems in production
- Strong Python engineering skills; comfort with async code, API design, and type-checked codebases
- Hands-on experience with RAG in production: chunking strategies, hybrid retrieval, reranking, and hallucination mitigation
- Practical knowledge of transformer architectures, prompt engineering limits, and when to fine-tune vs. prompt vs. retrieve
- Experience deploying models on cloud infrastructure (AWS, GCP, or Azure) with Docker/Kubernetes
- BS/MS in Computer Science, a related technical field, or equivalent practical experience
- Bonus: experience with agent frameworks (LangGraph, CrewAI), GPU profiling, multimodal models, or contributions to open-source LLM tooling