Cogniify is building production-grade Generative AI applications and is hiring a hands-on engineer to design, build, and deploy systems that power LLM-powered products. This role in the San Francisco Bay Area (hybrid) focuses on core capabilities like RAG pipelines, intelligent agent workflows, and reliable integration with internal tools, external APIs, and enterprise applications.
You will work across the full lifecycle, from model and framework selection to evaluation, safety guardrails, deployment, and ongoing LLMOps monitoring.
Key Responsibilities
- Design and develop scalable LLM-powered applications using Python.
- Build RAG pipelines including document processing, embeddings, vector databases, semantic search, and reranking.
- Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration, and human approval steps.
- Integrate LLMs with internal systems, external APIs, databases, and enterprise applications.
- Evaluate and select foundation models based on accuracy, latency, cost, security, and business needs.
- Improve prompt quality, retrieval accuracy, response time, and token usage.
- Implement safety guardrails, output validation, access controls, and fallback mechanisms.
- Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability.
- Containerize applications with Docker and deploy on AWS, Azure, or GCP.
- Apply monitoring and LLMOps practices for model performance, cost, latency, errors, and production usage.
- Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions.
- Document technical architecture, design decisions, APIs, and operational processes.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience.
- 6-9 years of professional software-development experience, including strong hands-on experience with Python.
- Experience developing backend services and integrating REST APIs.
- Hands-on experience building LLM or Generative AI applications.
- Practical experience implementing RAG using embeddings, semantic search, and vector databases.
- Experience with LangChain, LangGraph, LlamaIndex, or a comparable LLM application framework.
- Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini, or Azure OpenAI.
- Working knowledge of vector databases including Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, or pgvector.
- Experience with Docker and deployment on at least one of: AWS, Azure, or GCP.
- Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation.
- Strong understanding of software engineering practices including Git, testing, debugging, and clean code.
- Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.
Technologies and Frameworks
- Programming and AI: Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking, Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama
- Frameworks: LangChain, LangGraph, LlamaIndex, LangSmith, Langfuse, Arize Phoenix, AutoGen, CrewAI, Semantic Kernel
- Observability/ML tools: MLflow, Weights & Biases
- Model access: OpenAI, Anthropic Claude, Gemini, Azure OpenAI
- Vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector
- Cloud and delivery: Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines, REST APIs
Benefits
- Unlimited PTO.
- Very generous parental leave, much above industry standards.
- Entrepreneurial culture where pushing limits and taking risks is everyday business.
- Open communication with management and company leadership.
- Small, dynamic teams = massive impact.
- Medical, Dental and Vision coverage for employees.
- Access to Disability & Life insurance.
- Mental health and wellbeing support.
- Annual bonus program.
- Employer Stock Purchase Program (ESPP).
- Yearly team building experiences.
- Mentorship and sponsorship opportunities.
- Manager resources and support.
Success in the Role
- Production-ready AI applications that are accurate, secure, and maintainable.
- RAG systems that retrieve relevant information and reduce hallucinations.
- AI-agent workflows that reliably complete business tasks and integrate with existing systems.
- Measurable improvements in response quality, latency, and inference cost.
- Clear monitoring of application performance, usage, errors, and model behavior.
Salary range: USD 150,000 - 170,000 per year (US East/West Coast).
Work location: Hybrid remote, San Francisco Bay Area, CA.