Senior Generative AI Engineer — RAG & AI Agents (Hybrid SF)

Cogniify

San Francisco (CA)

Hybrid

USD 150,000 - 170,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Very generous parental leave
Annual bonus program
ESPP
Team building experiences

Job summary

Cogniify in the San Francisco Bay Area is hiring a hands-on engineer to design, build, and deploy LLM-powered systems for production-grade AI applications. You will work across the full lifecycle from model choice to deployment and monitoring in a hybrid work setting.

The role emphasizes building RAG pipelines, intelligent agent workflows, and robust integrations with internal tools and APIs, with a focus on accuracy, latency, cost, and security while collaborating with product and engineering

Qualifications

  • Bachelor's or Master's degree in CS, Engineering, or a related discipline, or equivalent practical software-development experience.
  • 6-9 years of professional software-development experience, including strong hands-on experience with Python.
  • Experience developing backend services and integrating REST APIs.
  • Hands-on experience building LLM or Generative AI applications.
  • Practical experience implementing RAG using embeddings, semantic search, and vector databases.
  • Experience with LangChain, LangGraph, LlamaIndex, or a comparable LLM application framework.
  • Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini, or Azure OpenAI.
  • Working knowledge of vector databases including Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, or pgvector.
  • Experience with Docker and deployment on at least one of AWS, Azure, or GCP.
  • Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation.
  • Strong understanding of software engineering practices including Git, testing, debugging, and clean code.
  • Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.

Responsibilities

  • Design and develop scalable LLM-powered applications using Python.
  • Build RAG pipelines including document processing, embeddings, vector databases, semantic search, and reranking.
  • Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration, and human approval steps.
  • Integrate LLMs with internal systems, external APIs, databases, and enterprise applications.
  • Evaluate and select foundation models based on accuracy, latency, cost, security, and business needs.
  • Improve prompt quality, retrieval accuracy, response time, and token usage.
  • Implement safety guardrails, output validation, access controls, and fallback mechanisms.
  • Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability.
  • Containerize applications with Docker and deploy on AWS, Azure, or GCP.
  • Apply monitoring and LLMOps practices for model performance, cost, latency, errors, and production usage.
  • Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions.
  • Document technical architecture, design decisions, APIs, and operational processes.

Education

Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline

Job description

Cogniify in the San Francisco Bay Area is hiring a hands-on engineer to design, build, and deploy LLM-powered systems for production-grade AI applications. You will work across the full lifecycle from model choice to deployment and monitoring in a hybrid work setting.

The role emphasizes building RAG pipelines, intelligent agent workflows, and robust integrations with internal tools and APIs, with a focus on accuracy, latency, cost, and security while collaborating with product and engineering

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GenAI Engineer — Production-Grade LLM & RAG
Senior GenAI Engineer — Production-Grade LLM & RAG

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior GenAI Architect — Production LLM & RAG Expert
Senior GenAI Architect — Production LLM & RAG Expert

TECHNEPTUNE CONSULTING INC • United States

Hybrid
USD 250,000 - 300,000
Senior Agentic AI Engineer — Hybrid (Azure LLMs & RAG)
Senior Agentic AI Engineer — Hybrid (Azure LLMs & RAG)

Creative Solutions Services, LLC • Houston (TX)

Hybrid
USD 170,062,000 - 220,722,000
Medical, Vision, and Dental Insurance
401k Retirement Fund
GenAI Engineer - RAG & Autonomous Agent Lead
GenAI Engineer - RAG & Autonomous Agent Lead

MACHINE LEARNING TECHNOLOGIES LLC • Engineer Springs (CA)

On-site
USD 96,000 - 207,000
Senior AI Platform Engineer (Hybrid)
Senior AI Platform Engineer (Hybrid)

People In AI • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior Generative AI Engineer
Senior Generative AI Engineer

Cogniify • San Francisco (CA)

Hybrid
USD 150,000 - 170,000
Unlimited PTO
Very generous parental leave
Annual bonus program
+2
Senior Generative AI Engineer — LLMs, RAG & AI Systems
Senior Generative AI Engineer — LLMs, RAG & AI Systems

Synechron • Charlotte (NC)

On-site
USD 105,000 - 108,000
Competitive compensation
Work abroad opportunities
Paid annual leave
+7
AI Agent Systems Architect — Hybrid (SF)
AI Agent Systems Architect — Hybrid (SF)

Goliath-Partners • San Francisco (CA)

Hybrid
USD 220,000 - 275,000
Ownership potential
Hybrid work model
Career growth in AI
Lead AI Analytics Scientist – RAG & LLMs (Hybrid)
Lead AI Analytics Scientist – RAG & LLMs (Hybrid)

Partner Company • United States

Hybrid
USD 136,000 - 192,000
Hybrid work environment
Senior AI Platform Engineer — LLM & RAG Systems
Senior AI Platform Engineer — LLM & RAG Systems

SupportFinity™ • San Francisco (CA)

Hybrid
USD 160,000 - 235,000
Health benefits including medical, dental, and vision
401(k) retirement plan
Learning and development budget
+1