Senior Generative AI Engineer

Cogniify

San Francisco (CA)

Hybrid

USD 150,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Very generous parental leave
Annual bonus program
ESPP
Team building experiences

Job summary

Cogniify in the San Francisco Bay Area is hiring a hands-on engineer to design, build, and deploy LLM-powered systems for production-grade AI applications. You will work across the full lifecycle from model choice to deployment and monitoring in a hybrid work setting.

The role emphasizes building RAG pipelines, intelligent agent workflows, and robust integrations with internal tools and APIs, with a focus on accuracy, latency, cost, and security while collaborating with product and engineering

Qualifications

  • Bachelor's or Master's degree in CS, Engineering, or a related discipline, or equivalent practical software-development experience.
  • 6-9 years of professional software-development experience, including strong hands-on experience with Python.
  • Experience developing backend services and integrating REST APIs.
  • Hands-on experience building LLM or Generative AI applications.
  • Practical experience implementing RAG using embeddings, semantic search, and vector databases.
  • Experience with LangChain, LangGraph, LlamaIndex, or a comparable LLM application framework.
  • Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini, or Azure OpenAI.
  • Working knowledge of vector databases including Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, or pgvector.
  • Experience with Docker and deployment on at least one of AWS, Azure, or GCP.
  • Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation.
  • Strong understanding of software engineering practices including Git, testing, debugging, and clean code.
  • Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.

Responsibilities

  • Design and develop scalable LLM-powered applications using Python.
  • Build RAG pipelines including document processing, embeddings, vector databases, semantic search, and reranking.
  • Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration, and human approval steps.
  • Integrate LLMs with internal systems, external APIs, databases, and enterprise applications.
  • Evaluate and select foundation models based on accuracy, latency, cost, security, and business needs.
  • Improve prompt quality, retrieval accuracy, response time, and token usage.
  • Implement safety guardrails, output validation, access controls, and fallback mechanisms.
  • Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability.
  • Containerize applications with Docker and deploy on AWS, Azure, or GCP.
  • Apply monitoring and LLMOps practices for model performance, cost, latency, errors, and production usage.
  • Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions.
  • Document technical architecture, design decisions, APIs, and operational processes.

Education

Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline

Job description

Cogniify is building production-grade Generative AI applications and is hiring a hands-on engineer to design, build, and deploy systems that power LLM-powered products. This role in the San Francisco Bay Area (hybrid) focuses on core capabilities like RAG pipelines, intelligent agent workflows, and reliable integration with internal tools, external APIs, and enterprise applications.

You will work across the full lifecycle, from model and framework selection to evaluation, safety guardrails, deployment, and ongoing LLMOps monitoring.

Key Responsibilities
  • Design and develop scalable LLM-powered applications using Python.
  • Build RAG pipelines including document processing, embeddings, vector databases, semantic search, and reranking.
  • Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration, and human approval steps.
  • Integrate LLMs with internal systems, external APIs, databases, and enterprise applications.
  • Evaluate and select foundation models based on accuracy, latency, cost, security, and business needs.
  • Improve prompt quality, retrieval accuracy, response time, and token usage.
  • Implement safety guardrails, output validation, access controls, and fallback mechanisms.
  • Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability.
  • Containerize applications with Docker and deploy on AWS, Azure, or GCP.
  • Apply monitoring and LLMOps practices for model performance, cost, latency, errors, and production usage.
  • Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions.
  • Document technical architecture, design decisions, APIs, and operational processes.
Requirements
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience.
  • 6-9 years of professional software-development experience, including strong hands-on experience with Python.
  • Experience developing backend services and integrating REST APIs.
  • Hands-on experience building LLM or Generative AI applications.
  • Practical experience implementing RAG using embeddings, semantic search, and vector databases.
  • Experience with LangChain, LangGraph, LlamaIndex, or a comparable LLM application framework.
  • Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini, or Azure OpenAI.
  • Working knowledge of vector databases including Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, or pgvector.
  • Experience with Docker and deployment on at least one of: AWS, Azure, or GCP.
  • Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation.
  • Strong understanding of software engineering practices including Git, testing, debugging, and clean code.
  • Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.
Technologies and Frameworks
  • Programming and AI: Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking, Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama
  • Frameworks: LangChain, LangGraph, LlamaIndex, LangSmith, Langfuse, Arize Phoenix, AutoGen, CrewAI, Semantic Kernel
  • Observability/ML tools: MLflow, Weights & Biases
  • Model access: OpenAI, Anthropic Claude, Gemini, Azure OpenAI
  • Vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector
  • Cloud and delivery: Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines, REST APIs
Benefits
  • Unlimited PTO.
  • Very generous parental leave, much above industry standards.
  • Entrepreneurial culture where pushing limits and taking risks is everyday business.
  • Open communication with management and company leadership.
  • Small, dynamic teams = massive impact.
  • Medical, Dental and Vision coverage for employees.
  • Access to Disability & Life insurance.
  • Mental health and wellbeing support.
  • Annual bonus program.
  • Employer Stock Purchase Program (ESPP).
  • Yearly team building experiences.
  • Mentorship and sponsorship opportunities.
  • Manager resources and support.
Success in the Role
  • Production-ready AI applications that are accurate, secure, and maintainable.
  • RAG systems that retrieve relevant information and reduce hallucinations.
  • AI-agent workflows that reliably complete business tasks and integrate with existing systems.
  • Measurable improvements in response quality, latency, and inference cost.
  • Clear monitoring of application performance, usage, errors, and model behavior.

Salary range: USD 150,000 - 170,000 per year (US East/West Coast).

Work location: Hybrid remote, San Francisco Bay Area, CA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Generative AI Engineer - USA
Senior Generative AI Engineer - USA

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Medical, Dental and Vision coverage
+1
AI Engineer – Senior Software Engineer
AI Engineer – Senior Software Engineer

DataJobs • Randolph Township (NJ)

On-site
USD 50,000 - 600,000
AI Engineer (III)
AI Engineer (III)

BravoTECH • Richardson (TX)

Hybrid
USD 140,000 - 210,000
Health, dental, and vision insurance
Retirement savings plan with company 1
Paid time off and holidays
+1
Sr AI/ML Engineer
Sr AI/ML Engineer

Vizient • Irving (TX)

On-site
USD 102,400 - 179,000
Senior AI Engineer (Generative AI & Agentic Systems)
Senior AI Engineer (Generative AI & Agentic Systems)

WonderBotz • United States

Remote
USD 150,000 - 210,000
Lead Data Scientist Applied AI - USA Onsite (Santa Clara, CA)
Lead Data Scientist Applied AI - USA Onsite (Santa Clara, CA)

Dover • Santa Clara (CA), Northern (KY)

Hybrid
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Employee stock purchase program
+2
Gen. AI Engineer
Gen. AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000
GenAI Engineer (US)
GenAI Engineer (US)

Latitude • Seattle (WA)

Hybrid
USD 130,000 - 210,000
Medical, dental, vision insurance
401(k) retirement benefits
Paid time off and holidays
+3
Gen. AI Engineer
Gen. AI Engineer

Kaleidoscope Innovation, Inc. • Fort Worth (TX)

On-site
USD 100,000 - 160,000
AI Engineer - Sr Lead Software Engineer
AI Engineer - Sr Lead Software Engineer

DataJobs • Jersey City (NJ)

On-site
USD 176,000 - 260,000