Senior Generative AI Engineer - USA

Socket.dev

San Francisco (CA)

On-site

USD 150,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Generous parental leave
Medical, Dental and Vision coverage
ESPP

Job summary

Socket.dev is seeking a hands‑on Generative AI and LLM Engineer to design, build and deploy production‑grade AI applications. You will own solutions end‑to‑end, from requirements to monitoring, with a focus on LLMs, RAG pipelines and intelligent agent workflows.

The role requires strong Python backend development experience and practical knowledge of vector databases, API integrations, Docker and cloud platforms.

Qualifications

  • Bachelor's or Master's in Computer Science or related field.
  • 6–9 years of professional software development with Python.
  • Experience building backend services and REST APIs.
  • Hands‑on experience with LLMs and Generative AI.
  • Practical RAG using embeddings, vector databases and search.

Responsibilities

  • Design and develop scalable LLM‑powered applications using Python.
  • Build RAG pipelines with embeddings, vector databases and semantic search.
  • Develop AI agent workflows with tool calling and memory.
  • Integrate LLMs with internal systems and external APIs.
  • Evaluate foundation models for accuracy, latency and cost.
  • Improve prompt quality, retrieval accuracy and token efficiency.
  • Implement safety guardrails, access controls and fallback mechanisms.
  • Containerize apps with Docker and deploy on AWS, Azure or GCP.
  • Set up monitoring and LLMOps for performance and costs.
  • Collaborate with product and engineering to deliver reliable AI solutions.
  • Document architecture, APIs and operational processes.

Skills

Python
LLM/Generative AI
Backend development
API integrations
Docker
Cloud platforms
Vector databases
LangChain/LangGraph
Team communication

Education

Bachelor's or Master's in CS

Tools

Docker
Kubernetes
LangGraph
Hugging Face
PyTorch
vLLM/Ollama
MLflow

Job description

The Role

We are looking for a hands‑on Generative AI and LLM Engineer to design, build and deploy production‑grade AI applications. The role will focus on developing LLM‑powered products, Retrieval‑Augmented Generation (RAG) pipelines and intelligent agent workflows using Python, modern AI frameworks and cloud infrastructure.

You will own solutions from requirement understanding and architecture through development, deployment and production monitoring. The ideal candidate combines strong Python backend development experience with practical knowledge of LLMs, vector databases, API integrations, Docker and at least one major cloud platform.

What You Will Do
  • Design and develop scalable LLM‑powered applications using Python.

  • Build RAG pipelines using document processing, embeddings, vector databases, semantic search and reranking.

  • Develop AI‑agent and multi‑agent workflows with tool calling, memory, orchestration and human approval steps.

  • Integrate LLMs with internal systems, external APIs, databases and enterprise applications.

  • Evaluate and select suitable foundation models based on accuracy, latency, cost, security and business requirements.

  • Improve prompt quality, retrieval accuracy, response time and token usage.

  • Implement safety guardrails, output validation, access controls and fallback mechanisms.

  • Build automated evaluation frameworks to measure response quality, hallucination, relevance and reliability.

  • Containerize applications using Docker and deploy them on AWS, Azure or GCP.

  • Implement monitoring and LLMOps practices for model performance, cost, latency, errors and production usage.

  • Collaborate with product, engineering and business teams to convert requirements into reliable AI solutions.

  • Document technical architecture, design decisions, APIs and operational processes.

What Success Looks Like
  • Production‑ready AI applications that are accurate, secure and maintainable.

  • RAG systems that retrieve relevant information and reduce hallucinations.

  • AI‑agent workflows that reliably complete business tasks and integrate with existing systems.

  • Measurable improvements in response quality, latency and inference cost.

  • Clear monitoring of application performance, usage, errors and model behaviour.

What We're Looking For
  • Bachelor's or Master's degree in Computer Science, Engineering or a related discipline, or equivalent practical software‑development experience.

  • 6-9 years of professional software‑development experience, including strong hands‑on experience with Python.

  • Experience developing backend services and integrating REST APIs.

  • Hands‑on experience building LLM or Generative AI applications.

  • Practical experience implementing RAG using embeddings, semantic search and vector databases.

  • Experience with LangChain, LangGraph, LlamaIndex or a comparable LLM application framework.

  • Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini or Azure OpenAI.

  • Working knowledge of vector databases such as Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS or pgvector.

  • Experience with Docker and deployment on at least one cloud platform: AWS, Azure or GCP.

  • Understanding of prompt engineering, hallucination reduction, output validation and LLM evaluation.

  • Strong understanding of software engineering practices, Git, testing, debugging and clean code.

  • Ability to communicate technical solutions clearly to both technical and non‑technical stakeholders.

Technical Frameworks and Toolkit
  • Experience building agentic or multi‑agent workflows using LangGraph, AutoGen, CrewAI, Semantic Kernel or similar frameworks.

  • Experience with Hugging Face Transformers, PyTorch or fine‑tuning techniques such as LoRA or QLoRA.

  • Knowledge of model serving and inference frameworks such as vLLM, TGI or Ollama.

  • Experience with Kubernetes, CI/CD pipelines and infrastructure automation.

  • Familiarity with observability or LLMOps tools such as LangSmith, Langfuse, Arize Phoenix, MLflow or Weights & Biases.

  • Knowledge of reranking, hybrid search, chunking strategies and retrieval evaluation.

  • Experience implementing AI guardrails, PII protection, prompt‑injection prevention and responsible AI practices.

Salary Range

US East/West Coast: $150000 - $170000

Disclaimer

The base salary range is a guideline and may vary based on factors such as candidate experience, specialized skills, and geographical location. Actual compensation may include additional benefits and bonuses.

Perks And Benefits of Working With Us
  • Unlimited PTO.

  • Please ask us about our very generous parental leave, much above industry standards!

  • Entrepreneurial culture where pushing limits and taking risks is everyday business.

  • Open communication with management and company leadership.

  • Small, dynamic teams = massive impact.

  • Medical, Dental and Vision coverage for employees.

  • Access to Disability & Life insurance.

  • Mental health and wellbeing support.

  • Annual bonus program.

  • Employer Stock Purchase Program (ESPP).

  • Yearly team building experiences.

  • Mentorship and sponsorship opportunities.

  • Manager resources and support.

Cogniify is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other protected characteristic

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Scientist Applied AI - USA
Lead Data Scientist Applied AI - USA

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Annual bonus program
+3
Principal AI Scientist
Principal AI Scientist

Cerence AI • United States

On-site
USD 123,000 - 198,000
Annual bonus opportunity
Flexible Time Off
Choice-based medical coverage
+2
Sr. AI Architect
Sr. AI Architect

Ampcus, Inc • Chantilly (VA)

On-site
USD 180,000 - 260,000
AI Engineer
AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 180,000 - 260,000
GenAI Engineer (US)
GenAI Engineer (US)

Latitude • Seattle (WA)

Hybrid
USD 130,000 - 210,000
Medical, dental, vision insurance
401(k) retirement benefits
Paid time off and holidays
+3
Software Engineer AI/ML Systems - USA
Software Engineer AI/ML Systems - USA

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Entrepreneurial culture
+10
LLM Engineer (Remote)
LLM Engineer (Remote)

Cognizant • Louisville (KY)

Remote
USD 81,000 - 142,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan and contributions
+3
AI/ML Engineer 2
AI/ML Engineer 2

Day & Zimmermann Company • Philadelphia

On-site
USD 101,000 - 166,000
Medical/Rx coverage
Dental and vision coverage
100% paid maternity leave
+2
Jr. AI Developer
Jr. AI Developer

Ultimate Staffing • Plano (TX)

On-site
USD 65,000 - 95,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Egen • United States

On-site
USD 120,000 - 160,000
Comprehensive health insurance
Paid vacation
401(k) employer match
+1