Job Title: Technology Architect | Cloud Platform | Google Cloud - Architecture – Gen AI Engineer
Work Location & Reporting Address: Charlotte, NC 28202 (Onsite-Hybrid. LOCAL CANDIDATES ONLY!!!)
Contract Duration: 12 months
Maximum Vendor Rate: * per hour max
Target Start Date: 01 Jul 2026
Visa Independence: Yes
Must Have Skills
- GEN AI
- Agentic AI
- VLLM
- Fast API
- REST API
- MCD
- Lang Graph
- Lang Chain
- Graph RAG
- ML Ops
- Python
- ML
- Data Science
- RAG
- LLM
Nice to Have Skills
Detailed Job Description
We are seeking a highly skilled Generative AI Engineer with a strong Python background to design, develop, and deploy cutting‑edge AI solutions. The ideal candidate will have hands‑on experience with Large Language Models (LLMs), prompt engineering, and Gen AI frameworks, alongside expertise in building scalable AI applications and developing Agentic AI solutions.
Key Responsibilities
- Design and implement Generative AI models for text, image, or multimodal applications.
- Develop prompt engineering strategies and embedding‑based retrieval systems.
- Integrate Gen AI capabilities into web applications and enterprise workflows.
- Build agentic AI applications with context engineering and MCP tools.
Required Skills & Qualifications
- 7+ years of hands‑on experience in AI, Data Science, ML, and Gen AI.
- 2+ years of strong hands‑on experience in Agentic AI, VLLM’s, Gen AI, Lang Chain, Lang Graph, RAG, LLM Oops, and AI Services in GCP and Azure.
- Strong experience designing and deploying Retrieval‑Augmented Generation (RAG) pipelines.
- Strong MLOps/LLMOps experience with CI/CD automation.
- Extensive experience with LangChain, LangGraph, and agentic AI patterns including routing, memory, multi‑agent orchestration, guardrails, and failure recovery.
- Experience in cloud‑native engineering across AWS (SageMaker, Lambda, ECS/Fargate, S3, API Gateway, Step Functions) and GCP (Vertex AI) for scalable AI delivery.
- Experience developing microservices and API development using FastAPI, REST APIs, Pydantic/JSON schemas, Docker, and Kubernetes for low‑latency serving.
- Strong experience with vector databases and semantic search technologies including Pinecone, FAISS, ChromaDB, and embedding lifecycle management.
- Strong proficiency in Python and AI/ML frameworks (PyTorch, TensorFlow).
- Hands‑on experience using session and memory for building multi‑agent systems along with using MCP tools.
- Hands‑on experience with LLMs, transformers, and Hugging Face ecosystem.
- Knowledge and experience with vector databases and RAG technique for semantic search.
- Familiarity with cloud AI services (AWS SageMaker, Azure OpenAI, GCP Vertex AI).
- Understanding of MLOps practices for scalable AI deployment.
- Strong experience in working with LLM fine‑tuning with LoRA, QLoRA, PEFT.
- Strong experience designing advanced RAG systems using Pinecone, FAISS, Weaviate, Chroma, hybrid retrieval, and custom embeddings.
- Strong experience in designing end‑to‑end LLMOps/MLOps pipelines using MLflow, DVC, SageMaker Pipelines, Vertex AI Pipelines, and GitHub Actions.
- Experience using cloud‑native AI systems on AWS (SageMaker, Lambda, EKS, EC2, Step Functions, S3, Glue) and GCP Vertex AI, supporting high‑volume inference and secure enterprise operations.
- Experience developing multi‑agent orchestration workflows using LangGraph and CrewAI for tool‑calling, validation agents, automated reasoning, and workflow supervision.
Minimum Years of Experience
10+ years
Interview Process
Face to face interview