The Senior AI Platform & Solutions Engineer designs, builds, implements, and maintains the enterprise AI platform, including cloud infrastructure, AI services, data pipelines, integrations, and Agentic AI solutions.
Reporting to the Director of AI, this senior hands‑on role combines AWS cloud and AI platform engineering with Generative AI development to securely develop, deploy, integrate, monitor, and scale enterprise AI solutions.
This role confirms and implements AWS-based AI infrastructure, reusable AI services and APIs, RAG capabilities, data pipelines, webhooks, integrations, and AI agents. The engineer also develops and optimizes AI solutions using LLMs, prompt and context engineering, tool calling, agent orchestration, multimodal AI, and computer vision.
The Senior AI Platform & Solutions Engineer partners with AI leadership, software development, IT infrastructure, cybersecurity, enterprise systems, and business stakeholders to deliver scalable, secure, maintainable, production‑ready AI capabilities.
AI Platform Architecture & Engineering
- Design, build, implement, and maintain the enterprise AI platform and services
- Develop scalable architecture for AI applications, agents, models, knowledge bases, and use cases
- Build reusable AI services, APIs, components, and integration patterns
- Establish standards and reference architectures for enterprise AI development
- Design AI solutions for scalability, reliability, maintainability, security, and cost efficiency
- Evaluate emerging AI technologies and recommend adoption where appropriate
- Design and implement AWS infrastructure for enterprise AI applications and services
- Develop production-ready architectures using AWS services such as Bedrock, Lambda, API Gateway, S3, ECS/EKS, IAM, and CloudWatch
- Configure and maintain model endpoints, infrastructure, networking, storage, and compute resources
- Implement Infrastructure as Code and repeatable provisioning
- Build containerized workloads with Docker and appropriate orchestration tools
- Implement monitoring, logging, alerting, scaling, backup, and disaster recovery
- Optimize AI infrastructure performance, availability, latency, and cloud costs
Agentic AI & AI Application Development
- Build production AI agents and Agentic AI workflows
- Develop single- and multi-agent architectures for business needs
- Implement reasoning, planning, tool/function calling, memory, and orchestration
- Enable agents to securely interact with systems, APIs, databases, and applications
- Build human-in-the-loop workflows and controls for autonomous processes
- Develop copilots, assistants, intelligent search, document intelligence, and automation solutions
- Integrate agents into enterprise applications and workflows
- Develop enterprise applications using LLMs and Generative AI
- Optimize prompt and context engineering strategies
- Implement model selection and routing for performance, security, latency, and cost
- Develop structured output, tool-use, and reasoning workflows
- Evaluate fine-tuning and model adaptation when appropriate
- Improve model reliability, quality, accuracy, and consistency
Retrieval-Augmented Generation (RAG) & Knowledge Systems
- Design and implement enterprise RAG architectures
- Build ingestion pipelines for structured and unstructured content
- Develop document processing, chunking, embedding, indexing, and retrieval strategies
- Implement semantic, vector, keyword, and hybrid search
- Maintain vector databases and AI knowledge repositories
- Optimize retrieval quality, context relevance, and response accuracy
- Support access controls and data-level permissions
- Develop safeguards to reduce hallucinations and unreliable responses
- Design scalable data pipelines for AI applications and services
- Build ETL/ELT and ingestion workflows for structured, semi-structured, and unstructured data
- Connect databases, documents, APIs, repositories, and cloud storage to AI services
- Develop transformation, validation, metadata extraction, and synchronization processes
- Implement event-driven and near-real-time processing as needed
- Ensure pipelines are reliable, observable, maintainable, and secure
APIs, Webhooks & Enterprise Integrations
- Design REST APIs and webhook integrations
- Build event-driven architectures linking AI with enterprise applications
- Integrate AI solutions with internal and third-party systems
- Implement secure API authentication and authorization
- Support OAuth, SAML, Microsoft Entra ID, service accounts, and identity standards
- Create reusable integration patterns for enterprise AI consumption
AI Security, Governance & Responsible AI
- Implement security controls across AI platforms, applications, agents, and pipelines
- Design authentication, authorization, RBAC, secrets management, and encryption
- Implement AI guardrails and content safety controls
- Support AI governance and Responsible AI standards
- Enable audit logging and traceability for AI and agent activity
- Partner with cybersecurity and compliance teams to address AI risks
- Protect enterprise data and sensitive information across AI workflows
AI Evaluation, Testing & Observability
- Develop automated evaluation and testing for AI applications and agents
- Define metrics for accuracy, relevance, hallucination, latency, reliability, and cost
- Implement regression testing for prompts, RAG systems, agents, and models
- Build monitoring and observability for AI applications and workflows
- Implement tracing for model calls, tool use, actions, failures, and performance
- Troubleshoot production AI issues and perform root cause analysis
DevOps & Production Operations
- Develop CI/CD pipelines for AI applications and infrastructure
- Implement source control, testing, deployment, and release practices
- Support releases, troubleshooting, incident response, and maintenance
- Maintain architecture, deployment, API, and operations documentation
- Improve platform reliability, developer experience, and deployment efficiency
- Support AI solutions using images, documents, video, and other multimodal data
- Develop or integrate computer vision capabilities such as classification, object detection, OCR, and visual analysis
- Evaluate vision models against business requirements
- Integrate computer vision into AI applications, agents, and workflows
Qualifications
Required:
- Bachelor’s degree in Computer Science, AI, ML, Software Engineering, IT, Data Engineering, or related field
- 8+ years of software, cloud, data, ML engineering, or related experience
- 3+ years building production AI, ML, or Generative AI solutions
- Strong hands‑on AWS solution design and implementation experience
- Experience with AWS services such as Bedrock, Lambda, API Gateway, S3, IAM, CloudWatch, ECS/EKS, or similar cloud technologies
- Strong Python and software engineering skills
- Experience with REST APIs, webhooks, and enterprise integrations
- Experience building data ingestion and processing pipelines
- Experience developing LLM-based applications
- Experience with RAG solutions
- Experience building AI agents, tool‑calling workflows, or Agentic AI applications
- Experience with containers, CI/CD, and cloud‑native development
- Understanding of enterprise security, identity, authentication, authorization, and data protection
- Ability to architect independently while remaining hands‑on in implementation and support
Preferred:
- Experience with Bedrock, OpenAI, Claude, Gemini, Llama, or other foundation model platforms
- Experience with AWS AI/ML services and enterprise AWS architectures
- Experience with LangGraph, LangChain, Semantic Kernel, CrewAI, AutoGen, OpenAI Agents SDK, or similar frameworks
- Experience with Model Context Protocol (MCP)
- Experience with vector databases and search tools such as OpenSearch, pgvector, Pinecone, Weaviate, or Milvus
- Experience with fine‑tuning, model adaptation, embeddings, or Hugging Face
- Experience with computer vision, OCR, object detection, or multimodal AI
- Experience with PyTorch, TensorFlow, OpenCV, YOLO, or related ML tools
- Experience with AI evaluation tools such as Promptfoo, DeepEval, RAGAS, or LangSmith
- Experience with Terraform or other Infrastructure as Code tools
- Experience with Docker and Kubernetes
- AWS Solutions Architect, AWS Machine Learning, or similar cloud certification
- Experience building AI solutions in engineering, construction, field services, utilities, industrial, or project-based services environments