This hybrid role in Dallas, TX focuses on designing, building, and deploying autonomous and semi-autonomous AI agents on AWS. You will work with goal-driven, tool-using, multi-step systems that combine AWS AI/ML services with LLMs, while integrating enterprise systems and emphasizing production-grade safety and reliability.
Responsibilities
- Design and develop agentic AI systems using Amazon Bedrock and other foundation models, including Claude and Titan.
- Build autonomous and semi-autonomous AI agents that perform multi-step reasoning, planning, tool usage, action execution, and human-in-the-loop collaboration.
- Integrate with Amazon SageMaker for custom model training, fine-tuning, evaluation, and experimentation.
- Implement Retrieval Augmented Generation (RAG) solutions using Amazon OpenSearch, Amazon Aurora, DynamoDB, and vector databases such as FAISS or Pinecone.
- Optimize inference cost, latency, scalability, and reliability across AI workloads.
- Implement guardrails, validation layers, and human-in-the-loop controls to support safe, reliable, and predictable AI behavior.
- Address hallucination mitigation, prompt injection risks, bias concerns, and model misuse scenarios.
- Ensure compliance with enterprise AI governance, security, and regulatory standards.
- Implement comprehensive logging, monitoring, observability, and audit trails for AI systems.
- Support production deployments, incident resolution, and root cause analysis for AI services.
- Continuously improve agent performance, reliability, robustness, and usability through iteration, monitoring, and feedback.
- Collaborate with architecture, DevOps, data engineering, security, and business teams to deliver end-to-end AI solutions.
Requirements
- Bachelor’s degree in computer science, Engineering, Technology, or a related field, with 8–12 years of overall IT experience and at least 3+ years of hands‑on experience building agentic AI solutions using AWS cloud‑native services.
- Strong expertise in core AWS services: AWS Lambda, Step Functions, EventBridge, S3, DynamoDB or Aurora, OpenSearch, Amazon Bedrock, and Amazon SageMaker.
- Advanced proficiency in Python for building scalable AI systems and automation workflows.
- Proven experience leading the technical design and implementation of complex AI agent architectures and components.
- Hands‑on experience designing RAG solutions and applying model fine‑tuning techniques.
- Demonstrated ability to manage performance, scalability, reliability, and cost optimization of AI workloads in production environments.
- Proven experience conducting code reviews, leading knowledge‑sharing sessions, and making critical decisions on technologies, architectures, and frameworks.
- Healthcare domain experience is desirable, along with knowledge of AI safety, governance, compliance, and responsible AI practices.
- AWS certifications such as Solutions Architect or Machine Learning Specialty are preferred; experience with LangChain and LangGraph is a plus.
- Excellent communication skills with the ability to collaborate effectively with technical and non-technical stakeholders.
Required Skills & Technologies
- AWS, Amazon Bedrock, Amazon SageMaker, Amazon OpenSearch, Amazon Aurora, DynamoDB
- AWS Lambda, Step Functions, EventBridge, S3
- Python
- Claude, Titan
- FAISS, Pinecone
- LangChain, LangGraph
Experience
- Minimum experience: 3 years
- Education: Bachelor’s degree
Location