Work Hours: Must work Pacific Time (PST/PDT) business hours
Employment Type: Long-Term Contract
About the Role
We are seeking a hands-on Applied AI Engineer to design, build, evaluate, and productionize LLM-powered agents and AI workflows.
This is a production engineering role rather than a research-only or proof-of-concept position. You will turn business use cases into reliable AI systems using an AWS-native platform with Amazon Bedrock as the foundation-model layer.
Engineers will work within standardized platform templates, CI/CD pipelines, guardrails, and evaluation frameworks while building AI applications that are reliable, observable, measurable, and ready for production.
Key Responsibilities
- Design, build, and deploy LLM-powered agents, workflows, and applications using Amazon Bedrock.
- Work with Bedrock capabilities including Bedrock Agents, Knowledge Bases, Guardrails, and multiple foundation models.
- Build agentic workflows incorporating RAG, tool/function calling, APIs, and enterprise data sources.
- Develop using standardized AI platform templates, CI/CD pipelines, guardrails, and evaluation harnesses.
- Build and maintain LLM evaluation suites as an integral part of the engineering lifecycle.
- Work with golden datasets, regression testing, and LLM-as-a-judge evaluation approaches.
- Implement prompt and context engineering strategies for reliable production behavior.
- Develop structured outputs, retries, fallbacks, and graceful degradation mechanisms.
- Integrate AI agents with enterprise applications and data through APIs and AWS services such as Lambda, Step Functions, SQS/SNS, and API Gateway.
- Instrument AI applications for quality, latency, token consumption, cost, and operational telemetry.
- Use services such as CloudWatch and Bedrock invocation metrics to monitor production systems.
- Collaborate directly with business users to rapidly iterate, demonstrate solutions, gather feedback, and deliver production functionality.
Required Qualifications
- 8+ years of software engineering experience.
- At least 2.5+ years of experience building production LLM applications, including agents, RAG pipelines, and tool/function calling.
- Strong programming skills in Python and/or TypeScript.
- Strong API engineering experience, including API design, versioning, authentication, and error handling.
- Strong practical experience with prompt engineering and context engineering.
- Hands-on experience developing LLM evaluation approaches including golden datasets, LLM-as-a-judge, and regression suites.
- Hands-on AWS experience.
- Production experience with Amazon Bedrock strongly preferred.
- Candidates with strong production experience using OpenAI/Anthropic APIs combined with AWS may also be considered if they can ramp quickly on Bedrock.
- Experience building reliable, production-grade AI systems rather than only prototypes or research projects.
- Ability to work full Pacific Time (PST/PDT) business hours.
Preferred Qualifications
- Strong information-retrieval/RAG experience, including chunking strategies, embeddings, hybrid search, and re-ranking.
- Experience with OpenSearch, pgvector, or Amazon Bedrock Knowledge Bases.
- Healthcare data experience involving PHI, PII, or HIPAA-aware engineering.
- LLM cost optimization experience, including model routing, prompt caching, batch inference, and throughput optimization.
- Infrastructure-as-Code experience using Terraform or AWS CDK.