Function
Cloud & Data Engineering
Our Company
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.
Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.
Job description
Meet Our Team
Join a forward-thinking Engineering and AI Platform team focused on building the next generation of enterprise AI solutions. Our team is pioneering agentic AI ecosystems powered by AWS Bedrock, AgentCore, MCP servers, and modern cloud-native technologies.
As a Senior AWS AgentCore Platform Engineer, you'll work alongside Cloud Architects, AI Engineers, Platform Engineers, and Security specialists to establish scalable, secure, and observable AI platforms. You'll play a critical role in defining the operational foundation that enables enterprise teams to deploy AI agents with confidence, governance, and efficiency.
This is an exciting opportunity to shape enterprise AI infrastructure, drive innovation in LLMOps, and influence platform standards across multiple business units.
What You'll Be Doing
AI Platform Observability & Reliability
- Design and implement enterprise-grade observability solutions for AI agent ecosystems built on AWS Bedrock, AgentCore, and MCP servers.
- Assess and optimize CloudWatch, X-Ray, Bedrock logging, and AgentCore tracing capabilities against agentic workflow requirements.
- Conduct gap analyses and implement observability solutions using Dynatrace and other monitoring platforms.
- Develop distributed tracing frameworks for AI workloads, including: LLM decision paths
- Tool invocations
- Sub-agent interactions
- MCP server communications
- Build structured logging frameworks to support troubleshooting, governance, and performance optimization.
- Design post-deployment validation pipelines for AI agents and MCP servers, including deployment health monitoring and registration verification.
Cost Governance & Optimization
- Architect cost visibility and governance frameworks across AI workloads.
- Extend cloud tagging strategies to include agent runtimes, vector databases, MCP services, and Bedrock token consumption.
- Develop cost allocation models to provide spending transparency by team, department, and application.
- Build dashboards and reporting solutions for AI platform cost tracking and forecasting.
- Configure AWS Budgets, automated alerts, anomaly detection, and optimization recommendations.
- Deliver automated cost reporting through Microsoft Teams and email channels.
Monitoring & Incident Management
- Define enterprise monitoring standards and alerting frameworks across AI platform services.
- Create and manage P1-P4 alerting strategies covering: Deployment failures
- Runtime exceptions
- Tool invocation errors
- MCP connectivity issues
Security & Platform Governance
- Assess IAM architectures and multi-team access models for enterprise-scale AI environments.
- Design Attribute-Based Access Control (ABAC) frameworks to support secure multi-team isolation.
- Evaluate Cedar policy engine capabilities within AgentCore for fine-grained authorization models.
- Develop reusable Terraform modules to enforce governance, security, and compliance standards.
- Identify scalability risks and implement secure platform design patterns for enterprise AI adoption.
Platform Engineering & Automation
- Build and maintain Infrastructure-as-Code solutions using Terraform.
- Design and enhance CI/CD pipelines supporting AI platform deployments.
- Collaborate with engineering, security, architecture, and business stakeholders in Agile environments.
- Drive platform standardization, automation, and operational excellence initiatives.
What You'll Bring to the Team
Required Qualifications
- 8+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud services including: IAM CloudWatch AWS Lambda AWS Bedrock Cloud-native monitoring and governance services
- Hands-on experience implementing observability and distributed tracing solutions using tools such as: Dynatrace Jaeger Honeycomb OpenTelemetry
- Experience designing and managing Infrastructure-as-Code using Terraform.
- Strong background buil