Turn this role into an interview — a resume and cover letter built around what this employer wants.
Teradata is seeking a Senior Principal Architect of Cloud Platform Engineering to lead the evolution of its cloud platform into an agentic era. You will define long‑term architecture across AWS, Azure, and GCP, with multi‑tenant models and agent‑driven execution frameworks, enabling autonomous AI systems to run as first‑class compute consumers.
You will guide provisioning, memory management, security, and observability while driving measurable improvements in latency, cost, and reliability
At Teradata, we believe that people thrive when empowered with better information. Teradata Autonomous Knowledge Platform activates enterprise intelligence by unifying data, knowledge and business context to achieve tangible outcomes. With Teradata, organizations can provide agents with full context for impact when it matters. Our solution lets businesses connect and scale on premises, in the cloud, or through a hybrid approach. Teradata delivers real business value with AI.
As Sr. Principal Architect of Cloud Platform Engineering, you will lead the evolution of Teradata's cloud deployment and platform technologies into the agentic era—designing and executing the technical strategy for a cloud platform engineered from the ground up to serve autonomous AI systems as first‑class compute consumers. You will define long‑term cloud architecture across AWS, Azure, and GCP with multi‑tenant and single‑tenant models, agent‑driven execution frameworks, and infrastructure primitives that agentic systems require: sub‑second APIs, persistent agent memory, agent‑scoped identity, fine‑grained cost controls, and deterministic audit trails. Your platform decisions will directly enable Teradata's engineering teams to adopt AI‑native development practices where autonomous agents become core participants in software delivery, while your leadership directly impacts the ability to scale globally, onboard agentic applications to production reliably, and operate infrastructure serving both human users and agents with equal rigor.
Success means delivering a highly automated, secure, and resilient cloud platform architected natively for autonomous AI systems
and the engineering teams that deploy them:
Sub‑second agent API latencies; dynamic scaling for agent micro‑queries; agent‑driven provisioning and self‑healing
infrastructure; reliable agentic workload execution with <5% cost variance from workload unpredictability
Identity and cost control systems that govern both human and agent actors; observable, debuggable agent execution; per
execution cost tracking and circuit breaking; deterministic audit trails for agentic decision reasoning
Engineering teams shipping features authored, tested, and reviewed by agentic systems; measurable velocity and quality
improvements from autonomous coding agents; reduced manual toil in infrastructure management and code verification
Reduced provisioning times; strong SLA adherence for both human and agent workloads; high platform availability (>99. 95%);
optimized cloud spend with per‑execution controls; incident response accelerated by agent‑driven diagnostics and remediation
Proven experience leading cloud platform, infrastructure, or SaaS engineering teams at scale, with deep expertise in cloud
native architectures (serverless, event-driven, agent‑driven execution models)
Hands‑on experience building and productionizing agentic applications at scale in production environments (agents writing
code, managing infrastructure, making autonomous decisions), with deep familiarity in agentic frameworks (LangGraph, Claude
API with tool_use, MCP servers, agent orchestration, memory management) and understanding how to architect infrastructure
for their unique demands (sub‑second latency, persistent memory, execution cost controls)
Proven ability to architect for agentic workload patterns: understanding agent failure modes (hallucinations, cost drift, goal
misalignment) and designing guardrails, observability, cost controls, and deterministic audit trails into platform primitives;
balancing reliability, performance, security, and cost across diverse workloads from bulk analytical scans to high‑frequencyagent micro‑queries
Direct, hands‑on experience building systems where autonomous agents operate as core decision‑makers and infrastructure
consumers (not just code‑assist tools); deep understanding of LLM capabilities and limitations in production—token budgets,
latency requirements, hallucination handling, determinism, reproducibility, cost predictability
Ability to architect for agent auditability and debugging: designing systems that produce interpretable agent reasoning trails,
enable reasoning replay, and establish accountability for agentic decisions; financial acumen for agent compute optimization—
understanding per‑execution spend tracking, cost circuit breaking, and the unique cost profiles of agentic vs. traditional
workloads
Hands‑on leadership experience with AWS, Azure, and/or GCP at scale; expertise in IaC, CI/CD, and platform automation
(Terraform, Jenkins, GitHub Actions, deployment orchestration); strong understanding of observability, incident management,
DR, and SLA‑driven operations extended to non‑deterministic agent workloads
A security‑first mindset embedding IAM, agent identity, delegated authority, and intent‑scoped governance into platform design;
ability to define KPIs across deployment frequency, provisioning time, latency, error rates, infrastructure efficiency, agent
execution cost per task, and agentic system reliability
Demonstrated success building and scaling high‑performing teams through periods of significant technical and organizational
change; proven ability to guide engineering teams on agentic thinking—shifting from "AI as tool" to "AI as engineer"—with
corresponding trust, autonomy, and verification frameworks
Ability to articulate the transition to agentic infrastructure across technical and non‑technical audiences; comfort with pioneering
new patterns where agentic infrastructure is nascent and requires novel solutions for cost control, identity, latency, and
auditability
Agentic systems are moving from research labs into production. Cloud platforms not designed for agent workloads will struggle with
their unique demands: sub‑millisecond latency sensitivity, unpredictable concurrency, deterministic auditability, and per‑execution