Principal AI Architect/Engineer

PepsiCo

Purchase (NY)

On-site

USD 180,000 - 206,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus 15%
Paid time off
Comprehensive benefits

Job summary

PepsiCo is seeking an AI Observability Architect to design and deploy a production-grade observability platform for AI agents, LLM workflows, and physical AI systems. You will own end-to-end telemetry, tracing, and quality signals across heterogeneous agent frameworks and platforms.

The role leads safety, security, and Responsible AI governance, partnering with RAI, data science, and platform teams to deliver scalable observability, policy enforcement, and incident response across cloud

Qualifications

  • Senior technical leader responsible for enterprise-grade AI observability platform design, deployment, and operation.
  • Owns OTEL integration across agentic platforms for unified observability and governance.
  • Leads safety, security, RAI, and red-teaming initiatives in observability workflows.

Responsibilities

  • Define enterprise observability architecture for AI agents, LLMs, and physical AI systems.
  • Build and operate telemetry pipelines with metrics, logs, traces, and real-time streaming (Kafka).
  • Instrument OTEL across platforms like Salesforce AgentForce, ServiceNow, and Microsoft Agent 365.
  • Deliver dashboards, alerts, SLOs/SLAs, runbooks, and RCA tooling to improve reliability.
  • Develop security observability playbooks and audit-ready traceability for AI decisions.

Skills

AI Observability
Distributed Systems
OpenTelemetry
Agentic AI Frameworks
Safety & Red Teaming
Reinforcement Learning
Python
MCP
A2A/UCP/AP2 Protocols
Cloud & Kubernetes

Education

Bachelor's or Master's in CS/AI/Data Science
12+ years technology experience
PhD a plus for research-heavy domains

Tools

Kafka
Terraform
Kubernetes
Azure/AWS/GCP
Salesforce/ServiceNow integrations

Job description

The AI Observability Architect is a senior technical leader responsible for designing, deploying, and operating an enterprise-grade, production-ready AI observability platform that spans the full spectrum of modern agentic AI — from large language model (LLM) workflows and multi-agent orchestration to physical AI systems, reinforcement learning harnesses, multi-modal pipelines, and agentic marketplaces. This role serves as the strategic and engineering authority for end-to-end telemetry, tracing, safety, and quality signals across heterogeneous agent frameworks and platforms.

The architect leads the convergence of AI observability with safety & security (including red teaming), Responsible AI (RAI), data science, physical AI, memory/skills engineering, agent fleet management, self-evolving harnesses, reinforcement learning, agent-to-agent protocols (A2A, UCP, AP2), and continuous quality engineering — making this a uniquely broad and high-impact role within the AI Solutions & Platforms organization

The role also owns OpenTelemetry (OTEL) integration across third-party agentic platforms (Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and others), enabling unified observability and governance at enterprise scale.

Responsibilities
  • Agentic AI Observability Architecture at Scale (30%)
  • Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems — covering planner/executor loops, tool/function calls, RAG retrieval chains, and memory/state transitions.
  • Build and operate unified telemetry pipelines incorporating metrics, logs, distributed traces, semantic/vector signals, and real-time event streaming (Kafka) at enterprise scale.
  • Instrument OpenTelemetry (OTEL) across heterogeneous platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and internal frameworks — delivering protocol-level observability for agent ecosystems including MCP, A2A, UCP, and AP2.
  • Design and implement observability for Agent Fleets, multi-modal pipelines, physical AI systems, and self-evolving reinforcement learning harnesses — including signal capture for reward shaping and policy evaluation.
  • Deliver dashboards, alerting, SLO/SLA management, incident runbook automation, and RCA tooling that drive measurable reliability improvements and reduce MTTR across agentic services.
  • Establish cost telemetry and FinOps observability for AI workloads — token consumption, inference cost allocation, and GPU/compute efficiency across cloud environments (Azure, AWS, GCP.)
  • Safety, Security & Red Teaming (15%)
  • Lead observability-driven red team exercises targeting agentic AI systems — instrumenting attack surfaces, adversarial prompt injection vectors, model evasion attempts, and multi-agent trust boundary failures.
  • Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy violation events, PII exposure risks, prompt leakage, and agent hallucination rates.
  • Partner with Security and RAI teams to embed threat modeling, zero-trust agent authentication, and behavioral anomaly detection into the observability platform.
  • Instrument secure policy enforcement layers across agent-to-agent communication protocols (A2A, UCP, AP2) and maintain audit-ready traceability for all AI decision events.
  • Develop and maintain a Security Observability Playbook covering incident classification, escalation paths, and forensic trace retention policies for agentic AI systems.
Responsible AI (RAI) & Governance (10
  • %)Integrate RAI signal capture — fairness, bias detection, explainability, and safety metrics — directly into observability pipelines, making compliance measurable and audit-read
  • y.Deliver governance dashboards that surface RAI compliance posture across all active AI agents and LLM deployments, aligned with global regulatory standard
  • s.Support risk assessments, gap analyses, and governance frameworks with real-time observability insights — enabling proactive risk mitigation rather than reactive audit response
  • s.Collaborate with RAI CoE and Legal/Compliance teams to define data retention, consent logging, and model decision traceability standards embedded in the telemetry architecture.
Quality Engineering for Agentic Solutions — Post Go-Live & Continuous QE (10%)
Memory, Skills, MCP & Harness Engineering Observability (10%)
Data Science Observability & Hardcore Python Engineering (5%)
Agentic Marketplace, Registry & Ecosystem Observability (5%)
Integration, Deployment & CI/CD Automation (5%)
Product Delivery & Stakeholder Collaboration (10%)
Compensation and Benefits:
  • The expected compensation range for this position is between $180,000 - $206,750.
  • Bonus based on performance and eligibility target payout is 15% of annual salary paid out annually.
  • Paid time off subject to eligibility, including paid parental leave, vacation, sick, and bereavement.
  • In addition to salary, PepsiCo offers a comprehensive benefits package to support our employees and their families, subject to elections and eligibility: Medical, Dental, Vision, Disability, Health, and Dependent Care Reimbursement Accounts, Employee Assistance Program (EAP), Insurance (Accident, Group Legal, Life), Defined Contribution Retirement Plan.
  • Bachelor's or Master's degree in Computer Science, AI/ML, Data Science, Software Engineering, or a related field (PhD a plus for research-heavy domains).
  • 12+ years in technology with deep experience in enterprise observability, distributed systems, platform engineering, or AI/ML infrastructure.
  • 5+ years in a senior/principal or architect-level role with demonstrated ownership of complex, cross-functional technical programs.
Core Technical Qualifications
  • AI Observability & Distributed Systems: Expert-level knowledge of observability primitives (metrics, logs, traces, events) applied to LLM/ML/agentic systems; hands-on OpenTelemetry (OTEL) instrumentation including custom exporters, semantic conventions, and trace propagation across agent/tool boundaries.
  • Agentic AI Frameworks: Direct experience with agentic AI platforms, multi-agent orchestration, LLM-based workflow design, and agent lifecycle management at production scale.
  • Safety, Security & Red Teaming: Demonstrated experience conducting red team exercises against AI systems; knowledge of adversarial attack patterns, prompt injection, model evasion, and multi-agent trust boundary failures; ability to design safety telemetry pipelines.
  • Memory, Skills & MCP: Working knowledge of agent memory architectures (episodic, semantic, working memory), Model Context Protocol (MCP), skill registries, and context injection patterns — with ability to design observability for these layers.
  • Agent-to-Agent Protocols: Familiarity with A2A (Agent-to-Agent), UCP (Universal Communication Protocol), and AP2 patterns; ability to implement protocol-level observability and policy enforcement.
  • Reinforcement Learning & Self-Evolving Harnesses: Understanding of RL training loops, reward signal capture, policy evaluation, and harness instrumentation for continuously improving agent systems.
  • Physical AI & Multi-Modal Systems: Experience or strong familiarity with observability for physical AI pipelines (robotics, edge inference, sensor fusion) and multi-modal models (vision, audio, text).
  • Data Science & Python Engineering: Proficiency in Python at a senior engineering level; experience with statistical anomaly detection, time-series analysis, and data pipeline design applied to observability data at scale.
  • Platform Integrations (OTEL / Enterprise): Hands-on experience integrating OTEL with enterprise agentic platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, or similar; strong understanding of enterprise integration patterns and API design.
  • Cloud & Infrastructure: Cloud fluency across Azure, AWS, and GCP; proficiency in Kubernetes, service mesh, IaC (Terraform/Bicep), and CI/CD tooling; experience with event streaming platforms (Kafka, Event Hubs).
  • Quality Engineering for AI: Experience designing continuous quality frameworks (CQE) for agentic solutions including eval harnesses, regression detection, quality gates, and SLA-backed quality benchmarking.
  • Responsible AI (RAI): Familiarity with RAI principles — fairness, bias detection, explainability, and safety — and ability to operationalize RAI signal capture within production observability pipelines.
  • Agentic Marketplace & Registry: Experience or strong familiarity with agent marketplace architectures, capability registries, and platform governance — ideally with observability or monitoring responsibilities for marketplace-registered components.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Architect/Engineer
Principal AI Architect/Engineer

PepsiCo • Chicago (IL)

On-site
USD 180,000 - 207,000
Bonus 15% of annual salary
Paid time off
Comprehensive benefits
+1
Senior Principal AI Architect/Engineer
Senior Principal AI Architect/Engineer

PepsiCo • Plano (TX)

On-site
USD 123,500 - 206,750
Bonus up to 15% of annual salary
Comprehensive benefits package including Medical, Dental, Vision
Paid time off including parental leave and vacation
Principal AI Architect/Engineer
Principal AI Architect/Engineer

PepsiCo • Plano (TX)

On-site
USD 110,700 - 185,250
Comprehensive benefits package
Paid time off including parental leave
AI Observability Architect for Enterprise Agentic Systems
AI Observability Architect for Enterprise Agentic Systems

PepsiCo • Purchase (NY)

On-site
USD 180,000 - 207,000
Bonus 15%
Paid time off
Comprehensive benefits
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Plano (TX)

On-site
USD 130,000 - 160,000
Senior AI Engineer
Senior AI Engineer

PepsiCo • Chicago (IL)

On-site
USD 110,000 - 186,000
Comprehensive benefits package
Paid time off including parental leave
Performance bonus
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Purchase (NY)

On-site
USD 120,000 - 160,000
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Chicago (IL)

On-site
USD 130,000 - 160,000
Sr Agentic AI Engineer
Sr Agentic AI Engineer

PepsiCo • Chicago (IL)

On-site
USD 106,000 - 179,000
Comprehensive benefits package
Paid time off
Bonus based on performance
AI Solutions Architect
AI Solutions Architect

PepsiCo • Plano (TX)

On-site
USD 110,700 - 185,250
Paid time off including parental leave, vacation, and sick leave
Comprehensive benefits package including medical, dental, vision
Bonus based on performance, target payout 12%