Principal AI Architect/Engineer

PepsiCo

Chicago (IL)

On-site

USD 180,000 - 206,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus 15% of annual salary
Paid time off
Comprehensive benefits
Tuition support

Job summary

PepsiCo seeks a senior AI Observability Architect to design, deploy, and operate an enterprise-grade observability platform spanning LLM workflows, multi-agent orchestration, and physical AI systems.

Leads integration of OTEL across major platforms, builds end-to-end telemetry pipelines, and drives safety, security, and RAI governance through measurable signals and dashboards.

Qualifications

  • Bachelor's or Master's degree in Computer Science, AI/ML, Data Science, Software Engineering, or a related field.
  • 12+ years in technology with deep experience in enterprise observability, distributed systems, platform engineering, or AI/ML infrastructure.
  • 5+ years in a senior/principal or architect-level role with ownership of complex cross-functional programs.

Responsibilities

  • Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems.
  • Build and operate unified telemetry pipelines including metrics, logs, traces, and real-time event streaming at enterprise scale.
  • Instrument OpenTelemetry across heterogeneous platforms; deliver protocol-level observability for agent ecosystems.
  • Design observability for Agent Fleets, multi-modal pipelines, physical AI systems, and RL harnesses.
  • Deliver dashboards, alerting, SLO/SLA management, and RCA tooling to improve reliability and reduce MTTR.
  • Establish cost telemetry and FinOps observability for AI workloads across cloud environments.

Skills

Observability
OpenTelemetry
Agent frameworks
Safety & Red Teaming
Reinforcement Learning
Python

Education

Bachelor's or Master's in CS/AI/DS/SE

Tools

OpenTelemetry
Kafka
Salesforce AgentForce
ServiceNow
Microsoft Agent 365

Job description

The AI Observability Architect is a senior technical leader responsible for designing, deploying, and operating an enterprise-grade, production-ready AI observability platform that spans the full spectrum of modern agentic AI — from large language model (LLM) workflows and multi-agent orchestration to physical AI systems, reinforcement learning harnesses, multi-modal pipelines, and agentic marketplaces. This role serves as the strategic and engineering authority for end-to-end telemetry, tracing, safety, and quality signals across heterogeneous agent frameworks and platforms.

The architect leads the convergence of AI observability with safety & security (including red teaming), Responsible AI (RAI), data science, physical AI, memory/skills engineering, agent fleet management, self-evolving harnesses, reinforcement learning, agent-to-agent protocols (A2A, UCP, AP2), and continuous quality engineering — making this a uniquely broad and high-impact role within the AI Solutions & Platforms organization

The role also owns OpenTelemetry (OTEL) integration across third-party agentic platforms (Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and others), enabling unified observability and governance at enterprise scale.

Responsibilities
  • Agentic AI Observability Architecture at Scale (30%)
  • Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems — covering planner/executor loops, tool/function calls, RAG retrieval chains, and memory/state transitions.
  • Build and operate unified telemetry pipelines incorporating metrics, logs, distributed traces, semantic/vector signals, and real-time event streaming (Kafka) at enterprise scale.
  • Instrument OpenTelemetry (OTEL) across heterogeneous platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and internal frameworks — delivering protocol-level observability for agent ecosystems including MCP, A2A, UCP, and AP2.
  • Design and implement observability for Agent Fleets, multi-modal pipelines, physical AI systems, and self-evolving reinforcement learning harnesses — including signal capture for reward shaping and policy evaluation.
  • Deliver dashboards, alerting, SLO/SLA management, incident runbook automation, and RCA tooling that drive measurable reliability improvements and reduce MTTR across agentic services.
  • Establish cost telemetry and FinOps observability for AI workloads — token consumption, inference cost allocation, and GPU/compute efficiency across cloud environments (Azure, AWS, GCP.)
  • Safety, Security & Red Teaming (15%)
  • Lead observability-driven red team exercises targeting agentic AI systems — instrumenting attack surfaces, adversarial prompt injection vectors, model evasion attempts, and multi-agent trust boundary failures.
  • Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy violation events, PII exposure risks, prompt leakage, and agent hallucination rates.
  • Partner with Security and RAI teams to embed threat modeling, zero-trust agent authentication, and behavioral anomaly detection into the observability platform.
  • Instrument secure policy enforcement layers across agent-to-agent communication protocols (A2A, UCP, AP2) and maintain audit-ready traceability for all AI decision events.
  • Develop and maintain a Security Observability Playbook covering incident classification, escalation paths, and forensic trace retention policies for agentic AI systems.
Responsible AI (RAI) & Governance (10
  • %)Integrate RAI signal capture — fairness, bias detection, explainability, and safety metrics — directly into observability pipelines, making compliance measurable and audit-read
  • y.Deliver governance dashboards that surface RAI compliance posture across all active AI agents and LLM deployments, aligned with global regulatory standard
  • s.Support risk assessments, gap analyses, and governance frameworks with real-time observability insights — enabling proactive risk mitigation rather than reactive audit response
  • s.Collaborate with RAI CoE and Legal/Compliance teams to define data retention, consent logging, and model decision traceability standards embedded in the telemetry architecture.
Quality Engineering for Agentic Solutions — Post Go-Live & Continuous QE (10%)
Memory, Skills, MCP & Harness Engineering Observability (10%)
Data Science Observability & Hardcore Python Engineering (5%)
Agentic Marketplace, Registry & Ecosystem Observability (5%)
Integration, Deployment & CI/CD Automation (5%)
Product Delivery & Stakeholder Collaboration (10%)
Compensation and Benefits:
  • The expected compensation range for this position is between $180,000 - $206,750.
  • Bonus based on performance and eligibility target payout is 15% of annual salary paid out annually.
  • Paid time off subject to eligibility, including paid parental leave, vacation, sick, and bereavement.
  • In addition to salary, PepsiCo offers a comprehensive benefits package to support our employees and their families, subject to elections and eligibility: Medical, Dental, Vision, Disability, Health, and Dependent Care Reimbursement Accounts, Employee Assistance Program (EAP), Insurance (Accident, Group Legal, Life), Defined Contribution Retirement Plan.
  • Bachelor's or Master's degree in Computer Science, AI/ML, Data Science, Software Engineering, or a related field (PhD a plus for research-heavy domains).
  • 12+ years in technology with deep experience in enterprise observability, distributed systems, platform engineering, or AI/ML infrastructure.
  • 5+ years in a senior/principal or architect-level role with demonstrated ownership of complex, cross-functional technical programs.
Core Technical Qualifications
  • AI Observability & Distributed Systems: Expert-level knowledge of observability primitives (metrics, logs, traces, events) applied to LLM/ML/agentic systems; hands-on OpenTelemetry (OTEL) instrumentation including custom exporters, semantic conventions, and trace propagation across agent/tool boundaries.
  • Agentic AI Frameworks: Direct experience with agentic AI platforms, multi-agent orchestration, LLM-based workflow design, and agent lifecycle management at production scale.
  • Safety, Security & Red Teaming: Demonstrated experience conducting red team exercises against AI systems; knowledge of adversarial attack patterns, prompt injection, model evasion, and multi-agent trust boundary failures; ability to design safety telemetry pipelines.
  • Memory, Skills & MCP: Working knowledge of agent memory architectures (episodic, semantic, working memory), Model Context Protocol (MCP), skill registries, and context injection patterns — with ability to design observability for these layers.
  • Agent-to-Agent Protocols: Familiarity with A2A (Agent-to-Agent), UCP (Universal Communication Protocol), and AP2 patterns; ability to implement protocol-level observability and policy enforcement.
  • Reinforcement Learning & Self-Evolving Harnesses: Understanding of RL training loops, reward signal capture, policy evaluation, and harness instrumentation for continuously improving agent systems.
  • Physical AI & Multi-Modal Systems: Experience or strong familiarity with observability for physical AI pipelines (robotics, edge inference, sensor fusion) and multi-modal models (vision, audio, text).
  • Data Science & Python Engineering: Proficiency in Python at a senior engineering level; experience with statistical anomaly detection, time-series analysis, and data pipeline design applied to observability data at scale.
  • Platform Integrations (OTEL / Enterprise): Hands-on experience integrating OTEL with enterprise agentic platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, or similar; strong understanding of enterprise integration patterns and API design.
  • Cloud & Infrastructure: Cloud fluency across Azure, AWS, and GCP; proficiency in Kubernetes, service mesh, IaC (Terraform/Bicep), and CI/CD tooling; experience with event streaming platforms (Kafka, Event Hubs).
  • Quality Engineering for AI: Experience designing continuous quality frameworks (CQE) for agentic solutions including eval harnesses, regression detection, quality gates, and SLA-backed quality benchmarking.
  • Responsible AI (RAI): Familiarity with RAI principles — fairness, bias detection, explainability, and safety — and ability to operationalize RAI signal capture within production observability pipelines.
  • Agentic Marketplace & Registry: Experience or strong familiarity with agent marketplace architectures, capability registries, and platform governance — ideally with observability or monitoring responsibilities for marketplace-registered components.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Architect/Engineer
Principal AI Architect/Engineer

PepsiCo • Purchase (NY)

On-site
USD 180,000 - 207,000
Bonus 15%
Paid time off
Comprehensive benefits
Senior Principal AI Architect/Engineer
Senior Principal AI Architect/Engineer

PepsiCo • Plano (TX)

On-site
USD 123,500 - 206,750
Bonus up to 15% of annual salary
Comprehensive benefits package including Medical, Dental, Vision
Paid time off including parental leave and vacation
Principal AI Architect/Engineer
Principal AI Architect/Engineer

PepsiCo • Plano (TX)

On-site
USD 110,700 - 185,250
Comprehensive benefits package
Paid time off including parental leave
AI Observability Architect for Enterprise Agentic Systems
AI Observability Architect for Enterprise Agentic Systems

PepsiCo • Purchase (NY)

On-site
USD 180,000 - 207,000
Bonus 15%
Paid time off
Comprehensive benefits
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Plano (TX)

On-site
USD 130,000 - 160,000
Senior AI Engineer
Senior AI Engineer

PepsiCo • Chicago (IL)

On-site
USD 110,000 - 186,000
Comprehensive benefits package
Paid time off including parental leave
Performance bonus
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Chicago (IL)

On-site
USD 130,000 - 160,000
Senior AI Engineer, Architect
Senior AI Engineer, Architect

PepsiCo • Purchase (NY)

On-site
USD 120,000 - 160,000
Sr Agentic AI Engineer
Sr Agentic AI Engineer

PepsiCo • Chicago (IL)

On-site
USD 106,000 - 179,000
Comprehensive benefits package
Paid time off
Bonus based on performance
AI Solutions Architect
AI Solutions Architect

PepsiCo • Plano (TX)

On-site
USD 110,700 - 185,250
Paid time off including parental leave, vacation, and sick leave
Comprehensive benefits package including medical, dental, vision
Bonus based on performance, target payout 12%