AI Platform Architect & Lead — Scale & Governance

Intercontinental Exchange (ICE)

Pleasanton (CA)

On-site

USD 149,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Intercontinental Exchange (ICE) is hiring an AI Platform Engineering Lead in California to guide the AI Platform Operations team, align architecture with business goals, and ensure secure, scalable AI/ML infrastructure. You will shape governance, manage budgets, and partner with stakeholders to drive 24/7 production readiness across LLM, MCP, and agentic capabilities.

The role demands deep expertise in GPU clusters, vector stores, RAG pipelines, and memory architectures, with a strong focus on

Qualifications

  • 8+ years in IT infrastructure or platform engineering roles.
  • 3+ years in technical leadership or management positions.
  • 1+ years hands-on experience with Kubernetes in production.
  • Direct experience with GPU infrastructure (NVIDIA preferred).
  • 2+ years experience using CUDA.
  • 1+ years experience using MCPs.
  • 2+ years experience with vector databases and embedding infrastructure.
  • 2+ years experience with RAG pipeline design and deployment.
  • 2+ years experience with agent memory patterns (in-context, external stores, retrieval-augmented memory).
  • 1+ years experience with agentic AI systems using orchestration frameworks.
  • 2+ years experience with semantic search, embedding models, and ANN search techniques.
  • 3+ years working with workflow/orchestrion automation tools.
  • Experience managing teams of 5+ technical staff.
  • Demonstrated success in vendor management and contract negotiation.
  • Strong executive communication and presentation skills.
  • Understanding of AI/ML workloads and infrastructure requirements.
  • Experience with enterprise monitoring and observability tools.
  • Ability to work in a service-oriented team environment.
  • Project Management, organization, and time management.
  • Customer focused, and dedicated to the best possible user experience.
  • Communicate effectively with both technical and business resources.
  • Fluent speaking, reading, and writing in English

Responsibilities

  • Define platform strategy, roadmap, and capability evolution.
  • Establish governance frameworks, policies, and exception processes.
  • Manage team budget, CapEx planning, and vendor relationships.
  • Build and lead the AI Platform Operations team.
  • Define the architecture, standards, and governance for AI/ML infrastructure, including GPU cluster design, compute resource planning, security controls, and observability across the platform.
  • Drive the strategy, standards, and governance for AI-enabled workflow automation across LLM, MCP, and agentic capabilities, ensuring the platform scales securely and operates with consistency across the enterprise.
  • Define the architecture and design standards for vector store infrastructure supporting RAG pipelines, agent memory, and semantic search across the enterprise.
  • Establish design patterns and best practices for RAG workflow implementation, including ingestion strategies, chunking approaches, embedding model selection, and retrieval optimization.
  • Architect agent memory frameworks, defining standards for short-term context, long-term persistent memory, and episodic memory patterns across AI platform workloads.
  • Drive the architecture and governance of Agentic AI systems, including multi-agent orchestration design and tool-use pipeline standards.
  • Define the strategy and architecture for hosting and managing MCP servers across the platform, including deployment topology, security boundaries, and integration standards.
  • Establish governance frameworks and policies for MCP server lifecycle management, versioning, and access control.
  • Evaluate and select MCP server tooling and vendors; manage relationships and roadmap alignment.
  • Serve as executive liaison for platform matters.
  • Own major incident management and executive communication.
  • Drive continuous improvement and platform maturity initiatives.
  • Align platform capabilities with enterprise architecture standards.
  • Respond to and assist in production operations in a 24/7 environment.
  • Provide technical analysis, resolve problems, and propose solutions.
  • Provide support to, and coordinate with, developers, operations staff, release engineers, and end-users
  • Educate and mentor team members and operations staff

Skills

Kubernetes
GPU infra
Vector DB
RAG pipelines
Agent memory
Orchestration
Vendor management

Tools

CUDA
MCPs
Observability tools

Job description

Intercontinental Exchange (ICE) is hiring an AI Platform Engineering Lead in California to guide the AI Platform Operations team, align architecture with business goals, and ensure secure, scalable AI/ML infrastructure. You will shape governance, manage budgets, and partner with stakeholders to drive 24/7 production readiness across LLM, MCP, and agentic capabilities.

The role demands deep expertise in GPU clusters, vector stores, RAG pipelines, and memory architectures, with a strong focus on

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineering Lead: Scale Secure AI Infra
AI Platform Engineering Lead: Scale Secure AI Infra

ICE • Chicago (IL)

On-site
USD 133,000 - 155,000
AI Platform Lead: Training & Inference at Scale
AI Platform Lead: Training & Inference at Scale

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
AI Platform Engineering Lead: Scale Secure AI Infra
AI Platform Engineering Lead: Scale Secure AI Infra

ICE • Atlanta (GA)

On-site
USD 133,000 - 180,000
AI Platform Engineer - Scalable Infra & LLM Ops
AI Platform Engineer - Scalable Infra & LLM Ops

International Materials, LLC • Delray Beach (FL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lead AI Platform Engineer — LLMs, Agents & RAG
Lead AI Platform Engineer — LLMs, Agents & RAG

Equinix • Redwood City (CA)

On-site
USD 142,000 - 212,000
Employee Assistance Program
US Benefits: health, life, disability,
retirement
Lead Engineer, Platform Engineering - AI
Lead Engineer, Platform Engineering - AI

ICE • New York (NY)

On-site
USD 149,000 - 180,000
Healthcare coverage
401(k) plan
Life insurance
Lead Engineer, Platform Engineering - AI
Lead Engineer, Platform Engineering - AI

ICE • Atlanta (GA)

On-site
USD 133,000 - 180,000
Healthcare coverage
401(k) plan
Life insurance
+1
Lead Engineer, Platform Engineering - AI
Lead Engineer, Platform Engineering - AI

ICE • Chicago (IL)

On-site
USD 133,000 - 155,000
Healthcare coverage (medical, dental and vision)
401(k) plan
Life insurance
+1
Senior AI Platform Engineer
Senior AI Platform Engineer

Mlg-Capital • Goerke's Corners (WI)

On-site
USD 150,000 - 210,000
AI Platform Architect — Lead Agentic Systems
AI Platform Architect — Lead Agentic Systems

Ridgeline • New York (NY)

On-site
USD 253,000 - 348,000
Multi-agent systems experience
Finance domain experience
Open-source contributions