Engineer, Platform Engineering - AI

ICE Clear Europe Limited

Georgia

Hybrid

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

ICE Clear Europe Limited is seeking an AI Platform Engineer to implement and optimize GPU cluster infra, container tooling, and AI-enabled workflows. You will deploy vector stores, RAG pipelines, and MCP servers, while ensuring secure, scalable operations in a containerized environment.

You will join the AI Platform Operations team, translating architecture decisions into reliable infrastructure, with 24/7 production support and cross-team collaboration.

Qualifications

  • Bachelor's degree preferred.
  • 3+ years in infrastructure engineering, systems administration, or DevOps.
  • 3+ years scripting and automation (Python, Ansible, GitOps).
  • 3+ years hands-on Kubernetes in production.
  • 2+ years Linux administration.
  • Direct experience with GPU infrastructure (NVIDIA preferred).
  • 1+ years CUDA exposure.
  • 1+ years MCPs exposure.
  • 1+ years vector DBs and embedding infra.
  • 1+ years RAG pipeline design and deployment.
  • 1+ years agent memory patterns (context, external stores).
  • 1+ years agentic AI systems orchestration frameworks.
  • 1+ years semantic search and embedding models.
  • 1+ years workflow/orchestration automation tools.
  • Experience with enterprise monitoring & observability tools.
  • Ability to work in a service-oriented team environment.
  • PM, organization and time management.
  • Customer-focused with a strong user experience mindset.
  • Clear communication with technical and business resources.
  • Fluent English (spoken/written).

Responsibilities

  • Deploy, configure, and maintain GPU clusters and related infra.
  • Design, build, and maintain AI workflow automation platform.
  • Manage NVIDIA drivers, CUDA toolkits, and container runtimes.
  • Create and maintain ML framework container images (PyTorch, TensorFlow).
  • Implement monitoring, alerting, and observability for GPUs.
  • Maintain vector store infra for RAG pipelines and memory.
  • Develop end-to-end RAG workflows: ingestion, chunking, embeddings, retrieval.
  • Tune agent memory: short/long-term memory and episodic retrieval.
  • Deploy and operate Agentic AI systems and orchestration frameworks.
  • Host and maintain MCP servers within container platform.
  • Manage MCP configs, versioning, access controls, and integration.
  • Monitor MCP health, performance, and incidents; perform RCA.
  • Develop automation to improve platform reliability and efficiency.
  • Provide L2/L3 support and vendor escalation.
  • Implement security controls: network policies, RBAC, secrets.
  • Execute change requests and document technical details.
  • Support production operations in a 24/7 environment.
  • Coordinate with developers, operations, release engineers and end-users.
  • Educate and mentor team members and ops staff.
  • Participate in weekly on-call rotation for after-hours support.

Skills

Python
Ansible
GitOps
Kubernetes
Scripting

Education

Bachelor's degree (preferred)

Tools

NVIDIA GPUs
CUDA
Container runtimes
Vector databases
RAG tooling
MCP servers
CI/CD tooling

Job description

Job Purpose

We are on a mission as a team. We are problem solvers and partners, always starting with our customers to solve their challenges and create opportunities. Our start-up roots keep us nimble, flexible, and moving fast. We take ownership and make decisions. We all work for one company and work together to drive growth across the business. We engage in robust debates to find the best path, and then we move forward as one team. We take pride in what we do, acting with integrity and passion, so that our customers can perform better. We are experts and enthusiasts - combining ever-expanding knowledge with leading technology to consistently deliver results, solutions and opportunities for our customers and stakeholders. Every day we work toward transforming global markets.

The AI Platform Engineer is responsible for the technical implementation, maintenance, and optimization of AI/ML infrastructure. This hands-on role focuses on GPU cluster deployment, container image management, platform tooling development, and deep technical troubleshooting. In addition, the engineer deploys and maintains AI-enabled workflow automation tools across LLM, MCP, and agentic capabilities, ensuring these systems operate efficiently and securely within a containerized architecture. This includes deploying and maintaining vector store infrastructure, implementing end-to-end RAG workflows, tuning agent memory systems, and hosting and managing MCP servers. The engineer also deploys and operates Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines. The engineer serves as a core technical contributor on the AI Platform Operations team, translating architectural decisions into working infrastructure and enabling advanced, automated workflows across the platform.

Responsibilities
  • Deploy, configure, and maintain GPU clusters and associated infrastructure
  • Designing, building, and maintaining the workflow automation platform that uses AI capabilities (LLM/MCP/Agentic capabilities)
  • Manage NVIDIA driver versions, CUDA toolkits, and container runtimes
  • Build and maintain approved container images with ML frameworks (PyTorch, TensorFlow, etc.)
  • Implement monitoring, alerting, and observability for GPU infrastructure
  • Deploy and maintain vector store infrastructure for RAG pipelines, agent memory, and semantic search
  • Implement and maintain end-to-end RAG workflows, including document ingestion, chunking, embedding generation, and retrieval optimization
  • Maintain and tune agent memory systems, including short-term context windows, long-term persistent memory stores, and episodic memory retrieval patterns
  • Deploy, operate, and maintain Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines
  • Deploy, host, and maintain MCP servers within the containerized platform infrastructure
  • Manage MCP server configurations, versioning, access controls, and integration with agentic workflows
  • Monitor MCP server health, performance, and availability; respond to incidents and perform root cause analysis
  • Develop automation and tooling to improve platform reliability and efficiency
  • Provide L2/L3 technical support and vendor escalation for complex issues
  • Implement security controls including network policies, RBAC, and secrets management
  • Execute change requests and maintain technical documentation
  • Respond to and assist in production operations in a 24/7 environment
  • Provide technical analysis, resolve problems, and propose solutions
  • Provide support to, and coordinate with, developers, operations staff, release engineers, and end-users
  • Educate and mentor team members and operations staff
  • Participate in a weekly on-call rotation for after-hours support
Knowledge and Experience
  • Bachelor's degree preferred.
  • 3+ years in infrastructure engineering, systems administration, or DevOps
  • 3+ years in scripting and automation skills (Python, Ansible, GitOps)
  • 3+ years hands-on experience with Kubernetes in production
  • 2+ years experience with Linux administration
  • Direct experience with GPU infrastructure (NVIDIA preferred)
  • 1+ years experience using CUDA
  • 1+ years experience using MCPs
  • 1+ years experience with vector databases and embedding infrastructure
  • 1+ years experience with RAG pipeline design and deployment
  • 1+ years experience with agent memory patterns (in-context, external stores, retrieval-augmented memory)
  • 1+ years experience with agentic AI systems using orchestration frameworks
  • 1+ years experience with semantic search, embedding models, and ANN search techniques
  • 1+ years working with workflow/orchestrion automation tools
  • Experience with enterprise monitoring and observability tools
  • Ability to work in a service-oriented team environment
  • Project Management, organization, and time management
  • Customer focused, and dedicated to the best possible user experience
  • Communicate effectively with both technical and business resources
  • Fluent speaking, reading, and writing in English
Desired Knowledge and Experience
  • 1+ years of experience with AI developer toolkits (NVIDIA drivers, CUDA, cuDNN, and NCCL)
  • 1+ years of experience with Run:AI, NVIDIA AI Enterprise, or DGX systems
  • 1+ years of experience with n8n
  • 1+ years of experience with GitHub Actions

Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineer, Platform Engineering - AI
Engineer, Platform Engineering - AI

Intercontinental Exchange Holdings, Inc. • Atlanta (GA)

On-site
USD 150,000 - 230,000
Agentic AI Platform Senior Account Manager UK&I
Agentic AI Platform Senior Account Manager UK&I

NVIDIA Corporation • United States

Remote
GBP 130,000 - 200,000
Lead Principal Engineer, Enterprise Agentic AI Platform
Lead Principal Engineer, Enterprise Agentic AI Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
Senior Staff Software Engineer - Agentic Automation
Senior Staff Software Engineer - Agentic Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
AI Engineer
AI Engineer

Madison-Davis, LLC • New York (NY)

On-site
USD 140,000 - 210,000
AI Platform Engineer
AI Platform Engineer

The Intersect Group • Atlanta (GA)

On-site
USD 120,000 - 170,000
Senior Staff Software Engineer - Agentic Automation
Senior Staff Software Engineer - Agentic Automation

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Staff DevOps Engineer
Staff DevOps Engineer

Newmark • Chicago (IL)

On-site
USD 150,000 - 190,000
Hybrid options
Competitive compensation
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000