AI Platform Engineer: GPU Infra & MLOps

ICE Clear Europe Limited

Georgia

Hybrid

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ICE Clear Europe Limited is seeking an AI Platform Engineer to implement and optimize GPU cluster infra, container tooling, and AI-enabled workflows. You will deploy vector stores, RAG pipelines, and MCP servers, while ensuring secure, scalable operations in a containerized environment.

You will join the AI Platform Operations team, translating architecture decisions into reliable infrastructure, with 24/7 production support and cross-team collaboration.

Qualifications

  • Bachelor's degree preferred.
  • 3+ years in infrastructure engineering, systems administration, or DevOps.
  • 3+ years scripting and automation (Python, Ansible, GitOps).
  • 3+ years hands-on Kubernetes in production.
  • 2+ years Linux administration.
  • Direct experience with GPU infrastructure (NVIDIA preferred).
  • 1+ years CUDA exposure.
  • 1+ years MCPs exposure.
  • 1+ years vector DBs and embedding infra.
  • 1+ years RAG pipeline design and deployment.
  • 1+ years agent memory patterns (context, external stores).
  • 1+ years agentic AI systems orchestration frameworks.
  • 1+ years semantic search and embedding models.
  • 1+ years workflow/orchestration automation tools.
  • Experience with enterprise monitoring & observability tools.
  • Ability to work in a service-oriented team environment.
  • PM, organization and time management.
  • Customer-focused with a strong user experience mindset.
  • Clear communication with technical and business resources.
  • Fluent English (spoken/written).

Responsibilities

  • Deploy, configure, and maintain GPU clusters and related infra.
  • Design, build, and maintain AI workflow automation platform.
  • Manage NVIDIA drivers, CUDA toolkits, and container runtimes.
  • Create and maintain ML framework container images (PyTorch, TensorFlow).
  • Implement monitoring, alerting, and observability for GPUs.
  • Maintain vector store infra for RAG pipelines and memory.
  • Develop end-to-end RAG workflows: ingestion, chunking, embeddings, retrieval.
  • Tune agent memory: short/long-term memory and episodic retrieval.
  • Deploy and operate Agentic AI systems and orchestration frameworks.
  • Host and maintain MCP servers within container platform.
  • Manage MCP configs, versioning, access controls, and integration.
  • Monitor MCP health, performance, and incidents; perform RCA.
  • Develop automation to improve platform reliability and efficiency.
  • Provide L2/L3 support and vendor escalation.
  • Implement security controls: network policies, RBAC, secrets.
  • Execute change requests and document technical details.
  • Support production operations in a 24/7 environment.
  • Coordinate with developers, operations, release engineers and end-users.
  • Educate and mentor team members and ops staff.
  • Participate in weekly on-call rotation for after-hours support.

Skills

Python
Ansible
GitOps
Kubernetes
Scripting

Education

Bachelor's degree (preferred)

Tools

NVIDIA GPUs
CUDA
Container runtimes
Vector databases
RAG tooling
MCP servers
CI/CD tooling

Job description

ICE Clear Europe Limited is seeking an AI Platform Engineer to implement and optimize GPU cluster infra, container tooling, and AI-enabled workflows. You will deploy vector stores, RAG pipelines, and MCP servers, while ensuring secure, scalable operations in a containerized environment.

You will join the AI Platform Operations team, translating architecture decisions into reliable infrastructure, with 24/7 production support and cross-team collaboration.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer: GPU, Kubernetes & RAG Workflows
AI Platform Engineer: GPU, Kubernetes & RAG Workflows

Intercontinental Exchange Holdings, Inc. • Atlanta (GA)

On-site
USD 150,000 - 230,000
AI Infrastructure Engineer — Cloud, GPU & MLOps
AI Infrastructure Engineer — Cloud, GPU & MLOps

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
AI Infrastructure Engineer: GPU-HPC Clusters & MLOps
AI Infrastructure Engineer: GPU-HPC Clusters & MLOps

Freelio • Northern (KY)

Hybrid
USD 90,000 - 230,000
Senior LLMOps Platform Engineer — GPU AI Infra
Senior LLMOps Platform Engineer — GPU AI Infra

Quantum Technologies. LLC • Jersey City (NJ), Northern (KY)

Hybrid
USD 120,000 - 165,000
AI Platform Engineer — Azure, GPUs, Secure Infra
AI Platform Engineer — Azure, GPUs, Secure Infra

Motion Recruitment Partners, LLC • Miami (FL)

Hybrid
USD 120,000 - 160,000
Hybrid work environment
AI Platform Engineer - Scalable Infra & LLM Ops
AI Platform Engineer - Scalable Infra & LLM Ops

International Materials, LLC • Delray Beach (FL), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
GPU Platform Engineer for ML Ops on Kubernetes (GCP)
GPU Platform Engineer for ML Ops on Kubernetes (GCP)

SmartRecruiters, Inc. • Paris (TX)

On-site
USD 103,000 - 148,000
E-learning platform
Game library
Clubs and activities
ML Infrastructure Engineer – GPU Compute Platform
ML Infrastructure Engineer – GPU Compute Platform

Remanence • Paris (TX)

Hybrid
USD 124,000 - 186,000
Visa sponsorship
Relocation support
Hybrid work setup
+1