AI Platform Engineer

Park Place Technologies in

Highland Heights (OH)

On-site

USD 90,000 - 130,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Park Place Technologies AI Platform Engineer is responsible for deploying and managing the computing and storage backbone powering AI models. Working on-site with AI Data Engineering, IT Infrastructure, Security, and business stakeholders, this role ensures reliability, performance, and scalability of our AI/ML infrastructure.

Responsibilities include provisioning CPU/GPU clusters, IaC with Terraform/Ansible, CI/CD pipelines and Kubernetes, data stores for LLMs, monitoring, and cross-functional

Qualifications

  • 2+ years in infrastructure/DevOps/platform engineering with some AI/ML exposure.
  • Proficiency in cloud computing, container orchestration, and infrastructure-as-code.
  • Knowledge of networking and distributed storage.
  • Experience with monitoring/observability tools.
  • Ability to troubleshoot complex systems and collaborate cross-functionally.
  • Foundational understanding of AI/ML concepts.

Responsibilities

  • Provision and maintain CPU/GPU compute clusters and high-bandwidth storage for AI/ML.
  • Develop and maintain IaC modules using Terraform and Ansible; participate in reviews.
  • Develop and maintain CI/CD pipelines and Kubernetes-based orchestration.
  • Set up and monitor data stores for LLMs and RAG pipelines; coordinate with data teams.
  • Monitor performance and logs; implement improvements with guidance from senior engineers.
  • Collaborate with AI Data Engineering, IT Infrastructure, Security, and business leaders.

Skills

Terraform
Ansible
Kubernetes
Flux
CI/CD
AI/ML

Tools

Terraform
Ansible
Kubernetes
Flux

Job description

The AI Platform Engineer is responsible for deploying and managing the computing and storage platform backbone that powers AI models. Working on-site alongside AI Data Engineering, IT Infrastructure, Security, and business stakeholders, this role is central to the reliability, performance, and scalability of our AI/ML infrastructure.

Responsibilities:
  • Provision and maintain CPU and GPU compute clusters, high-bandwidth network equipment, and scalable, high-throughput storage systems for AI/ML training and inference, working on-site and virtually to coordinate hands-on configuration and troubleshooting with infrastructure teams.
  • Develop and maintain scripts and infrastructure-as-code (IaC) modules for automated configuration of AI/ML infrastructure using tools such as Terraform and Ansible; participate in in-person and virtual code and design reviews with the platform engineering team.
  • Develop and maintain CI/CD pipelines and container orchestration platforms (Kubernetes, Flux); troubleshoot pipeline failures and elevate to more senior engineers or other teams as needed, collaborating in person and virtually to resolve issues quickly and effectively.
  • Set up and monitor data stores (file systems, block storage, traditional and vector databases) to support LLMs and RAG pipelines, with regular on-site coordination with AI Data Engineering teams to ensure data infrastructure meets evolving model requirements.
  • Monitor system performance and log metrics to ensure reliability and uptime; implement system improvements under guidance of more senior engineers, including attending in-person and virtual planning and incident review sessions.
  • Collaborate in person with AI Data Engineering, IT Infrastructure, Security, and Business Leaders to optimize infrastructure for performance and cost, including participating in cross-functional working sessions, architecture reviews, and stakeholder briefings.
  • Deploy and maintain Model Context Protocol (MCP) connectors and ensure secure, scalable infrastructure for MCP integrations, coordinating directly with security and engineering teams to validate configurations and address emerging requirements.
Basic Qualifications:
  • At least two (2) years of experience working in an infrastructure, DevOps, or platform engineering role, with at least some exposure to AI/ML workloads.
  • Proficiency in cloud computing, container orchestration, and infrastructure-as-code.
  • Knowledge of networking and distributed storage;
  • Experience with monitoring and observability tools to track performance and system health.
  • Ability to troubleshoot complex systems and collaborate in person with Data Engineering, Infrastructure, AI Engineering, Security, and Business Leaders.
  • Strong problem-solving, communication, and teamwork skills; comfortable working in a collaborative, on-site team environment.
  • Foundational understanding of AI/ML concepts.
Preferred Qualifications:
  • Knowledge of GPU acceleration.
Travel:
  • 10%
AI Platform Engineer (Architecture)
AI Platform Engineer

The AI Platform Engineer is responsible for deploying and managing the computing and storage platform backbone that powers AI models. Working on-site alongside AI Data Engineering, IT Infrastructure, Security, and business stakeholders, this role is central to the reliability, performance, and scalability of our AI/ML infrastructure.

Responsibilities:
  • Provision and maintain CPU and GPU compute clusters, high-bandwidth network equipment, and scalable, high-throughput storage systems for AI/ML training and inference, working on-site and virtually to coordinate hands-on configuration and troubleshooting with infrastructure teams.
  • Develop and maintain scripts and infrastructure-as-code (IaC) modules for automated configuration of AI/ML infrastructure using tools such as Terraform and Ansible; participate in in-person and virtual code and design reviews with the platform engineering team.
  • Develop and maintain CI/CD pipelines and container orchestration platforms (Kubernetes, Flux); troubleshoot pipeline failures and elevate to more senior engineers or other teams as needed, collaborating in person and virtually to resolve issues quickly and effectively.
  • Set up and monitor data stores (file systems, block storage, traditional and vector databases) to support LLMs and RAG pipelines, with regular on-site coordination with AI Data Engineering teams to ensure data infrastructure meets evolving model requirements.
  • Monitor system performance and log metrics to ensure reliability and uptime; implement system improvements under guidance of more senior engineers, including attending in-person and virtual planning and incident review sessions.
  • Collaborate in person with AI Data Engineering, IT Infrastructure, Security, and Business Leaders to optimize infrastructure for performance and cost, including participating in cross-functional working sessions, architecture reviews, and stakeholder briefings.
  • Deploy and maintain Model Context Protocol (MCP) connectors and ensure secure, scalable infrastructure for MCP integrations, coordinating directly with security and engineering teams to validate configurations and address emerging requirements.
Basic Qualifications:
  • At least two (2) years of experience working in an infrastructure, DevOps, or platform engineering role, with at least some exposure to AI/ML workloads.
  • Proficiency in cloud computing, container orchestration, and infrastructure-as-code.
  • Knowledge of networking and distributed storage;
  • Experience with monitoring and observability tools to track performance and system health.
  • Ability to troubleshoot complex systems and collaborate in person with Data Engineering, Infrastructure, AI Engineering, Security, and Business Leaders.
  • Strong problem-solving, communication, and teamwork skills; comfortable working in a collaborative, on-site team environment.
  • Foundational understanding of AI/ML concepts.
Preferred Qualifications:
  • Knowledge of GPU acceleration.
Travel:
  • 10%

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities
This employer is required to notify all applicants of their rights pursuant to federal employment laws.For further information, please review the Know Your Rights notice from the Department of Labor.IT/IS

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Planet Pharma • Indianapolis (IN)

On-site
USD 130,000 - 170,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Magen Financial LLC • Miami (FL)

Hybrid
USD 140,000 - 190,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Kansas Ag Connection • Kansas City (KS)

On-site
USD 140,000 - 190,000
AI Platform Engineer
AI Platform Engineer

Jobgether • Germany (OH)

On-site
USD 162,000 - 208,000
Fully remote opportunity
Equity ownership program
Technology allowance
+1
Head of AI Platform Engineering, Execution Services
Head of AI Platform Engineering, Execution Services

Soni • New York (NY)

On-site
USD 144,000 - 240,000
Sr. Systems Engineer - AI
Sr. Systems Engineer - AI

Dairy Farmers of America • Kansas City (KS)

On-site
USD 98,000 - 139,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
AI System Developer I
AI System Developer I

Tallgrass MLP Operations, LLC in • Houston (TX)

On-site
USD 65,000 - 90,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Capstone Investment Advisors • New York (NY)

On-site
USD 120,000 - 150,000