AI Platform Engineer: Scale GPU AI for Researchers

The Johns Hopkins University Applied Physics Laboratory

Laurel (MD)

On-site

USD 85,000 - 165,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The Johns Hopkins University Applied Physics Laboratory (APL) seeks an AI Solutions Engineer to bridge researchers and a scalable AI computing platform. You will enable AI workloads from experimentation to large-scale GPU execution, working with Jupyter, containers, Kubernetes, and AI frameworks.

You will collaborate with researchers to move AI workloads into production, configure platform services, and optimize GPU utilization across distributed environments.

Qualifications

  • Bachelor of Science degree or equivalent years of related professional work experience.
  • At least three (3) years experience using container-based orchestration (e.g. Kubernetes) and run-time environments (e.g. Docker).
  • Hands-on experience with Jupyter environments and GPU-accelerated AI frameworks such as PyTorch, TensorFlow, Hugging Face, or similar technologies.
  • Working knowledge of Kubernetes and GPU computing concepts sufficient to configure, integrate, and troubleshoot AI applications running in shared computing environments.
  • Experience supporting researchers or developers with model training, fine-tuning, inference, workload optimization, and troubleshooting.
  • Proficiency with Python and shell scripting for automation, troubleshooting, and platform integration.
  • Working knowledge of machine learning concepts and model development workflows.
  • Strong problem-solving, communication, collaboration, prioritization, and continuous-learning skills.
  • Eligibility to obtain Interim Secret level security clearance and ultimately Secret level clearance; U.S. citizenship required.

Responsibilities

  • Own the technical configuration and evolution of the AI platform by evaluating new capabilities and establishing best practices.
  • Partner with Linux and Kubernetes admins on deployments, upgrades, and troubleshooting across AI platform and computing environment.
  • Provide hands-on technical assistance to researchers and data scientists using GPU-accelerated AI/ML platforms for model development, training, fine-tuning, and inference.
  • Troubleshoot AI workloads across the stack including Python environments, Jupyter, containers, Kubernetes, GPU scheduling, storage, networking, and AI frameworks.
  • Help researchers optimize GPU utilization and scale workloads from single-GPU to multi-GPU and multi-node distributed execution.
  • Develop and maintain reusable environments, workflows, automation, documentation, and operational guidance to improve platform reliability and researcher productivity.
  • Occasional weekend and after-hours work may be required to meet business needs.

Skills

Kubernetes
Docker
Python
Jupyter
GPU scheduling
PyTorch
TensorFlow
Shell scripting
ML concepts

Education

Bachelor's degree

Tools

NVIDIA CUDA
NVIDIA GPU Operator
NCCL
Kubeflow
MLflow
Weights & Biases

Job description

The Johns Hopkins University Applied Physics Laboratory (APL) seeks an AI Solutions Engineer to bridge researchers and a scalable AI computing platform. You will enable AI workloads from experimentation to large-scale GPU execution, working with Jupyter, containers, Kubernetes, and AI frameworks.

You will collaborate with researchers to move AI workloads into production, configure platform services, and optimize GPU utilization across distributed environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer: Scale GPU AI & Research
AI Platform Engineer: Scale GPU AI & Research

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 85,000 - 165,000
AI/ML Platform Engineer – Build Scalable AI Systems
AI/ML Platform Engineer – Build Scalable AI Systems

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 85,000 - 195,000
AI Solutions Engineer
AI Solutions Engineer

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 85,000 - 165,000
Impact AI/ML Scientist — Analytic Capabilities
Impact AI/ML Scientist — Analytic Capabilities

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
AI/ML Data Scientist for National-Impact Research
AI/ML Data Scientist for National-Impact Research

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
AI Solutions Engineer
AI Solutions Engineer

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 85,000 - 165,000
Senior Software Engineer - AI, Data Pipelines & Cloud
Senior Software Engineer - AI, Data Pipelines & Cloud

Talentify • Laurel (MD)

On-site
USD 100,000 - 245,000
AI-Powered Reverse Engineer & Workflow Architect
AI-Powered Reverse Engineer & Workflow Architect

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 100,000 - 245,000
AI Assurance Engineer - Mission-Critical Systems
AI Assurance Engineer - Mission-Critical Systems

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 100,000 - 245,000
AI/ML Platform Engineer: Build Production-Ready AI Systems
AI/ML Platform Engineer: Build Production-Ready AI Systems

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 85,000 - 195,000
Education assistance
Retirement contributions
Work/life balance