AI Solutions Engineer

The Johns Hopkins University Applied Physics Laboratory

Laurel (MD)

On-site

USD 85,000 - 165,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

The Johns Hopkins University Applied Physics Laboratory (APL) seeks an AI Solutions Engineer to bridge researchers and a scalable AI computing platform. You will enable AI workloads from experimentation to large-scale GPU execution, working with Jupyter, containers, Kubernetes, and AI frameworks.

You will collaborate with researchers to move AI workloads into production, configure platform services, and optimize GPU utilization across distributed environments.

Qualifications

  • Bachelor of Science degree or equivalent years of related professional work experience.
  • At least three (3) years experience using container-based orchestration (e.g. Kubernetes) and run-time environments (e.g. Docker).
  • Hands-on experience with Jupyter environments and GPU-accelerated AI frameworks such as PyTorch, TensorFlow, Hugging Face, or similar technologies.
  • Working knowledge of Kubernetes and GPU computing concepts sufficient to configure, integrate, and troubleshoot AI applications running in shared computing environments.
  • Experience supporting researchers or developers with model training, fine-tuning, inference, workload optimization, and troubleshooting.
  • Proficiency with Python and shell scripting for automation, troubleshooting, and platform integration.
  • Working knowledge of machine learning concepts and model development workflows.
  • Strong problem-solving, communication, collaboration, prioritization, and continuous-learning skills.
  • Eligibility to obtain Interim Secret level security clearance and ultimately Secret level clearance; U.S. citizenship required.

Responsibilities

  • Own the technical configuration and evolution of the AI platform by evaluating new capabilities and establishing best practices.
  • Partner with Linux and Kubernetes admins on deployments, upgrades, and troubleshooting across AI platform and computing environment.
  • Provide hands-on technical assistance to researchers and data scientists using GPU-accelerated AI/ML platforms for model development, training, fine-tuning, and inference.
  • Troubleshoot AI workloads across the stack including Python environments, Jupyter, containers, Kubernetes, GPU scheduling, storage, networking, and AI frameworks.
  • Help researchers optimize GPU utilization and scale workloads from single-GPU to multi-GPU and multi-node distributed execution.
  • Develop and maintain reusable environments, workflows, automation, documentation, and operational guidance to improve platform reliability and researcher productivity.
  • Occasional weekend and after-hours work may be required to meet business needs.

Skills

Kubernetes
Docker
Python
Jupyter
GPU scheduling
PyTorch
TensorFlow
Shell scripting
ML concepts

Education

Bachelor's degree

Tools

NVIDIA CUDA
NVIDIA GPU Operator
NCCL
Kubeflow
MLflow
Weights & Biases

Job description

Description

Do you thrive in a fast-paced, dynamic environment?

Do you have a passion for AI systems and computing platforms that enable them?

Are you a continuous learner who enjoys solving complex technical problems and working collaboratively with researchers, developers, and infrastructure teams? If so, we're looking for someone like you to join our team at APL.

The AI Solutions Engineer combines AI platform engineering with hands-on researcher enablement. You will serve as a technical bridge between researchers and the organization's scalable AI computing platform. You will also work directly with researchers to move AI workloads from experimentation to large-scale GPU execution, providing expertise across Jupyter, containers, Kubernetes, GPU scheduling, distributed training, and AI frameworks.

As an AI Solutions Engineer, you will ...

  • Own the technical configuration and evolution of the AI platform by evaluating new capabilities, determining appropriate adoption, and establishing application-level configurations, standards, and best practices.
  • Partner with Linux and Kubernetes administrators on platform deployments, upgrades, infrastructure integration, and troubleshooting issues that span the AI platform and underlying computing environment.
  • Provide hands-on technical assistance to researchers and data scientists using GPU-accelerated AI/ML platforms for model development, training, fine-tuning, and inference.
  • Troubleshoot AI workloads across the stack, including Python environments, Jupyter, containers, Kubernetes, GPU scheduling and allocation, storage, networking, and AI frameworks.
  • Help researchers optimize GPU utilization and scale workloads from interactive or single-GPU experimentation to multi-GPU and multi-node distributed execution.
  • Develop and maintain reusable environments, workflows, automation, documentation, and operational guidance that improve platform reliability, usability, and researcher productivity.
Qualifications

You'll meet our minimum qualifications if you...

  • Hold a Bachelor of Science degree or equivalent years of related professional work experience.
  • Have at least three (3) year experience using container-based orchestration (e.g. Kubernetes) and run-time environments (e.g. Docker).
  • Have hands-on experience with Jupyter environments and GPU-accelerated AI frameworks such as PyTorch, TensorFlow, Hugging Face, or similar technologies.
  • Have working knowledge of Kubernetes and GPU computing concepts sufficient to configure, integrate, and troubleshoot AI applications running in shared computing environments.
  • Have experience supporting researchers or developers with model training, fine-tuning, inference, workload optimization, and troubleshooting.
  • Have proficiency with Python and shell scripting for automation, troubleshooting, and platform integration.
  • Have working knowledge of machine learning concepts and model development workflows.
  • Can demonstrate strong problem-solving, communication, collaboration, prioritization, and continuous-learning skills.
  • Are able to work effectively with leadership to prioritize competing tasks.
  • Are able to obtain Interim Secret level security clearance by your start date and can ultimately obtain Secret level clearance. If selected, you will be subject to a government security clearance investigation and must meet the requirements for access to classified information. Eligibility requirements include U.S. citizenship.
Desired Qualifications
  • Experience with AI workload orchestration and GPU scheduling platforms such as Run:ai, Slurm, Kubernetes-based GPU scheduling, or equivalent technologies.
  • Experience with NVIDIA GPU platforms and technologies such as CUDA, NVIDIA GPU Operator, NCCL, or NVIDIA Container Toolkit.
  • Experience supporting distributed AI training across multiple GPUs and/or multiple compute nodes.
  • Experience building and supporting containerized Jupyter environments and creating reusable AI/ML development environments.
  • Experience with AI/ML experiment tracking and lifecycle platforms such as MLflow, Weights & Biases, Kubeflow, or similar technologies.
  • Understanding of high-performance storage and networking considerations for GPU-intensive AI workloads.
  • Familiarity with a cloud computing platform (AWS, Azure or GCP).
Special Working Conditions
  • Occasional weekend and other after-hours work may be required to handle and/or complete critical work to meet APL's business needs.
About Us
Why Work at APL?

The Johns Hopkins University Applied Physics Laboratory (APL) brings world-class expertise to our nation's most critical defense, security, space and science challenges. While we are dedicated to solving complex challenges and pioneering new technologies, what makes us truly outstanding is our culture. We offer a vibrant, welcoming atmosphere where you can bring your authentic self to work, continue to grow, and build strong connections with inspiring teammates.

At APL, we celebrate our differences of perspectives and encourage creativity and bold, new ideas. Our employees enjoy generous benefits, including a robust education assistance program, unparalleled retirement contributions, and a healthy work/life balance. APL's campus is located in the Baltimore-Washington metro area. Learn more about our career opportunities at https://www.jhuapl.edu/careers.

All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, physical or mental disability, genetic information, veteran status, occupation, marital or familial status, political opinion, personal appearance, or any other characteristic protected by applicable law.APL is committed to providing reasonable accommodation to individuals of all abilities, including those with disabilities. If you require a reasonable accommodation to participate in any part of the hiring process, please contact Accessibility@jhuapl.edu.

The referenced pay range is based on JHU APL's good faith belief at the time of posting. Actual compensation may vary based on factors such as geographic location, work experience, market conditions, education/training and skill level with consideration for internal parity. For salaried employees scheduled to work less than 40 hours per week, annual salary will be prorated based on the number of hours worked. APL may offer bonuses or other forms of compensation per internal policy and/or contractual designation. Additional compensation may be provided in the form of a sign-on bonus, relocation benefits, locality allowance or discretionary payments for exceptional performance. APL provides eligible staff with a comprehensive benefits package including retirement plans, paid time off, medical, dental, vision, life insurance, short-term disability, long-term disability, flexible spending accounts, education assistance, and training and development. Applications are accepted on a rolling basis.

Minimum Rate

$85,000 Annually

Maximum Rate

$165,000 Annually

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Solutions Engineer
AI Solutions Engineer

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 85,000 - 165,000
Senior Software Engineer
Senior Software Engineer

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 100,000 - 245,000
Generous benefits
Career development opportunities
2027 PhD Graduate – Artificial Intelligence and Machine Learning (AI/ML) Research Scientist
2027 PhD Graduate – Artificial Intelligence and Machine Learning (AI/ML) Research Scientist

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
Education assistance
Retirement plans
Paid time off
Data Scientist
Data Scientist

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 290,000
2027 PhD Graduate - Artificial Intelligence and Machine Learning (AI/ML) Research Scientist
2027 PhD Graduate - Artificial Intelligence and Machine Learning (AI/ML) Research Scientist

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
Education assistance
Retirement contributions
Work-life balance
Senior Artificial Intelligence Researcher
Senior Artificial Intelligence Researcher

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 290,000
Data Scientist
Data Scientist

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 290,000
2027 Internship - Tactical Intelligence Systems
2027 Internship - Tactical Intelligence Systems

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 47,000 - 100,000
2027 Graduate - Developer - Cyber-Physical Systems
2027 Graduate - Developer - Cyber-Physical Systems

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 85,000 - 165,000
Education assistance
Retirement contributions
Healthy work/life balance
2027 Internship - AI & Data Science Intern - Analytic Capabilities
2027 Internship - AI & Data Science Intern - Analytic Capabilities

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 31,000 - 66,000
On-site internship