AI Solutions Engineer

Johns Hopkins Applied Physics Lab

Laurel (MD)

On-site

USD 85,000 - 165,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Johns Hopkins Applied Physics Laboratory seeks an AI Solutions Engineer to bridge researchers and a scalable AI computing platform. You will configure and evolve the AI stack, collaborate on platform deployments, and assist researchers with GPU-accelerated workloads for training, inference, and deployment.

Ideal candidates have 3+ years in Kubernetes/Docker, hands-on Jupyter, and experience with PyTorch, TensorFlow, or similar frameworks.

Qualifications

  • Bachelor of Science degree or equivalent professional experience.
  • 3+ years of container-based orchestration (Kubernetes) and runtime environments (Docker).
  • Hands-on experience with Jupyter environments and GPU-accelerated AI frameworks (PyTorch, TensorFlow, Hugging Face).
  • Knowledge of Kubernetes and GPU computing for configuring and troubleshooting AI apps in shared environments.
  • Experience supporting researchers with model training, fine-tuning, inference, and optimization.
  • Proficiency in Python and shell scripting for automation and platform integration.
  • Familiarity with ML workflows and research-to-production transitions.
  • Strong problem-solving, communication, collaboration, and prioritization skills.
  • Ability to obtain Interim Secret, eventually Secret clearance; U.S. citizenship required.

Responsibilities

  • Own the AI platform configuration and evolution, adopting new capabilities and setting standards.
  • Collaborate with Linux/Kubernetes admins on deployments, upgrades, and integration.
  • Provide hands-on support to researchers using GPU-accelerated AI platforms for model work.
  • Troubleshoot AI workloads across Python, Jupyter, containers, Kubernetes, and GPU components.
  • Help scale workloads from single-GPU experiments to multi-GPU multi-node runs.
  • Develop reusable environments, workflows, automation, and documentation to boost reliability.

Skills

Python
Shell scripting
GPU scheduling
AI frameworks
Distributed training
Research collaboration
Problem solving

Education

Bachelor of Science in CS or related field

Tools

Docker
Kubernetes
Jupyter

Job description

Description

Do you thrive in a fast-paced, dynamic environment?

Do you have a passion for AI systems and computing platforms that enable them?

Are you a continuous learner who enjoys solving complex technical problems and working collaboratively with researchers, developers, and infrastructure teams? If so, we're looking for someone like you to join our team at APL.

TheAI Solutions Engineercombines AI platform engineering with hands-on researcher enablement. You will serve as a technical bridge between researchers and the organization's scalable AI computing platform. You will also work directly with researchers to move AI workloads from experimentation to large-scale GPU execution, providing expertise across Jupyter, containers, Kubernetes, GPU scheduling, distributed training, and AI frameworks.

As an AI Solutions Engineer, you will …
  • Own the technical configuration and evolution of the AI platform by evaluating new capabilities, determining appropriate adoption, and establishing application-level configurations, standards, and best practices.
  • Partner with Linux and Kubernetes administrators on platform deployments, upgrades, infrastructure integration, and troubleshooting issues that span the AI platform and underlying computing environment.
  • Provide hands-on technical assistance to researchers and data scientists using GPU-accelerated AI/ML platforms for model development, training, fine-tuning, and inference.
  • Troubleshoot AI workloads across the stack, including Python environments, Jupyter, containers, Kubernetes, GPU scheduling and allocation, storage, networking, and AI frameworks.
  • Help researchers optimize GPU utilization and scale workloads from interactive or single-GPU experimentation to multi-GPU and multi-node distributed execution.
  • Develop and maintain reusable environments, workflows, automation, documentation, and operational guidance that improve platform reliability, usability, and researcher productivity.
Qualifications

You’ll meet our minimum qualifications if you…

  • Hold a Bachelor of Science degree or equivalent years of related professional work experience.
  • Have at least three (3) year experience using container-based orchestration (e.g. Kubernetes) and run-time environments (e.g. Docker).
  • Have hands- on experience with Jupyter environments and GPU-accelerated AI frameworks such as PyTorch, TensorFlow, Hugging Face, or similar technologies.
  • Have working knowledge of Kubernetes and GPU computing concepts sufficient to configure, integrate, and troubleshoot AI applications running in shared computing environments.
  • Have experience supporting researchers or developers with model training, fine-tuning, inference, workload optimization, and troubleshooting.
  • Have proficiency with Python and shell scripting for automation, troubleshooting, and platform integration.
  • Have working knowledge of machine learning concepts and model development workflows.
  • Can demonstrate strong problem-solving, communication, collaboration, prioritization, and continuous-learning skills.
  • Are able to work effectively with leadership to prioritize competing tasks.
  • Are able to obtain Interim Secret level security clearance by your start date and can ultimately obtain Secret level clearance. If selected, you will be subject to a government security clearance investigation and must meet the requirements for access to classified information. Eligibility requirements include U.S. citizenship.
Desired Qualifications:
  • Experience with AI workload orchestration and GPU scheduling platforms such as Run:ai, Slurm, Kubernetes-based GPU scheduling, or equivalent technologies.
  • Experience with NVIDIA GPU platforms and technologies such as CUDA, NVIDIA GPU Operator, NCCL, or NVIDIA Container Toolkit.
  • Experience supporting distributed AI training across multiple GPUs and/or multiple compute nodes.
  • Experience building and supporting containerized Jupyter environments and creating reusable AI/ML development environments.
  • Experience with AI/ML experiment tracking and lifecycle platforms such as MLflow, Weights & Biases, Kubeflow, or similar technologies.
  • Understanding of high-performance storage and networking considerations for GPU-intensive AI workloads.
  • Familiarity with a cloud computing platform (AWS, Azure or GCP).
Special Working Conditions:
  • Occasional weekend and other after-hours work may be required to handle and/or complete critical work to meet APL's business needs.
About Us

The Johns Hopkins University Applied Physics Laboratory (APL) brings world-class expertise to our nation’s most critical defense, security, space and science challenges. While we are dedicated to solving complex challenges and pioneering new technologies, what makes us truly outstanding is our culture. We offer a vibrant, welcoming atmosphere where you can bring your authentic self to work, continue to grow, and build strong connections with inspiring teammates.

At APL, we celebrate our differences of perspectives and encourage creativity and bold, new ideas. Our employees enjoy generous benefits, including a robust education assistance program, unparalleled retirement contributions, and a healthy work/life balance. APL’s campus is located in the Baltimore-Washington metro area. Learn more about our career opportunities at https://www.jhuapl.edu/careers.

All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, physical or mental disability, genetic information, veteran status, occupation, marital or familial status, political opinion, personal appearance, or any other characteristic protected by applicable law. APL is committed to providing reasonable accommodation to individuals of all abilities, including those with disabilities. If you require a reasonable accommodation to participate in any part of the hiring process, please contact Accessibility@jhuapl.edu.

The referenced pay range is based on JHU APL’s good faith belief at the time of posting. Actual compensation may vary based on factors such as geographic location, work experience, market conditions, education/training and skill level with consideration for internal parity. For salaried employees scheduled to work less than 40 hours per week, annual salary will be prorated based on the number of hours worked. APL may offer bonuses or other forms of compensation per internal policy and/or contractual designation. Additional compensation may be provided in the form of a sign-on bonus, relocation benefits, locality allowance or discretionary payments for exceptional performance. APL provides eligible staff with a comprehensive benefits package including retirement plans, paid time off, medical, dental, vision, life insurance, short-term disability, long-term disability, flexible spending accounts, education assistance, and training and development. Applications are accepted on a rolling basis.

Minimum Rate

$85,000 Annually

Maximum Rate

$165,000 Annually

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Solutions Engineer
AI Solutions Engineer

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 85,000 - 165,000
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 105,000 - 245,000
Data Scientist
Data Scientist

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 105,000 - 290,000
2027 BS/MS Graduate - AI/ML Research Engineer
2027 BS/MS Graduate - AI/ML Research Engineer

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 85,000 - 165,000
2027 PhD Graduate – Artificial Intelligence and Machine Learning (AI/ML) Research Scientist
2027 PhD Graduate – Artificial Intelligence and Machine Learning (AI/ML) Research Scientist

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
Education assistance
Retirement plans
Paid time off
2027 Internship - Tactical Intelligence Systems
2027 Internship - Tactical Intelligence Systems

Johns Hopkins Applied Physics Laboratory • Laurel (MD)

On-site
USD 47,000 - 100,000
Data Scientist
Data Scientist

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 290,000
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities
2027 PhD Graduate - AI/ML Data Scientist/Engineer - Analytic Capabilities

The Johns Hopkins University Applied Physics Laboratory • Laurel (MD)

On-site
USD 105,000 - 245,000
2027 Internship - Tactical Intelligence Systems
2027 Internship - Tactical Intelligence Systems

Johns Hopkins Applied Physics Lab • Laurel (MD)

On-site
USD 31,000 - 66,000