VLM Engineer

Aqua IT

Springfield (VA)

On-site

USD 150,000 - 190,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Aqua IT is seeking a seasoned ML Engineer to design and run fine-tuning pipelines for Vision-Language Models on domain imagery. You will build multimodal evaluation frameworks and scalable AWS-based training infrastructure for distributed fine-tuning of large models.

You will curate geospatial datasets, collaborate on architectures and adapters, and optimize inference. Strong Python, PyTorch, and HuggingFace expertise are essential for success.

Qualifications

  • 5+ years of professional ML engineering experience with a focus on deep learning.
  • 1+ years of hands-on experience fine-tuning large foundation models (LLMs or VLMs).
  • Experience with parameter-efficient fine-tuning methods (LoRA, QLoRA, adapters).
  • Familiarity with supervised fine-tuning, instruction tuning, and RLHF/DPO alignment techniques.
  • 4+ years of advanced Python development for ML workloads.

Responsibilities

  • Design and execute fine-tuning pipelines for Vision-Language Models on domain-specific imagery datasets.
  • Develop evaluation frameworks for multimodal model performance across image understanding, VQA, and spatial reasoning.
  • Build scalable training infrastructure on AWS (SageMaker, EC2 GPU instances) for distributed fine-tuning of large multimodal models.
  • Engineer data pipelines to curate, annotate, and transform geospatial imagery datasets for supervised and instruction-tuning workflows.
  • Collaborate with applied scientists and solutions architects to iterate on architectures, adapter strategies (LoRA/QLoRA), and inference optimization.

Skills

Python development
PyTorch
HuggingFace ecosystem
Distributed training
Model fine-tuning
CI/CD for ML

Tools

AWS SageMaker
DeepSpeed
FSDP
Megatron
TensorRT
ONNX

Job description

  • Design and execute fine-tuning pipelines for Vision-Language Models (VLMs) on domain-specific imagery datasets, including data preprocessing, training orchestration, and hyperparameter optimization
  • Develop and implement evaluation frameworks for multimodal model performance, including task-specific metrics for image understanding, visual question answering, and spatial reasoning
  • Build scalable training infrastructure on AWS (SageMaker, EC2 GPU instances) for distributed fine-tuning of large multimodal models
  • Engineer data pipelines for curating, annotating, and transforming geospatial imagery datasets into model-ready formats for supervised and instruction-tuning workflows
  • Collaborate with applied scientists and solutions architects to iterate on model architectures, adapter strategies (LoRA/QLoRA), and inference optimization techniques

Basic Requirements

  • TS/SCI with CI Poly required
  • 5+ years of professional machine learning engineering experience with a focus on deep learning
  • 1+ years of hands-on experience fine-tuning large foundation models (LLMs or VLMs)
  • Experience with parameter-efficient fine-tuning methods (LoRA, QLoRA, adapters)
  • Familiarity with supervised fine-tuning, instruction tuning, and RLHF/DPO alignment techniques
  • 4+ years of advanced Python development for ML workloads
  • Strong proficiency with PyTorch and the HuggingFace ecosystem (Transformers, PEFT, Datasets, Accelerate)
  • Experience with distributed training frameworks (DeepSpeed, FSDP, or Megatron)
  • 3+ years of experience with computer vision or multimodal models
  • Understanding of vision transformer architectures (ViT, CLIP, LLaVA-family models, or similar
  • Experience processing and augmenting image datasets at scale
  • 3+ years of experience with AWS ML infrastructure

    SageMaker Training jobs, Processing jobs, and endpoint deployment

    GPU instance selection, multi-node training, and cost optimization on EC2 (P4/P5/G5/G6e), S3 data management for large-scale training datasets

  • 2+ years of experience building ML evaluation pipelines Automated benchmarking, metric computation, and result analysis
  • Experience with both quantitative metrics and qualitative/human evaluation approaches
  • Strong software engineering fundamentals (version control, testing, CI/CD for ML workflows)

Preferred Qualifications:

  • 2+ years of experience with geospatial or remote sensing imagery
  • Familiarity with electro-optical and SAR satellite imagery formats and characteristics
  • Understanding of geospatial metadata, coordinate systems, and imagery preprocessing
  • Experience with model quantization and inference optimization (vLLM, TensorRT, ONNX)
  • Experience with MLOps and experiment tracking tools (MLflow, Weights & Biases, SageMaker Experiments)
  • Familiarity with data annotation platforms and active learning workflows for imagery
  • Experience with containerized ML workflows (Docker, ECR, ECS/EKS)
  • 2+ years of experience with Authority to Operate (ATO) processes in government environments
  • Implementation of NIST 800-53 controls and security compliance for ML systems
  • Experience deploying models in air-gapped or disconnected environments
  • Familiarity with multimodal evaluation benchmarks (MMMU, MMBench, GQA, or domain-specific equivalents)
  • Publications or demonstrated contributions in computer vision, VLMs, or multimodal AI
  • Experience with synthetic data generation for training data augmentation
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Computer Vision & AI/ML Engineer Jobs
Computer Vision & AI/ML Engineer Jobs

AIToolboard • Springfield (MA)

On-site
USD 150,000 - 230,000
Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000
Applied Machine Learning Engineer & Computer Vision
Applied Machine Learning Engineer & Computer Vision

Motion Recruitment • Arlington (VA), Northern (KY)

On-site
USD 150,000 - 210,000
AI Developer
AI Developer

Salvo Software LLC • United States

Remote
USD 150,000 - 230,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Clearview AI • United States

On-site
USD 150,000 - 200,000
Medical, Dental, Vision
STD and LTD Plans
Staff Computer Vision AI/ML Engineer
Staff Computer Vision AI/ML Engineer

Vets Hired • Aurora (CO)

On-site
USD 180,000 - 270,000
Computer Vision & Machine Learning Engineer New Remote, US
Computer Vision & Machine Learning Engineer New Remote, US

Buzz Solutions, Inc. • Northern (KY)

Remote
USD 130,000 - 180,000
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

On-site
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallelwireless • United States

On-site
USD 140,000 - 210,000