Edge AI/Model Optimization Engineer

Nextgenfed

Aberdeen (MD)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NextGen is seeking a skilled Edge AI/Model Optimization Engineer to support deployment, optimization, and sustainment of AI capabilities on edge and tactical computing platforms. You will evaluate, tune, benchmark, and operationalize large language models and embedding models for constrained hardware such as the X9 Spider Mission Computer architecture.

The role requires collaboration with AI engineers, systems integrators, and mission stakeholders to ensure AI-enabled capabilities stay

Qualifications

  • Bachelor's degree in a technical discipline.
  • 5+ years of AI/ML deployment and edge computing experience.
  • Experience deploying LLMs and AI inference in constrained environments.
  • Familiarity with CUDA, TensorRT, ONNX Runtime, or equivalent.
  • Experience tuning runtime configurations (quantization, batching, memory).
  • Proven ability to benchmark AI workloads on edge hardware.
  • Proficiency with Linux, containers, and orchestration tools.

Responsibilities

  • Evaluate LLMs and AI inference solutions for quality, latency, and memory on embedded GPUs.
  • Tune runtime configurations for edge deployment and hardware constraints.
  • Collaborate with stakeholders to assess mission requirements and platform tradeoffs.
  • Benchmark workflows and model-serving architectures against hardware limits.
  • Recommend model selection and runtime tradeoffs for mission effectiveness.
  • Develop repeatable performance, stress tests, and deployment validation workflows.
  • Package, deploy, and sustain local model-serving components for edge environments.
  • Coordinate with engineering and integration teams to ensure reliable agent behavior after optimizations.
  • Support edge deployments in tactical or disconnected environments.
  • Train customer personnel on supported profiles, constraints, and sustainability practices.
  • Maintain documentation and benchmarking results.

Skills

AI model optimization
Edge computing
GPU acceleration
Python
Linux & containers
CI/CD

Education

Bachelor's degree in CS/Engineering

Tools

CUDA
TensorRT
ONNX Runtime
vLLM
Docker
Kubernetes

Job description

NextGen is seeking a highly motivated and technically skilled Edge AI/Model Optimization Engineer to support the deployment, optimization, and sustainment of AI and agentic AI capabilities within edge and tactical computing environments. This role focuses on evaluating, tuning, benchmarking, and operationalizing Large Language Models (LLMs), embedding models, and AI inference services for constrained hardware platforms, including the X9 Spider Mission Computer architecture and other edge compute systems supporting operational missions using ReadiChat.

ReadiChat is a mission‑focused, agentic AI platform designed to help organizations build, deploy, govern, and scale specialized AI agents for operational workflows. It combines AI agents, workflow orchestration, grounded knowledge, testing frameworks, and enterprise controls into a single collaborative workspace.

The ideal candidate will possess expertise in AI model optimization, GPU‑enabled edge computing, runtime performance tuning, and operational AI deployment. This role requires close collaboration with AI engineers, systems integrators, mission stakeholders, and operational users to ensure AI‑enabled capabilities remain performant, reliable, and mission‑effective within disconnected, degraded, intermittent, and low‑bandwidth environments.

Responsibilities
  • Evaluate candidate LLMs, embedding models, and AI inference solutions for quality, latency, memory utilization, reliability, and operational performance on embedded GPU‑enabled edge compute platforms, including the X9 Spider Mission Computer architecture.
  • Tune and optimize AI model runtime configurations for edge deployment, including quantization strategies, batching configurations, context window sizing, cache behavior, inference scheduling, and GPU memory utilization specific to operational edge hardware environments.
  • Collaborate with customer stakeholders to assess mission requirements and evaluate alternative edge compute platforms when operational demands exceed X9 Spider capabilities or when cost, performance, power, size, weight, or thermal tradeoffs require additional analysis.
  • Benchmark agentic AI workflows, inference pipelines, and model‑serving architectures against target hardware constraints and operational performance thresholds.
  • Recommend model‑selection, runtime, and configuration tradeoffs balancing mission effectiveness, latency, throughput, resource utilization, reliability, and operational sustainability.
  • Build and maintain repeatable performance and stress‑testing frameworks for evaluating latency, throughput, tool‑call overhead, failover behavior, degraded‑resource conditions, and disconnected operational scenarios on edge compute platforms.
  • Package, deploy, validate, and sustain local model‑serving components and inference services to support reliable operation within tactical and edge environments.
  • Collaborate with agent engineers, AI developers, and integration teams to validate that agent behavior, workflow reliability, and operational outcomes remain acceptable following model compression, quantization, runtime optimization, or hardware configuration changes.
  • Support deployment, troubleshooting, optimization, and sustainment activities for AI‑enabled applications operating in edge, airborne, tactical, or disconnected operational environments.
  • Train customer technical personnel on supported model profiles, operational constraints, runtime tuning considerations, deployment limitations, troubleshooting procedures, and platform sustainment best practices.
  • Maintain technical documentation, benchmarking results, model validation reports, deployment procedures, optimization baselines, configuration guides, and operational support materials.
  • Support DevSecOps and CI/CD activities associated with AI model packaging, deployment automation, runtime validation, and operational release processes.
Qualifications
  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, Data Science, Artificial Intelligence, or related technical discipline.
  • 5+ years of experience supporting AI/ML deployment, model optimization, edge computing, GPU acceleration, or AI inference operations.
  • Experience deploying and optimizing LLMs, embedding models, or AI inference pipelines within resource‑constrained or edge‑compute environments.
  • Experience with GPU‑enabled systems and inference optimization technologies such as CUDA, TensorRT, ONNX Runtime, vLLM, Ollama, or equivalent platforms.
  • Experience tuning AI runtime configurations including quantization, batching, caching, and memory optimization techniques.
  • Experience benchmarking AI models and operational workflows against hardware performance constraints.
  • Experience with Linux‑based systems, containerized deployments, and orchestration technologies such as Docker and Kubernetes.
  • Familiarity with Python and AI/ML deployment frameworks commonly used for edge inference and operational AI systems.
  • Strong analytical, troubleshooting, and performance optimization skills.
  • Ability to communicate technical findings and operational tradeoffs effectively to technical and non‑technical stakeholders.
  • Active Security Clearance is required.
Additional Qualifications
  • Experience supporting tactical, airborne, or mission‑command edge computing environments.
  • Familiarity with X9 Spider Mission Computer architectures or similar embedded GPU‑enabled mission systems.
  • Experience supporting AI‑enabled workflows within NGC2, AIDP, EMSCO, Lattice, or related operational ecosystems.
  • Experience with model quantization techniques such as INT8, FP16, GGUF, GPTQ, AWQ, or similar optimization approaches.
  • Familiarity with disconnected, degraded, intermittent, and low‑bandwidth (DDIL) operational environments.
  • Experience with hardware evaluation and performance trade studies for operational edge compute systems.

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Edge AI/Model Optimization Engineer
Edge AI/Model Optimization Engineer

NextGen Federal Systems • Aberdeen (MD)

On-site
USD 90,000 - 120,000
Equal Opportunity Employer
Edge AI/Model Optimization Engineer
Edge AI/Model Optimization Engineer

NextGen Federal Systems • Aberdeen (WA)

On-site
USD 130,000 - 190,000
Edge AI Optimization Engineer for Tactical Systems
Edge AI Optimization Engineer for Tactical Systems

Nextgenfed • Aberdeen (MD)

On-site
USD 140,000 - 210,000
Edge AI Architect: LLM & Inference Optimization
Edge AI Architect: LLM & Inference Optimization

NextGen Federal Systems • Aberdeen (WA)

On-site
USD 130,000 - 190,000
Edge AI Engineer: LLMs & Model Optimization
Edge AI Engineer: LLMs & Model Optimization

NextGen Federal Systems • Aberdeen (MD)

On-site
USD 90,000 - 120,000
Equal Opportunity Employer
Lead Edge AI/ML Engineer
Lead Edge AI/ML Engineer

Arcfield • Home Creek (VA)

On-site
USD 101,000 - 201,000
MCP/Integration Engineer
MCP/Integration Engineer

NextGen Federal Systems • Aberdeen (MD)

On-site
USD 90,000 - 120,000
Director Embedded AI Engineering
Director Embedded AI Engineering

Honeywell • Atlanta (GA)

On-site
USD 120,000 - 150,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
MCP/Integration Engineer
MCP/Integration Engineer

Nextgenfed • Aberdeen (MD)

On-site
USD 90,000 - 120,000