Engineering Manager, Agent Prompts & Evals

Tactical Edge

Washington (District of Columbia)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work setup

Job summary

Tactical Edge in Washington, DC is seeking a hands-on engineering manager to lead the team responsible for prompt engineering, model evaluation, and AI quality assurance across our AI products. You will own the systems that ensure our agents produce reliable, accurate, and safe outputs.

This player-coach role sets the technical direction, builds the evaluation infrastructure, and grows a team of specialists while continuing to write code and review prompts hands-on, with a strong focus on LLM

Qualifications

  • 5+ years engineering experience with 2+ years managing teams.
  • Deep experience with LLM prompting, evaluation, and optimization.
  • Familiarity with eval frameworks (Braintrust, Langsmith, or custom).
  • Production AI systems experience with observability and monitoring.
  • Understanding of model architectures, tokenization, and inference optimization.

Responsibilities

  • Manage and grow a team of prompt engineers and eval specialists.
  • Set standards for prompt design, evaluation methodology, and quality metrics across all products.
  • Design and operate eval pipelines that measure accuracy, safety, hallucination rates, and task completion across all AI agents.
  • Build automated benchmarks and regression suites.
  • Run A/B tests, cost-quality tradeoff analysis, and latency benchmarks.
  • Own the quality bar for AI outputs.
  • Implement red-teaming, adversarial testing, and safety evaluations.
  • Build dashboards for monitoring AI quality in production.
  • Work with product, engineering, and customer teams to translate requirements into evaluation criteria and prompt strategies.

Skills

LLM prompting
Evaluation
Team leadership
Observability
Monitoring
Prompt engineering
Cross-functional collaboration

Tools

Braintrust
Langsmith
Custom eval framework

Job description

We're looking for a hands-on engineering manager to lead the team responsible for prompt engineering, model evaluation, and AI quality assurance across all Tactical Edge products. You'll own the systems that ensure our AI agents produce reliable, accurate, and safe outputs.

This is a player-coach role — you set the technical direction, build the evaluation infrastructure, and hold the quality bar while growing a team of specialists.

Key Responsibilities
  • Manage and grow a team of prompt engineers and eval specialists.
  • Set standards for prompt design, evaluation methodology, and quality metrics across all products.
Prompt Engineering at Scale
  • Build and maintain prompt libraries, templates, and versioning systems.
  • Establish best practices for system prompts, few-shot examples, chain-of-thought reasoning, and tool-use instructions.
Evaluation Systems
  • Design and operate eval pipelines that measure accuracy, safety, hallucination rates, and task completion across all AI agents.
  • Build automated benchmarks and regression suites.
  • Run A/B tests, cost-quality tradeoff analysis, and latency benchmarks.
Quality & Safety
  • Own the quality bar for AI outputs.
  • Implement red-teaming, adversarial testing, and safety evaluations.
  • Build dashboards for monitoring AI quality in production.
Cross-functional Collaboration
  • Work with product, engineering, and customer teams to translate requirements into evaluation criteria and prompt strategies.
  • Technical manager who still writes code and reviews prompts hands-on.
  • Deep understanding of LLM behavior, failure modes, and edge cases.
  • Data-driven decision maker who builds systems to measure what matters.
  • Strong communicator who can translate AI quality concepts for non-technical stakeholders.
  • High bar for quality with a pragmatic approach to shipping.
Preferred Qualifications
  • 5+ years engineering experience with 2+ years managing teams.
  • Deep experience with LLM prompting, evaluation, and optimization.
  • Familiarity with eval frameworks (Braintrust, Langsmith, or custom).
  • Production AI systems experience with observability and monitoring.
  • Understanding of model architectures, tokenization, and inference optimization.
How We Work

Outcome-driven

Enterprise-first

Agentic by design

Systems that reason and act safely

Small teams, high ownership

Autonomy with accountability

What You'll Get
  • Work on real, production AI deployments
  • Enterprise-scale challenges and measurable impact
  • Cross-functional collaboration and high ownership
  • Flexible work setup where applicable
Hiring Process
  1. Intro call

    Fit + context

  2. Technical discussion

    Prompt engineering, eval design, model selection

  3. Leadership & systems interview

    Team management, cross-functional collaboration

  4. Final conversation

    Alignment + next steps

We value clarity, ownership, and thoughtful execution over buzzwords.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Prompting & AI Evaluation
Engineering Manager, Prompting & AI Evaluation

Tactical Edge • Washington

On-site
USD 180,000 - 240,000
Flexible work setup
Prompt and Evaluation Engineer
Prompt and Evaluation Engineer

Pop-Up Talent • United States

Hybrid
USD 140,000 - 180,000
Health insurance
401K with employer matching
Discretionary Time off
+1
Prompt Engineer
Prompt Engineer

Lockedinai • New York (NY)

Hybrid
USD 100,000 - 130,000
Meaningful early-stage equity
Impactful role in product development
Remote-first work culture
+1
Lead Prompt Engineer, AI Solutions and Development
Lead Prompt Engineer, AI Solutions and Development

Disney • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Senior Applied AI Engineer
Senior Applied AI Engineer

Level • Austin (TX)

On-site
USD 120,000 - 150,000
AI Engineer
AI Engineer

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available
AI Prompt & Agent Developer
AI Prompt & Agent Developer

ProLegion • San Francisco (CA)

On-site
USD 120,000 - 190,000
Principal Engineer - Context Engineering & LLM Optimization
Principal Engineer - Context Engineering & LLM Optimization

Hobbsnews • Charlotte (NC)

On-site
USD 130,000 - 160,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Prompt Engineer
Prompt Engineer

Odiin.AI • New York (NY)

On-site
USD 120,000 - 160,000