Prompt and Evaluation Engineer

Pop-Up Talent

United States

Hybrid

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401K with employer matching
Discretionary Time off
Paid holidays

Job summary

A leading technology recruitment firm is seeking a Prompt & Evaluation Engineer to design and evaluate prompts for AI models while collaborating across teams. The role emphasizes crafting effective prompts, developing evaluation frameworks, and integrating tools for reliability. Ideal candidates will bring 3–5 years of experience in applied AI or NLP, along with strong Python skills. This role offers a mid-senior level position with a full-time employment structure and opportunities for growth.

Qualifications

  • 3–5 years in applied AI, NLP, or related software engineering roles.
  • Strong Python skills for building eval pipelines.
  • Experience designing and iterating prompts for LLMs in production.

Responsibilities

  • Design, test, and refine prompts for LLMs.
  • Build and maintain evaluation frameworks.
  • Collaborate with teams to translate workflows into AI behaviors.

Skills

Applied AI
NLP
Python
Data analysis
LLM evaluation frameworks
Frameworks like LangChain
Strong communication skills

Tools

LangChain
LangGraph

Job description

Base Pay Range: $140,000.00/yr – $180,000.00/yr

Remote or Hybrid – U.S. Based

Reports to: Sr. Engineering

Prompt & Evaluation Engineer

WorkBoard's Strategy Execution Platform powers the digital operating rhythm for companies around the globe, providing clarity, alignment, and insights for growth.

We’re expanding our agentic AI capabilities and need a Prompt & Evaluation Engineer to design, optimize, evaluate, and scale the interaction layer between our AI models and enterprise users.

The Opportunity

As a Prompt & Evaluation Engineer, you'll specialize in crafting, testing, and evaluating prompts and structured workflows that make LLMs reliable, accurate, and contextually aware. You'll collaborate with product, engineering, and AI research teams to turn business requirements into effective, repeatable agentic behaviors.

What You'll Do

Prompt & Workflow Engineering

  • Design, test, and refine prompts for LLMs across a variety of enterprise use cases.
  • Develop structured multi‑step reasoning flows and evaluation harnesses for reliability.
  • Integrate tool‑calling strategies, safety checks, and contextual augmentation.

Evaluation & Reliability

  • Build and maintain evaluation frameworks to measure prompt effectiveness, accuracy, and safety.
  • Define success metrics and benchmarks for tool‑augmented LLM workflows.
  • Partner with research and product to scale continuous evaluation and feedback loops.

Tool & Framework Integration

  • Work with frameworks like LangChain, LangGraph, or similar to implement stateful agents.
  • Collaborate with engineers to align prompts with backend APIs and MCP‑based toolchains.
  • Ensure tool discovery and calling is consistent, robust, and observable in production.
  • Partner with Product, UX, and Engineering to translate workflows into AI behaviors.
  • Deliver at least one end‑to‑end agent conversation (e.g., get_objective) as a template for scaling.
  • Contribute to experimentation frameworks, logging, and evaluation dashboards.

Knowledge Sharing

  • Define best practices and guidelines for enterprise prompt engineering and evals.
  • Mentor engineers and PMs in prompt design and evaluation techniques.
What You Bring
  • At least 3–5 years in applied AI, NLP, or related software engineering roles.
  • Strong Python skills for building eval pipelines, data preprocessing, and experimentation.
  • Solid data experience: analyzing logs, designing benchmarks, and leveraging datasets to improve prompt quality.
  • Experience designing and iterating prompts for LLMs in production environments.
  • Hands‑on experience with LLM evaluation frameworks (e.g., LangSmith, custom eval harnesses).
  • Familiarity with frameworks like LangChain, LangGraph, or semantic caching methods.
  • Plus: knowledge of MCP toolchains and API‑to‑agent integration.
  • Strong communication skills and the ability to collaborate across disciplines.
  • A builder's mindset: creative, iterative, and outcome‑driven.
Within Your First 6 Months
  • 1 Month: Ramp up our agentic frameworks and learn existing prompt + eval libraries.
  • 3 Months: Own and deliver new prompt‑driven workflows with evaluation benchmarks.
  • 6 Months: Establish yourself as a go‑to expert in prompt engineering & evals, and publish internal best practices for scaling.
Our Values – We Live By the 4 Hs
  • Hungry for the opportunity
  • Intellectually honest
  • Operating as one happy team
  • Discretionary Time off & sick days
  • Paid holidaysHealth insurance
  • 401K with employer matching
  • Quarterly All‑Hands Meetings
  • And much more!

Seniority level: Mid‑Senior level

Employment type: Full‑time

Job function: Software Development

We are proud to be an equal opportunity workplace committed to building a team culture that celebrates learning, diversity, and inclusion. If you're hungry to grow your skills while growing a company, your sense of urgency matches the size of our market opportunity, and you value and enable teammates' contributions, then come join us! Our salary ranges are determined by role, level and geographic location. Within the range, individual pay is determined by additional factors, including job‑related skills, experience, and relevant education or training. Please note that the compensation details listed in US role postings reflect the base salary only, and do not include equity, or benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Prompt Engineer
Prompt Engineer

Internet Brands • El Segundo (CA)

On-site
USD 60,000 - 105,000
Health insurance
401(k) with company match
PTO and holidays
+1
Engineering Manager, Agent Prompts & Evals
Engineering Manager, Agent Prompts & Evals

Tactical Edge • Washington

On-site
USD 180,000 - 240,000
Flexible work setup
Senior Full Stack AI Engineer (Rapid Prototyping & Analytics )
Senior Full Stack AI Engineer (Rapid Prototyping & Analytics )

Prompt Health • United States

Hybrid
USD 160,000 - 220,000
Competitive salaries
Remote/hybrid environment
Flexible PTO
+4
Prompt Engineer
Prompt Engineer

The AI CEO • United States

Remote
USD 140,000 - 150,000
Senior Full Stack AI Engineer (Rapid Prototyping & Analytics )
Senior Full Stack AI Engineer (Rapid Prototyping & Analytics )

Launch Tennessee • United States

Remote
USD 160,000 - 220,000
Competitive salaries
Remote/hybrid environment
Flexible PTO
+4
Prompt Engineer
Prompt Engineer

Brooksource • Minnesota

On-site
USD 80,000 - 120,000
Prompt Engineer (AI Strategy & Analytics)
Prompt Engineer (AI Strategy & Analytics)

Mbi Llc • Harrisburg

On-site
USD 85,000 - 110,000
Prompt Engineer
Prompt Engineer

Steampunk • McLean (VA)

On-site
USD 115,000 - 140,000
Prompt Engineer
Prompt Engineer

Bright Vision Technologies • Novi (MI)

On-site
USD 76,000 - 99,000
Project Lion - Lead Prompt Engineer - United States (Remote, Part-Time)
Project Lion - Lead Prompt Engineer - United States (Remote, Part-Time)

Welo Global • United States

Remote
USD 140,000 - 200,000