LLM Evaluation Engineer

ThirdLaw | Runtime AI Safety

United States

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Market cash compensation
Above-market equity
Generous benefits

Job summary

ThirdLaw | Runtime AI Safety is seeking an experienced AI Engineer to design and build the evaluation layer for their platform, ensuring the safety and compliance of LLM outputs. This role requires deep expertise in machine learning systems, foundation models, and real-time evaluation logic. Ideal candidates will have 7+ years of relevant experience, strong Python skills, and familiarity with technologies like Hugging Face Transformers. Competitive market compensation and benefits are provided for the right fit.

Qualifications

  • 7+ years in ML systems or AI engineering roles, with experience in LLMs or NLP.
  • Deep understanding of foundation models like OpenAI and Claude.
  • Strong proficiency in Python and experience with vector search techniques.

Responsibilities

  • Build real-time evaluation logic for LLM outputs.
  • Implement evaluation strategies using semantic similarity and classification.
  • Prototype and productize small language models for various applications.

Skills

ML systems experience
Foundation models understanding
Python proficiency
Experience with semantic similarity
Hands-on with vector search

Tools

Hugging Face Transformers
LangChain
PyTorch
TensorFlow

Job description

About The Company

ThirdLaw is building the control layer for AI in the enterprise. As companies rush to adopt LLMs and AI agents, they face new safety, compliance, and operational risks that traditional observability tools were never designed to detect. Metrics like latency or cost don’t capture when a model makes a bad decision, leaks sensitive data, or behaves unpredictably. We help IT and Security teams answer the foundational question: "Is this OK?"—and take real‑time action when it’s not. Backed by top‑tier venture firms and trusted by forward‑looking enterprise design partners, we’re building the infrastructure to monitor, evaluate, and control AI behavior in real‑world environments—at runtime, where it matters.

About The Role

You’ll build the evaluation layer in the ThirdLaw platform—the part of the system that decides whether an LLM prompt, response, tool call, or agent behavior is acceptable. This includes designing and tuning guardrails, classifiers, and semantic judgment systems that operate in real time. You'll integrate foundation models, similarity search, rules engines, and prompt templates to power high‑precision, low‑latency policy enforcement. This is not an experimental ML research role—it’s a product‑critical engineering role. You’ll work with structured trace data, foundation models, and real‑world constraints to build AI safety systems that actually ship.

What You’ll Do
  • Design and build real‑time evaluation logic that determines whether LLM prompts or outputs violate enterprise policies.
  • Implement evaluation strategies using a mix of semantic similarity, foundation model scoring, rule‑based systems, and statistical checks.
  • Integrate model outputs with downstream enforcement actions (e.g., redaction, escalation, blocking).
  • Prototype, tune, and productize small language models and prompt templates for classification, labeling, or scoring.
  • Collaborate with data infrastructure engineers to connect evaluation logic with ingestion and storage layers.
  • Build tools to observe, debug, and improve evaluator performance across real‑world data distributions.
  • Define abstractions for reusable evaluation components that can scale across use cases.
Requirements
  • 7+ years of experience in ML systems or AI engineering roles, with at least 1–2 years working directly with LLMs, NLP pipelines, or semantic search.
  • Deep understanding of foundation models (e.g., OpenAI, Claude, Mistral, Llama) and how to work with them via APIs or open source.
  • Hands‑on experience with vector search (e.g., FAISS, Qdrant, Weaviate) and embeddings pipelines.
  • Proven ability to implement real‑time or near‑real‑time evaluation logic using semantic similarity, classifier scoring, or structured rules.
  • Strong in Python, with familiarity using libraries like Hugging Face Transformers, LangChain, and PyTorch or TensorFlow.
  • Ability to reason about model behavior, test prompt configurations, and debug complex decision logic in production.
Nice‑to‑Have
  • Experience with OpenTelemetry, Model Context Protocol (MCP), or structured tracing of multi‑agent or multi‑model pipelines.
  • Experience with red‑teaming, AI risk taxonomies, or safety audits for LLM‑based systems.
  • Based in or willing to spend time in the San Francisco Bay Area for in‑person collaboration.
Why Apply?

Our team is small and focused, valuing autonomy and real impact over titles and management. We need strong technical skills, a proactive mindset, and clear written communication, as much of our work is asynchronous. If you're organized, take initiative, and want to work closely with customers to shape our products, you'll fit in well here.

Finally, we pay market cash compensation and generally above‑market equity. The compensation package for this role is benchmarked using Carta Total Compensation and reflects real‑time market data for our company’s size, this role’s level, and your geographic location. We have well‑designed and generous benefits.

https://www.thirdlaw.io/

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend / Platform Engineer, AI Analytic Engines
Backend / Platform Engineer, AI Analytic Engines

ThirdLaw | Runtime AI Safety • United States

On-site
USD 120,000 - 150,000
Market cash compensation
Above-market equity
Generous benefits
AI Engineer - FL
AI Engineer - FL

LawPro.ai • United States

On-site
USD 140,000 - 230,000
Full-stack AI Application Engineer
Full-stack AI Application Engineer

ThirdLaw | Runtime AI Safety • Berkeley (CA)

On-site
USD 100,000 - 140,000
Full-stack AI Application Engineer
Full-stack AI Application Engineer

ThirdLaw, Inc. • Berkeley (CA)

Hybrid
USD 100,000 - 150,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000