A leading AI technology firm in Redwood City is seeking an LLM Evaluations Engineering Lead. In this full-time position, you will be responsible for building evaluation systems for agentic LLMs, ensuring improved performance and reliability. Ideal candidates have strong software engineering skills and deep understanding of evaluation methodologies for machine learning. Join to work on impactful AI systems with the autonomy to shape their development.
Qualifications
Strong experience building evaluation systems for ML models, preferably LLMs.
Deep understanding of agentic failure modes such as tool misuse and hallucinated evidence.
Comfortable operating between research environments and production systems.
Responsibilities
Build eval harnesses for agentic LLM systems, both offline and in-workflow.
Design evaluations for planning, execution, recovery, and safety.
Implement verifier-driven scoring and regression gates.
Turn evaluation failures into useful training signals.
Skills
Building evaluation systems for ML models
Python
Data pipelines
Test harnesses
Distributed execution
Reproducibility
Understanding of agentic failure modes
Reasoning about metrics
Job description
A leading AI technology firm in Redwood City is seeking an LLM Evaluations Engineering Lead. In this full-time position, you will be responsible for building evaluation systems for agentic LLMs, ensuring improved performance and reliability. Ideal candidates have strong software engineering skills and deep understanding of evaluation methodologies for machine learning. Join to work on impactful AI systems with the autonomy to shape their development.