Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.
Qualifications
Knowledge of LLM evaluation and data curation methods.
Experience in designing LLM benchmarking methods.
Ability to shift focus as new findings emerge.
Responsibilities
Own LLM evaluation processes and methods for benchmarks.
Generate synthetic data and conduct benchmarking.
Deliver scalable and reproducible production code.
Develop new benchmarking methods for safety and helpfulness.
Co-author academic papers and presentations.
Skills
LLM evaluation
Data curation techniques
Designing benchmarking methods
Adaptability and flexibility
Job description
Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.