Metric library by intent
The seven intent groups from the paper. Under each are its prototypical metrics, the representative name for a set of related metrics, with a short description and the other names the same idea goes by. Select a metric to see every related metric in the full list.
All 118 metrics
Every metric extracted from the 39 sources, with its definition, a reference, and how it is evaluated. These are the same metrics summarized in the library, shown in full. CHAI’s use-case metrics are summarized here at the construct level; the full set of 196 CHAI metrics by use case is in the downloadable dataset.
Sources reviewed
The 39 sources in the review (oldest first): 34 peer-reviewed papers, 4 preprints, and 1 white paper.
Methodology
How the review was conducted. The paper points readers here for full detail.
Search strategy
We searched PubMed, IEEE Xplore, ACM Digital Library, and Google Scholar for papers published between January 2023 and June 2025 using the terms: ("large language model" OR "LLM" OR "generative AI" OR "ChatGPT" OR "GPT-4") AND ("evaluat*" OR "benchmark*" OR "metric*" OR "assess*") AND ("health*" OR "clinical" OR "medical" OR "biomedical"). We also conducted backward and forward citation searches on included papers.
Inclusion criteria
- Proposes, describes, or validates one or more named evaluation metrics or frameworks for LLM outputs in a healthcare context
- Provides sufficient detail to extract metric definitions (name, what is measured, how it is scored)
- Published in English as a peer-reviewed paper, accepted preprint, or formally released white paper/guideline
Exclusion criteria
- Papers that only apply existing metrics without proposing or adapting them for healthcare LLM evaluation
- Editorials, commentaries, and opinion pieces without original metrics
- Papers focused exclusively on non-text modalities (imaging, genomics) with no LLM text-generation evaluation
Data extraction and coding
For each included source, we extracted every named evaluation metric together with its definition, intent group (what aspect of quality it measures), evaluator type (human, automated, LLM judge, or organizational), the framework or study it belongs to, and the citation. Metrics that appear under different names across sources were grouped under a prototypical name in the metric library. The seven intent groups (technical performance, safety and harm, quality of information, understanding and reasoning, expression and communication, user experience and trust, system and deployment) emerged inductively from the extracted metrics and were validated by two authors.
Suggest an addition
This is a living resource. If you know of a framework, benchmark, or metric set we should include, let us know.