The landscape of LLM evaluation metrics in healthcare

Azra Ismaila, Manvi Sa, Sumon Kanti Deya, Nabile Safdarb, Melissa Sousab, Selvi Ramalingamb

a Emory University, Atlanta, GA  ·  b Emory Healthcare, Atlanta, GA

Companion page for a rapid review presented at EFMI STC 2026. Use the tabs to browse the metric library, the full set of 118 metrics, the 39 reviewed sources, or our methodology. This is a living resource; you can suggest additions below.

Cite (upcoming): A. Ismail, Manvi S, S.K. Dey, N. Safdar, M. Sousa, S. Ramalingam. The Landscape of LLM Evaluation Metrics in Healthcare: A Rapid Review. In: Studies in Health Technology and Informatics (EFMI STC 2026). IOS Press, forthcoming.

Download the full dataset (Excel)

Metric library by intent

The seven intent groups from the paper. Under each are its prototypical metrics, the representative name for a set of related metrics, with a short description and the other names the same idea goes by. Select a metric to see every related metric in the full list.

All 118 metrics

Every metric extracted from the 39 sources, with its definition, a reference, and how it is evaluated. These are the same metrics summarized in the library, shown in full. CHAI’s use-case metrics are summarized here at the construct level; the full set of 196 CHAI metrics by use case is in the downloadable dataset.

Human a person rates it
Automated scored by a fixed procedure
LLM judge an LLM rates it
Organizational reviewed by a team
Some metrics use more than one.

Sources reviewed

The 39 sources in the review (oldest first): 34 peer-reviewed papers, 4 preprints, and 1 white paper.

    Methodology

    How the review was conducted. The paper points readers here for full detail.

    Search strategy

    We searched PubMed, IEEE Xplore, ACM Digital Library, and Google Scholar for papers published between January 2023 and June 2025 using the terms: ("large language model" OR "LLM" OR "generative AI" OR "ChatGPT" OR "GPT-4") AND ("evaluat*" OR "benchmark*" OR "metric*" OR "assess*") AND ("health*" OR "clinical" OR "medical" OR "biomedical"). We also conducted backward and forward citation searches on included papers.

    Inclusion criteria

    Exclusion criteria

    Data extraction and coding

    For each included source, we extracted every named evaluation metric together with its definition, intent group (what aspect of quality it measures), evaluator type (human, automated, LLM judge, or organizational), the framework or study it belongs to, and the citation. Metrics that appear under different names across sources were grouped under a prototypical name in the metric library. The seven intent groups (technical performance, safety and harm, quality of information, understanding and reasoning, expression and communication, user experience and trust, system and deployment) emerged inductively from the extracted metrics and were validated by two authors.

    Suggest an addition

    This is a living resource. If you know of a framework, benchmark, or metric set we should include, let us know.