APPRAISE-AI
APPRAISE-AI is a scoring tool that helps assess how good a study is when using artificial intelligence to support clinical decisions. It looks at things like how relevant the study is to real-world medicine, how trustworthy the data is, and whether results are clearly reported. It gives a score out of 100 so different AI studies can be compared fairly.
At a glance
Use when
Evaluating quality of AI-based clinical prediction studies during peer review, funding assessment, or systematic reviews
Avoid when
Assessing non-prediction AI applications (e.g., image generation) or non-clinical AI research
Inputs
Primary research studies on AI-based clinical prediction models, particularly in development, silent, or clinical trial phases
Outputs
A quantitative quality score (0–100) with domain-level and overall assessments, identifying strengths and weaknesses in AI studies
How it works
APPRAISE-AI is a quantitative evaluation tool designed to assess methodological and reporting quality of AI prediction models for clinical decision support. It evaluates studies across six domains: clinical relevance, data quality, methodological conduct, robustness of results, reporting quality, and reproducibility. The tool comprises 24 scored items with a maximum total score of 100. It was validated through interrater and intrarater reliability testing and correlated with expert scores, citation rates, QUADAS-2 bias assessment, and TRIPOD adherence. The tool was applied in a systematic review of sepsis prediction models.
- HTA domains
- Clinical Effectiveness, Organisational aspects, Patient and Social Aspects
- Categories
- Appraisal
- Assumptions
- Higher scores on the 24-item scale reflect better methodological and reporting quality; consistency in rater interpretation is achievable; quality correlates with citation rates and expert judgment
- Strengths
- Demonstrates high interrater and intrarater reliability; correlates strongly with expert judgment, TRIPOD adherence, and citation rates; provides a standardized, quantitative comparison across AI studies
- Limitations
- Validated only in the context of sepsis prediction models; may not generalize to all AI applications in healthcare; focuses on published reporting, potentially missing protocol-level flaws
- Also known as
- APPRAISE-AI Tool
Questions this answers
- › How clinically relevant is the AI model being studied?
- › How robust and transparent is the study's methodology and reporting?
- › How reproducible are the results of the AI study?
- › How does the overall quality of an AI study compare to others in the same field?
- › What aspects of an AI study are commonly underreported or weak?
- › How well does the study adhere to existing reporting guidelines like TRIPOD?
References & sources
Similar by meaning
Beta record. Generated from the primary source via AI extraction and independent audit, pending final human review.

