Mynd Healthcare

Report / Paper metadata

An external peer review instrument for Korean Medicine clinical practice guideline drafts (KM-EPRI): development and content validation

Researchers developed 117 prompts to help external reviewers examine Korean Medicine guideline drafts. A small expert panel endorsed their content, but the study did not test whether the prompts improve real reviews or make different reviewers more consistent.

What is the instrument for?

KM-EPRI is intended to structure review while a clinical practice guideline is still being developed. Its four modules address clinical questions, introductory material, development methods and recommendations. This differs from judging the overall quality of an already completed guideline. The instrument is a reference for free-text review, not a scoring system that automatically approves a guideline.

How were the prompts developed?

The team collected 3,630 external-review comments from 13 published guidelines developed under one Korean Medicine guideline manual. After excluding formatting and editorial comments, 2,073 content-related comments remained. Repeated points were combined within modules, then redundant or direct value-checking items were removed. The final instrument contains 117 items in 33 subdomains. An additional question about the module framework was evaluated separately and is not a 118th instrument item.

What did the expert panel assess?

All ten invited experts participated. They rated the proposed items on a five-point scale, with agreement defined in advance as at least 80% of the panel selecting one of the upper two categories. Every item met that threshold in the first round, so a second round was not conducted under the planned stopping rule. The average item content-validity index was 0.976. These numbers describe expert endorsement of the proposed content, not measured accuracy in reviewing an actual guideline.

Why do the agreement measures need care?

The reported Kendall concordance coefficient was low, at 0.196. The authors interpret this alongside the high ratings: limited variation between items can reduce a statistic concerned with relative ranking. That explanation should not be turned into a claim that practical reviewer reliability has been demonstrated. The study did not measure whether two reviewers using the instrument reach the same judgments on real drafts.

What limits transfer to other settings?

The panel and source comments came from the same specialist community, so development and validation were not fully independent. Patients and public representatives were not included. Common past criticisms received more attention than rarely raised problems, and the 117-item workload could be substantial. The study took place within one national guideline system; adoption elsewhere would need adaptation and testing rather than assuming that these prompts fit every clinical field.

What remains to be tested?

The paper leaves practical feasibility, inter-reviewer reliability and effects on review quality unmeasured. Construct and criterion validity also remain future work. Content validation is a useful development step, but it is not evidence that a reviewed guideline improves patient outcomes or that its treatment recommendations are correct. This explanation uses the full methodological paper and does not reproduce its item list, tables or figures.

Kind
Paper metadata
Identifier
PMC13642740

View original record ↗

Worldwide health knowledge / Healthcare records