- A version, prompt, repetition and rating protocol.
- A human-assessment and failure-handling record.
- CHART-aligned limits and sharing statements.
Keep the scope specific
CHART (2025) addresses studies evaluating generative chatbots providing health advice or responses. It is not a universal AI checklist or medical approval of a product. Identify the study design and applicable base reporting guidance too.
Patient involvement or personal data requires design-appropriate ethics, consent and safety assessment. An educational response evaluation does not establish clinical effectiveness or permission to treat.
Make inputs and versions traceable
Record model name, observable version, access date, interface, available settings, accessible system instructions, exact prompts and conversation context. Explicitly state unavailable information in closed services rather than guessing.
Explain question origins, population, subject, difficulty and language. Do not upload real patient information to external services without authorization. Record browsing or retrieval conditions.
Variable outputs and assessment criteria
Hypothetical design: 20 synthetic scenarios, three repetitions per scenario per configuration. This is not a recommended minimum or power calculation; it distinguishes questions, outputs and repetitions. Retain permitted full outputs for review.
Define accuracy, completeness, uncertainty, risk and citation criteria before rating. Do not describe outputs as independent scenarios; account for dependence and rater agreement.
Human evaluation and a failure example
A fictional response may sound fluent while inventing a source or omitting the need for individual assessment. Fluency and medical correctness are separate criteria.
Describe rater expertise, training, blinding where possible, disagreements and adjudication. Plan handling of harmful content. Do not present the examples as patient advice.
Report changes and limits
Report units of analysis, failed outputs, criteria, dates and configurations. Do not conceal discarded failures. A model update does not inherit the previous version’s evaluation result.
Explain restricted prompts, data, code or outputs and permissible alternatives. This guide and its case do not rank actual medical model performance.
Completed teaching worksheet
This is a hypothetical teaching case, not observed data, an actual review or a publication acceptance. Numbers illustrate decisions.
| Decision or record | Teaching example | Your project action |
|---|---|---|
| Scope | Health-response evaluation; synthetic case | Specify design and base guidance. |
| Version | Observable model and access time | State unknowns explicitly. |
| Unit | 20 scenarios × 3 repetitions | Report dependence and actual counts. |
| Rating | Accuracy, harm and citation separately | Define rules and expertise. |
| Failure | Invented citation in a fictional response | Document output and handling. |
Deliverables and completion checks
- A version, prompt, repetition and rating protocol.
- A human-assessment and failure-handling record.
- CHART-aligned limits and sharing statements.
Version and scope references: EQUATOR — CHART statement (2025) · EQUATOR Network — Reporting guidelines
Sources and further reading
Official sources for verification and further reading
- EQUATOR — CHART statement (2025) Source verification: 2026-10-06
- EQUATOR Network — Reporting guidelines
- ICMJE — Roles of authors and contributors

