Research guide

Reporting chatbot health-advice studies with CHART

Specify CHART scope, model versions, prompts, variable outputs, human assessment and harm controls in health-advice chatbot research.

Comparing a manuscript with source material and highlighted passages
Prepared by: Dr. Didgar Research Institute · Last revised: · 3 min read · Guide created:
Expected deliverables
  • A version, prompt, repetition and rating protocol.
  • A human-assessment and failure-handling record.
  • CHART-aligned limits and sharing statements.
Workbook and completed example

Keep the scope specific

CHART (2025) addresses studies evaluating generative chatbots providing health advice or responses. It is not a universal AI checklist or medical approval of a product. Identify the study design and applicable base reporting guidance too.

Patient involvement or personal data requires design-appropriate ethics, consent and safety assessment. An educational response evaluation does not establish clinical effectiveness or permission to treat.

Make inputs and versions traceable

Record model name, observable version, access date, interface, available settings, accessible system instructions, exact prompts and conversation context. Explicitly state unavailable information in closed services rather than guessing.

Explain question origins, population, subject, difficulty and language. Do not upload real patient information to external services without authorization. Record browsing or retrieval conditions.

Variable outputs and assessment criteria

Hypothetical design: 20 synthetic scenarios, three repetitions per scenario per configuration. This is not a recommended minimum or power calculation; it distinguishes questions, outputs and repetitions. Retain permitted full outputs for review.

Define accuracy, completeness, uncertainty, risk and citation criteria before rating. Do not describe outputs as independent scenarios; account for dependence and rater agreement.

Human evaluation and a failure example

A fictional response may sound fluent while inventing a source or omitting the need for individual assessment. Fluency and medical correctness are separate criteria.

Describe rater expertise, training, blinding where possible, disagreements and adjudication. Plan handling of harmful content. Do not present the examples as patient advice.

Report changes and limits

Report units of analysis, failed outputs, criteria, dates and configurations. Do not conceal discarded failures. A model update does not inherit the previous version’s evaluation result.

Explain restricted prompts, data, code or outputs and permissible alternatives. This guide and its case do not rank actual medical model performance.

Completed teaching worksheet

This is a hypothetical teaching case, not observed data, an actual review or a publication acceptance. Numbers illustrate decisions.

Completed teaching worksheet
Decision or recordTeaching exampleYour project action
ScopeHealth-response evaluation; synthetic caseSpecify design and base guidance.
VersionObservable model and access timeState unknowns explicitly.
Unit20 scenarios × 3 repetitionsReport dependence and actual counts.
RatingAccuracy, harm and citation separatelyDefine rules and expertise.
FailureInvented citation in a fictional responseDocument output and handling.

Deliverables and completion checks

  • A version, prompt, repetition and rating protocol.
  • A human-assessment and failure-handling record.
  • CHART-aligned limits and sharing statements.

Sources and further reading

Official sources for verification and further reading

This guide supports research learning and planning; align implementation with the actual design and institutional requirements. Editorial policy
Back to top