- Keep the clean-run notebook, environment and execution log.
Write an analytical narrative
Explain the purpose, data provenance, transformations, assumptions and interpretation alongside the code. Jupyter combines text, code and outputs; an R reporting workflow should preserve the same traceable reasoning.
Eliminate hidden session state
Out-of-order cells can reuse stale variables or session-dependent results. Restart the environment and run every step in order. Stored output may belong to an older code version, so prepare a clean execution before publishing.
Record inputs and environment
Document language and package versions, environment setup, random seeds and data locations. A fixed seed alone does not guarantee identical behavior across all hardware. Provide permitted data or clearly labeled teaching data and explain access restrictions. Keep credentials and confidential data out of repositories.
Organize and check the workflow
Move reusable logic into functions and add meaningful checks for ranges, record counts and known answers. Use documented relative paths. Separate raw data, processed data and outputs and preserve transformation decisions.
Deliver a rerunnable package
Include a README, execution command, environment description, sample or accessible inputs, executed notebook and expected outputs. A colleague should not need undocumented steps. HTML and PDF aid reading but do not replace the code and environment needed for rerunning.
Practical research checklist
- Complete a clean ordered run.
- Record versions and setup.
- Explain inputs, outputs and provenance.
- Include documentation and scientific checks.
Worked case and implementation decisions
The following is a fictional teaching case. Do not use its numbers or wording as actual study findings.
A notebook can show attractive output while relying on a variable created only in an earlier session. Restart the kernel and execute all cells in order: an undefined name reveals a state or ordering problem. Old displayed output is not reproduction evidence. Record data, environment, relevant random seeds and execution commands; reconcile figures and text against fresh output. Successful execution on one machine does not guarantee identical behaviour on every system.
| Stage | Teaching example | Verification question |
|---|---|---|
| Clean start | Fresh kernel and cleared outputs | Any hidden state? |
| Order | Run all cells from the top | Independent of prior clicks? |
| Environment | Library and data versions | Are setup commands explicit? |
| Randomness | Seed and system limits | Are determinism claims justified? |
| Delivery | Fresh output and execution log | Do document numbers agree? |
Exercise output: Keep the clean-run notebook, environment and execution log.
Sources and further reading
Official sources for verification and further reading

