Docs · 3 of 6

Reviewing evidence

Accept, edit, reject, or flag every extracted value from one forest-by-outcome view, and see exactly what each decision changes downstream.

The Evidence workbench is where you do the actual judgment of a review. Every value TrialExtract extracted from your trials is here, and your job is not to retype any of it. Your job is to decide, value by value, whether the machine got it right. Open a review and go to Evidence: the values are laid out grouped by outcome, with the forest plot as the structure rather than a spreadsheet of cells.

Why evidence is grouped by outcome

A meta-analysis is a set of outcomes, each measured across several studies. The workbench mirrors that: values are partitioned into one panel per outcome and measure family, so an outcome reported as a continuous change and again as a responder proportion becomes two panels, never one axis mixing incompatible measures. Inside a panel you see every study’s effect on the same scale. That is the view that lets you catch the value that sits impossibly far from its neighbours, which a flat table hides.

Click any value row to open its receipt in the column on the right: the source quote, the derivation, and whether independent verification ran. The receipt is how you check a value without leaving the review. See The receipt for what each part means.

The four decisions

Every value takes one of four decisions, and each has a one-key shortcut so you can move through a panel without reaching for the mouse:

  • Accept (A): the value is right as extracted.
  • Edit (E): the value is close but a number is wrong. The engine’s value stays on screen while you type the correction, so you are always correcting against what it read, not from memory.
  • Reject (R): the value does not belong in the analysis.
  • Needs fix (F): the value is flagged for follow-up before it can be trusted.

The chip on each row then reads Accepted, Edited, Rejected, or Needs fix. The workbench is keyboard-first: A / E / R / F to decide, arrow keys or J / K to move between rows, and number keys to pick a reason. Three clean accepts take about six seconds.

Reject and Needs fix need a reason

When you reject or flag a value, you are disagreeing with the extraction, and that disagreement has to be auditable later. So those two decisions cannot be submitted blank: you pick a coded reason from a fixed vocabulary, and the workbench blocks the submit until you do.

The reasons are:

  • Wrong value: the number itself is incorrect.
  • Wrong source: cited to the wrong table, figure, or passage.
  • Wrong derivation: the computation or method behind the value is wrong.
  • Missing data: a required input for this value is absent.
  • Unclear receipt: the provenance is insufficient to trust the value.
  • Other: none of the above; explain in the note.

Because the reason is coded, you can later query a review for every value you rejected as a wrong source, and your methods trail can say precisely why each excluded value was excluded. Accept and Edit do not require a reason.

What a decision changes

A decision is not cosmetic, but its effect lands downstream rather than on this screen. Every decision is recorded against the value, with the reviewer, the time, and, for a rejected or Needs fix, the coded reason you picked. Nothing is ever silently dropped.

The forest on this page stays labeled provisional. It plots and pools every value that carries a derivable effect, and it does not restate itself as confirmed when you accept a row. What your accepts and rejects actually build is the confirmed pool on the next surface: accepted values form the confirmed diamond, and the values you rejected or flagged are surfaced beside the forest there, each carrying the reason you gave. That is the line Pooling across studies builds on.

Filtering without lying about the count

The filter bar narrows the workbench by study type or outcome domain, and the count it shows is the real number of values that match. It is the same filter every surface of the review shares, so narrowing here narrows the quality view too. Nothing is padded and nothing is hidden behind the number, and it is honest about what it counts: values, not studies. One study contributes many values, so this count is never a stand-in for the study count in your write-up.