Docs · 4 of 6

The quality queue

Instead of scanning every extracted number, work the ones most likely to be wrong first. The quality queue ranks each value by how much it needs you, and shows its reasoning.

The quality queue is where you stop reading every number and let the review point you at the ones that need a person. It looks at every extracted value in the review, scores each by how much it needs your judgment, and sorts the worklist so the most suspicious values are the first ones you touch. Nothing here is a verdict. It is a triage surface: it tells you where to look and why, and you make the call.

Four kinds of extracted item share this one worklist: evidence values, comparison statistics, adverse events, and risk-of-bias assessments. They are ranked and queued together, so a suspicious harm count sits in line next to a suspicious effect size.

Flagged for you

At the top of the surface, Flagged for you holds the single value most in need of a human right now. Next to it is a score, and the score is never a black box. It is the sum of named signals, and every finding carries the exact list of signals that added up to it, heaviest first. You can always answer the question “why is this 95?” by reading the decomposition underneath.

The signals, in weight order, heaviest first, are:

  • Fabrication pressure: TrialExtract was forced to fabricate at least one value in this study, so every value from it earns elevated scrutiny.
  • Quarantined: TrialExtract held the value back because it could not verify it against the source.
  • Verification differs: the independent second read landed on a different number than the extraction.
  • CI-gating contradiction: significance the confidence interval does not support (see below).
  • Stale decision: a prior decision went stale because the value changed under it.
  • Derived · unverified: computed from arm data rather than reported, and not yet independently verified.
  • No source quote: verified, but no verbatim quote could be pinned to check against the page.
  • Imputed N: the per-arm N was carried from the randomized or analysed counts, not restated in the results table.
  • QC flagged: an automated quality check marked the value for review.

Each signal is worth a fixed number of points, and the exact points behind any score are always spelled out in the decomposition under it. The scariest signals carry the most, so they sort a value to the top. Only values with at least one signal become findings. A clean value never nags you, and it never appears here.

The eight queues

Below the hero, the same items are filed into named queues. A queue is just a filter over the worklist, and it always reports an honest count, including zero. They run scariest and most actionable first:

  • Differs on verify: the independent read disagrees with the extracted number.
  • Fabrication flag: TrialExtract was forced to fabricate at least one value in this study.
  • Quarantined: held back, because TrialExtract could not verify it against the source.
  • Flagged: engine QC flagged this value for review.
  • Derived · unverified: computed from arm data, not reported, and not yet verified.
  • Missing source quote: verified, but no verbatim source quote could be pinned.
  • Needs reconfirm: a prior decision went stale because the value changed under it.
  • Needs review: awaiting your decision: accept, edit, or flag.

The CI-gating contradiction

One signal is worth calling out because it catches a specific, easy-to-miss error. The engine records its own significance flag for a value. Separately, the value carries a confidence interval. When the flag says significant but the interval includes the no-effect value (0 for a difference, 1 for a ratio), the two disagree, and the honest reading is the interval: by CI-gating, it is not significant. The queue surfaces that contradiction with the interval spelled out in the reason line, so you see exactly what does not add up. The gate itself is the same rule the pooling surface uses to decide what counts as significant.

What to look at first

When you open the surface, it lands you on the highest-priority queue that actually has something in it, so your first click is real QC truth rather than a catch-all. Work each value from the queue into its decision the same way you would anywhere else: accept it, edit it, or flag it, with its receipt one click away for the evidence behind the number. Deciding a value settles the Needs review queue for it, and re-confirming a stale decision settles Needs reconfirm. The data-keyed signals, though, a fabrication flag or a quarantine or a verification that differs, are read straight from the value’s own data, so they persist through your decision until that data itself changes.

When you have worked a queue down, its count reads a plain zero and the worklist panel says so in words: Nothing in that queue. When no value anywhere still carries an open signal, the Flagged for you banner lifts away rather than hanging an empty frame on the screen. That cleared state is the goal, a real, earned result rather than a blank screen. For how each decision is recorded and how a value can come back needing reconfirmation, see reviewing evidence.