Automated data extraction for systematic reviews
Screening's done.
Now it's just you and forty PDFs.
Every mean, SD, and N, located and typed by hand. The SD that turns out to be an SEM. The subgroup that lives only in a figure. The one transposed digit that reaches print with your name on it. TrialExtract does that pass in minutes, and hands you a receipt for every number: the sentence it read, the formula it ran, the check that ran. You review instead of retype.
Founding cohort access. A free preview of any paper you bring, before you spend a credit.
Free preview on every paper before you spend a credit · exports to metafor & RevMan
The table you've been dreading. Already filled.
Every outcome from the PDFs, one row apiece: the native effect, the standardized SMD, and where each was read.
The receipt behind every number
The numbers are the machine's job. The judgment's yours.
Open any value and it answers for itself. Every cell in that table traces straight back to the paper: the exact sentence it was read from, the formula that computed it, and whether a second reading agreed.
Click a tab for the three shapes a number arrives in: reported and verified, recovered from the arm means, or recovered from a confidence interval. This is what you'd be signing off on.
Only per-arm means reported. The effect is computed, every step shown.
Table 1, AVLT Delayed Recall (# of words): Bacopa 7.6 (3.9) vs Placebo 6.9 (4.2) at +12 Weeks (F=5.4; 1,21; p=0.03*). Narrative: “Controlling for baseline cognitive deficit using the Blessed Orientation–Memory–Concentration test, Bacopa participants had enhanced AVLT delayed word recall memory scores relative to placebo.”
Not located in the quote (not restated here): 24 Bacopa n,24 placebo n
- 1 Reported arm summaries Bacopa: n=24, mean=7.6, SD=3.9 · placebo: n=24, mean=6.9, SD=4.2 N=48
- 2 Pooled SD (Cochrane 5.3.a) √[((24−1)·3.9² + (24−1)·4.2²) / (24+24−2)] s_pooled = 4.0528
- 3 Mean difference 7.6 − 6.9 Δ = 0.7
- 4 Cohen's d Δ / s_pooled = 0.7 / 4.0528 d = 0.1727
- 5 Hedges small-sample correction (Hedges & Olkin 1985) g = d · J, J = 1 − 3/(4·48−9) = 0.98361 g = 0.1699
- 6 Variance & 95% CI (Hedges & Olkin 1985) Var(g) = (N)/(n₁·n₂) + g²/(2N) = 0.0836; CI = g ± 1.96·√Var [−0.397, 0.737]
Independently re-derived: the point and 95% CI both match the reported values.
Independent verification has not run for this value. Provenance located the number in the source; that is not the same as a second reader agreeing with it.
The extraction phase
Screening ends. The grind begins.
Extraction by hand
- Two to three hours per trial: finding the right table, the right arm, the right timepoint
- Usually alone. If you can staff a second extractor, a reconciliation meeting on top of it
- SEM or SD? Median and IQR to convert? Endpoint or change score? Decided at 11pm, cell by cell
- Six weeks later, a grid you still have to spot-check against forty PDFs
With TrialExtract
- Upload your library. The first pass is done in minutes, not weeks
- Every value arrives with the exact sentence it came from, highlighted where it sits in the paper
- SEM to SD, median and IQR to mean and SD, effect sizes: converted with cited formulas, every step shown
- Second-extractor rigor without a second extractor. You review what the machine did, and decide
How it works
You bring the library. It brings back the month.
- 1
Import your studies
RIS, BibTeX, .nbib, CSV, or drop the PDFs. Open-access full texts are fetched for you.
- 2
Preview, free
It reads each paper and shows what's extractable: the tables, figures, and outcomes it found. No credit spent.
- 3
Run the extraction
Priced per paper, quoted up front. The quote is exactly what you're charged.
- 4
Review by outcome
Every value lands in a grid, organized by outcome. Open any one for its receipt, then accept, edit, or reject it.
- 5
Pool and export
DerSimonian–Laird random-effects pooling, then metafor, RevMan, or a workbook, each with the full audit trail attached.
Free on every paper
Bring your ugliest PDF. It reads the whole thing, free.
The scanned table. The SEM where you wanted an SD. The subgroup buried in a figure. Give it the paper you're sure will break it. It reads the whole thing: a plain-language synopsis, and every table, figure, and supplementary block it found, tagged by how it was read. Reading is free. You spend a credit only after you've watched it work.
A 12-week, double-blind, three-arm randomized controlled trial comparing standardized Bacopa monnieri extract (300 mg/day and 600 mg/day) with placebo in 60 healthy older adults. It reports memory, sustained-attention, and reaction-time outcomes with per-arm means and standard deviations, alongside adverse-event counts, primarily in Tables 2–4 and Figure 1.
Readable structure we found
The deliverable, in full
The whole table, with receipts.
Six weeks of typing, done. One row per outcome, per arm, per timepoint: the effect as the paper reported it, the standardized SMD with its confidence interval, and where each value was read. Derived values are marked as derived. A value the paper doesn't report enough to compute is shown as reported, never invented.
| Study | Run | Outcome | Contrast | Timepoint | Family | Native | SMD d [95% CI] | Status | Verify | Decision |
|---|---|---|---|---|---|---|---|---|---|---|
| Stough 2008 | — | Table 1 — Mental control (Wechsler Memory Scale) | SBME vs Placebo | 12 weeks (TREATMENT) | Continuous | Mean difference 0.60 score | 1.14 [0.43, 1.86] | derived | — | — |
| Calabrese 2008 | — | Table 1 — Heart rate (bpm) | Bacopa vs placebo | 6 weeks (TREATMENT) | Continuous | Mean difference −2.60 beats/min | −0.27 [−0.84, 0.29] | derived | — | — |
| Calabrese 2008 | — | Article narrative — Total adverse events (all reported AEs) | Bacopa vs placebo | 12 weeks (TREATMENT) | Continuous | Mean difference −5.00 N | — | reported | — | — |
| Stough 2008 | — | Table 1 — Percentage improvement category: > 21% (0–12 weeks, total score) | SBME vs Placebo | 12 weeks (TREATMENT) | Responder | Responder proportion 0.56 % | — | reported | — | — |
What you walk away with
Pooled, plotted, and ready to defend.
The forest plot you'd drop straight into the manuscript: two randomized bacopa memory trials, pooled into one random-effects estimate. And when the studies genuinely disagree, it shows you, instead of hiding it behind false precision.
Memory (Bacopa monnieri vs placebo)
Defensibility
Built for the day Reviewer 2 asks where a number came from.
Every value here answers for itself: the sentence it was read from, the formula that computed it, and its verification status. When a paper doesn't report a number cleanly, the derivation is shown, not hidden. Open any number and check it yourself, even two years on at revision, long after the details have left your head.
Suspect numbers stay out of your pool
Quality checks run on every extracted value. One that fails them is quarantined: surfaced with the reason, and excluded from every pooled estimate until you rule on it. Your forest plot never quietly includes a value you haven't seen flagged.
Your decisions, on the record
Accepts, edits, and rejections are stamped into the export. When a co-author asks whether you included the per-protocol arm, the answer is in the trail, not in anyone's memory.
Nothing enters silently
No value reaches your evidence table without a source. If a paper doesn't report enough to compute an effect, you see that stated. You'll never discover a guessed number at galley proofs.
Statistical rigor
It recovers the numbers a paper leaves out. No typing, no guessing.
When a paper reports an SEM instead of an SD, or a median where you needed a mean, TrialExtract recovers the real value with the Cochrane Handbook's own formula: every step shown on the receipt, nothing estimated, nothing assumed. Hedges' g applies its small-sample correction as a visible step. Pooling names its estimator on the plot. Math your methods section can cite, and a statistical reviewer can check line by line.
Derivation · recorded on the receipt
SE = (CIupper − CIlower) / (2 · z0.975)
SD = SE / √(1/n₁ + 1/n₂)
Cochrane Handbook §6.5.2.3 · obtaining SDs from CIs
g = J · d, J = 1 − 3/(4·df − 1)
Hedges & Olkin · small-sample correction, shown as a step
Pricing
Credits, not subscriptions.
Reviews are episodic, so your tooling shouldn't bill monthly. At two to three hours a paper, forty trials is over a month of work by hand. A credit takes one of them from PDF to a cited row you can review in minutes, each value carrying its own receipt. Buy a pack when a review starts. Nothing recurs.
Starter
$49
3 papers · $16.33 per paper
Pilot it on your own trials
Outcome matrix save 9%
$149
10 papers
A typical meta-analysis
Review pack save 19%
$399
30 papers
The full systematic review
Previews are free on every paper, so you see exactly what's extractable before a credit ever leaves your balance.
Questions reviewers ask
Before you trust it with a review.
- What study designs does it handle?
- Randomized controlled trials: parallel and crossover, continuous, dichotomous, and time-to-event outcomes, including the multi-arm and multi-dose shapes common in supplement and nutrition trials. RCTs are the one study type TrialExtract is built and tuned for, and that focus is why the receipts hold up.
- Will a journal accept tool-assisted extraction?
- What editors want from any extraction, by hand or tool-assisted, is transparency about who checked what. That's exactly what the audit trail and the methods paragraph give them: every value reviewed by a named person, on the record.
- What happens to my PDFs?
- They're used to run your extraction, stored for your review, and deleted on your schedule. Your library isn't training data.
For your manuscript
Ready for your methods section.
Report tool-assisted extraction the honest way. Paste this and set your reviewer count:
Outcome data were extracted using TrialExtract (Reseda Labs), which records a source quotation, derivation chain, and verification status for each value. All extracted values were reviewed by [N] reviewer(s), and pooled estimates were computed using DerSimonian–Laird random-effects models.
Your evidence table, by this afternoon.
Every number leaves with its receipt, so it holds up whenever it's questioned. Founding seats are limited and go in the order they're claimed, so the earlier you're in, the sooner you're running.
Claim your founding seat before it's taken: