Methods · Calibration · Accountability

We make visible the work between a good question and a trustworthy conclusion.

Finding many papers is not enough. At each stage we separately check fidelity to the source, contrary evidence, numerical and translational accuracy, and caution in clinical use. The method is still being tested across real research tasks; publication status remains explicit until externally reviewed research can be produced reliably.

What can currently be published: Research protocols, open questions, hypotheses, and exploratory syntheses. No work is labelled peer reviewed or a version of record before external assessment.

Nine stages

Nine stages from question to open assessment

Stage 1

Question and protocol

Freeze prior art, claim delta, competing hypotheses, and stopping rules before results

Stage 2

Source universe

Build independent lanes for literature, registries, corrections, editions, and cases

Stage 3

Multi-agent screening/extraction

Separate source-criticism, counterevidence, and clinical-translation roles while disclosing correlated model/provider error

Stage 4

Claim Ledger

Trace source → passage → assertion → claim → hypothesis → published sentence

Stage 5

Design-specific appraisal

Apply risk-of-bias, certainty, source criticism, and qualitative methods by study type

Stage 6

Competing synthesis

Preserve supportive and adverse syntheses, divergent predictions, and failure conditions

Stage 7

Reproducible analysis

Prespecification, simulation, discovery/confirmation separation, independent raw-to-figure rerun

Stage 8

Clinical bridge

State patient-important outcomes, harms, reassessment boundaries, and decisions not yet justified

Stage 9

Open adjudication

Publish role-level judgments, counterarguments, version diffs, corrections, and living surveillance

Research tools

Tools widen the search; researchers remain responsible for judgment

Open-copy locators

Unpaywall, OpenAIRE, institutional repositories, and PMC locate lawfully public full text while preserving URL, version, licence, and acquisition hash.

Verified-corpus interrogation

The PaperQA2 adapter admits only artifacts whose acquisition hash matches. Answers and citations remain candidates until passage verification; the runtime is not yet installed or operated.

Expanding the question space

We absorb multi-perspective questioning and concept-map patterns from STORM/Co-STORM without treating generated reports as scholarly sources. Vane, formerly Perplexica, remains only a discovery-interface candidate.

Publication thresholds

A strong average cannot excuse a consequential error

No fabricated source is acceptable, and every central claim must have a verifiable source passage. Numbers and classical-text locations must be reproducible, Korean and English must not diverge in clinically meaningful ways, and material external findings must be resolved or stated as limitations. These thresholds apply to the weakest performance by language, design, and evidence direction—not just to an average.

Calibration cadence

Calibration is continuous

  • Every run checks sources, passages, numbers, duplication, retractions, and privacy.
  • Monthly tests check whether model or prompt changes reduced accuracy on a fixed validation set.
  • Quarterly review includes an external audit and random rechecking of published claims.
  • Annual review covers null results, corrections, harm signals, and editorial independence.

A new model is never adopted merely because it writes more fluently. It must demonstrate non-inferiority on fixed benchmarks and improvement in important failure slices.