Question and protocol
Freeze prior art, claim delta, competing hypotheses, and stopping rules before results
Methods · Calibration · Accountability
Finding many papers is not enough. At each stage we separately check fidelity to the source, contrary evidence, numerical and translational accuracy, and caution in clinical use. The method is still being tested across real research tasks; publication status remains explicit until externally reviewed research can be produced reliably.
Nine stages
Freeze prior art, claim delta, competing hypotheses, and stopping rules before results
Build independent lanes for literature, registries, corrections, editions, and cases
Separate source-criticism, counterevidence, and clinical-translation roles while disclosing correlated model/provider error
Trace source → passage → assertion → claim → hypothesis → published sentence
Apply risk-of-bias, certainty, source criticism, and qualitative methods by study type
Preserve supportive and adverse syntheses, divergent predictions, and failure conditions
Prespecification, simulation, discovery/confirmation separation, independent raw-to-figure rerun
State patient-important outcomes, harms, reassessment boundaries, and decisions not yet justified
Publish role-level judgments, counterarguments, version diffs, corrections, and living surveillance
Research tools
Unpaywall, OpenAIRE, institutional repositories, and PMC locate lawfully public full text while preserving URL, version, licence, and acquisition hash.
The PaperQA2 adapter admits only artifacts whose acquisition hash matches. Answers and citations remain candidates until passage verification; the runtime is not yet installed or operated.
We absorb multi-perspective questioning and concept-map patterns from STORM/Co-STORM without treating generated reports as scholarly sources. Vane, formerly Perplexica, remains only a discovery-interface candidate.
Publication thresholds
No fabricated source is acceptable, and every central claim must have a verifiable source passage. Numbers and classical-text locations must be reproducible, Korean and English must not diverge in clinically meaningful ways, and material external findings must be resolved or stated as limitations. These thresholds apply to the weakest performance by language, design, and evidence direction—not just to an average.
Calibration cadence
A new model is never adopted merely because it writes more fluently. It must demonstrate non-inferiority on fixed benchmarks and improvement in important failure slices.