01 / Claim
A paper says it works.
Now show the evidence.
Paper2Proof selects one bounded numerical claim, proposes a reproducible experiment, pauses when scientific judgment matters, and returns a content-addressed claim-to-evidence trail.
- 6
- specialized roles
- 12/12
- contract fixtures passed
- 0
- unapproved external writes
Evidence workspace
Choose a controlled demo or inspect a public paper.
No sign-up. No paper upload. Public sources only.
Ready for a source
One claim. One bounded run.
Every step accounted for.
Select a demo on the left to watch the agent workflow assemble real evidence.
Building evidence
Human decision required
02 / Bounded plan
Waiting for a policy-safe plan…
No code generated.
03 / Observed evidence
Agent and tool trace0 events
Designed for agency, not chat
The agents do the busywork.
Humans keep the judgment.
The workflow continues in the background, but crossing a scientific or safety boundary always creates a concrete decision with visible consequences.
- 01Intake & Risk
Normalizes public sources and treats every document as untrusted data.
- 02Claim Extraction
Links one numerical claim to exact source content and a declared tolerance.
- 03Reproduction Planner
Proposes one CPU-sized, networkless experiment under hard policy limits.
- 04Execution
Runs approved code only inside a declared isolated backend—not on the host.
- 05Verification & Critic
Audits limitations while deterministic code owns the numerical verdict.
- 06Report & Action
Packages provenance, timestamps, hashes, trace, and the next human decision.
“A successful run is evidence—not proof that a paper is true.”
Model ≠ authority. Deterministic policy owns URLs, code, budgets, state transitions, and verdict math.
Provenance ≠ decoration. Local, recorded, and AgentCore results are impossible to confuse in the schema.
Failure ≠ silence. Blocked and failed paths produce structured reasons instead of disappearing.