Agents for HumansResearch integrity

A paper says it works.
Now show the evidence.

Paper2Proof selects one bounded numerical claim, proposes a reproducible experiment, pauses when scientific judgment matters, and returns a content-addressed claim-to-evidence trail.

See the agent workflow
6
specialized roles
12/12
contract fixtures passed
0
unapproved external writes

Evidence workspace

Choose a controlled demo or inspect a public paper.

No sign-up. No paper upload. Public sources only.

Ready for a source

One claim. One bounded run.
Every step accounted for.

Select a demo on the left to watch the agent workflow assemble real evidence.

Designed for agency, not chat

The agents do the busywork.
Humans keep the judgment.

The workflow continues in the background, but crossing a scientific or safety boundary always creates a concrete decision with visible consequences.

  1. 01
    Intake & Risk

    Normalizes public sources and treats every document as untrusted data.

  2. 02
    Claim Extraction

    Links one numerical claim to exact source content and a declared tolerance.

  3. 03
    Reproduction Planner

    Proposes one CPU-sized, networkless experiment under hard policy limits.

  4. 04
    Execution

    Runs approved code only inside a declared isolated backend—not on the host.

  5. 05
    Verification & Critic

    Audits limitations while deterministic code owns the numerical verdict.

  6. 06
    Report & Action

    Packages provenance, timestamps, hashes, trace, and the next human decision.

“A successful run is evidence—not proof that a paper is true.”

Model ≠ authority. Deterministic policy owns URLs, code, budgets, state transitions, and verdict math.

Provenance ≠ decoration. Local, recorded, and AgentCore results are impossible to confuse in the schema.

Failure ≠ silence. Blocked and failed paths produce structured reasons instead of disappearing.