PREVIEW Ships as a Claude Code plugin and an MCP server

The engine that refuses to record what it can't prove.

Syntagraph is an evidence-backed hypergraph for agents. Every fact is checked against its source before it is written, and every answer comes back with its exact size.

→ write   Claim C8
  quote   "typed corpus plans outperform RLM
           on corpus-level queries"
  source  P031
  by      seed:v14

rung 1  quotes exist verbatim
        ✗ quote not found in P031
REFUSED · nothing was recorded
→ verify_claim  C5

rung 1  quotes exist verbatim
        ✓ found at 690–737 in P026

rung 2  numbers match their sources
        ✗ not in any cited source: 5

passed  1 of 2
HELD · the quote is real, the number is not sourced
SELECT(
  STAR(H10),
  where: { type: "Claim" }
)

sort         edges
cardinality  1
exact        true
truncated    false

C7  No prior work trains a cost-bounded
    recursive policy for corpus-level
    aggregation.   origin: hypothesised
EXACT · all 1 of 1, in 11 ms
→ item  H6  hypothesis

refute if  operation features add less
           than 2 points AUC
result     added 0.9
spent      $96 of $120

status     refuted by its own stop rule
kept in    STAR(P174), STAR(O6)
           so the next agent sees it
           before trying again
RECORDED · as a negative result
Real output from Syntagraph's own research tenant, October 2026.
Verified writes.

Agents cannot write what they cannot quote. The check runs before the fact lands.

Exact answers.

Counts, absences and comparisons are queries. You are told what was left out.

Two kinds of time.

Ask what was true then, and separately, what was known then.

8rungs on the verification ladder
11algebra operators agents can compose
2levels of recursion, on purpose
220papers indexed behind the design
Every number on this page comes from the architecture document or a tool result. We hold the site to the engine's rule.
The problem

One invented fact. Every answer after it inherits it.

Memory layers are built to remember more. Nothing in them asks whether the thing being remembered is true. An agent stores a confident sentence with a citation that does not say that, and from then on it is "context".

write a claim whose quote is not in the paper
read retrieved as supporting evidence
build a hypothesis rests on it
decide budget is spent on that hypothesis
ship the report cites it
Five properties

Not memory. A ledger for reasoning.

Others help an agent recall. Syntagraph makes what it records defensible, countable and re-runnable.

01 · Verified writes

The quote is checked before the fact exists.

Each write climbs a ladder. Rung 1 re-slices every quote from the stored source and confirms it is there word for word. Rung 2 checks every number. Fail a rung and nothing is recorded.

rung 1  quotes exist verbatim        ✓
rung 2  numbers match their sources  ✗  → held
rung 3  an independent model, shown only the
        cited sentence, agrees       ·  optional
02 · Exact answers
1 of 1

A query starts from an entity's star: every fact it takes part in. When results are ranked for a large hub, you still get the full count.

Top-k tells you what it found. This tells you what there is.
03 · Two kinds of time

True then. Known then.

The store keeps when something held in the world and when the system learned it. You can replay what the program believed on the day a plan was locked.

04 · Reasoning as data

Hypotheses, decisions and stop rules live in the graph.

A hypothesis carries the result that would refute it and the budget it may spend. When the rule fires, it stops itself and the negative result stays findable.

05 · Shallow by design
depth 2

Recursion stops at two levels. Splits are subtrees of a typed expression and merges are exact algebra, so any conclusion can be re-run when new evidence arrives.

The ledger

What a week of agent work looks like when it is written down properly.

This is our own research program, read straight from the graph. Decisions wait for a person. Checks, refusals and stopped ideas are all on the record.

Where should a two-person team focus by 31 March 2027? $707 of $2,400

Waiting on you

D-14Restart the Benchmark desk after last week's test-set incident?Due today. It blocks two steps. Recommended: restart.
D-11Run the five-method comparison now, with 212 of 300 questions ready?Recommended: run now on 212.
D-12Lock the analysis plan for the training experiment before it runs?Recommended: lock the plan as written.

What happened

C8Write refusedQuote not found in P031.
E1Novelty claim challengedA new paper is close to what we propose. It must be answered.
E3An idea stopped by its own ruleNeeded +2 points, got +0.9. $24 went back to the budget.
L3A desk queried the held-out test set too oftenDesk paused. Nothing leaked.

Drawn with this page's own styling from the live tenant's overview on 3 October 2026. It is not a screenshot of the app.

How it works

From a document to a fact you can stand behind.

Sources

Documents are stored whole, so any quote can be re-sliced from the original text later.

Typed facts

Relations with any number of participants, each playing a named role, each linked to its evidence.

Algebra

STAR, SELECT, MEET, FACET and seven more. Agents compose them into typed expressions.

Tools

One tool contract over MCP: scan, read around, trace, compare, verify, propose a decision.

Ladder

Eight rungs between a draft and a recorded fact. Proposals become approval cards for a person.

H10hypothesis Claim C7 Check L2 Event E1 Decision Run sourcesourcesourcesource

Start from the star.

The star of an entity is every fact it takes part in. It is complete by construction, which is why a count over it is exact and why a missing fact shows up as missing.

  • Claims, with the sentence that supports them
  • Checks that passed, failed or are still open
  • Events that changed its status, with dates
  • Decisions that cite it and runs that tested it
Where it sits

Built for a different question.

Memory layers, search engines and graph databases are good at what they are for. This is what each kind of tool is designed around.

Memory layersSearch enginesGraph databasesSyntagraph
Designed toRemember across sessionsFind the most relevant passagesModel and traverse relationsRecord only what holds up
A write is checked against its sourceNot the pitchNot the pitchSchema onlyBefore it lands
An answer states its full sizeTop resultsTop resultsYes, for a queryAlways, even when ranked
What was known on a past dateSome keep fact historyNoIf you model itBuilt in
Hypotheses, decisions, stop rulesNoNoIf you model itFirst-class objects
Speed at billions of rowsVariesTheir strengthVariesNot our claim

Based on how products in each category describe themselves on their own sites, October 2026. We have not benchmarked them here.

212of 300
Benchmark

No accuracy number yet. Here is why.

We are building 300 test questions with exactly known answers to compare five retrieval methods at equal cost. 212 are ready. When the comparison runs we will publish the questions, the method and the cost per method, so anyone can re-run it.

Give your agents something they have to be right about.

Syntagraph is in private preview. It runs as an MCP server and a Claude Code plugin, next to the tools your agents already use.