web3ctx · evidence

Measured,
not promised.

Every number on this page is regenerated from data at build time — nothing is typed by hand. Every measurement carries its protocol, its composition, and its limitations. The experiments that failed are published beside the ones that shipped.

Immutable citations

Can the answer be re-checked after it's given?

Share of citation URLs in each tool's audited payloads that are commit-pinned and re-fetchable. Counted over cached bytes — a property of the URL string. Frozen audit, 2026-08-13.

Figure 1 · immutable / total citation URLs · higher is better
web3ctx
142/144
context7
0/57
exa (web search)
7/213
firecrawl
25/338
ethskills (skill)
0/181

firecrawl publishes only with its sentence: incomplete — our harness overwrote 7 valid payloads; Firecrawl answered every question on the first fetch.

claude web search is MISSING, not passed. No key in that environment. A row silently absent reads as one that passed, so it stays in the table with no number.

exa is labelled “web search” — the only tool its public endpoint exposes. Its code-retrieval product is unmeasured.

Our own two non-immutable URLs, named: https://github.com/sponsors/wevm (a sponsor link inside wagmi's own docs, which we quoted) and https://github.com/wevm/viem/issues/2484 (the name of a test case in viem's source).

Counted over cached bytes, on 16 questions, and it is one count, not a verdict. The audit issues none and none is computable from it.

Head-to-head · 2026-08-20

Six real integration traps, eight arms, raw bytes on disk.

Publication-grade harness: raw HTTP response bytes written to disk before parsing, one tokenizer over full payloads on every arm, refusals void. Six verified integration footguns plus ten plain questions.

Figure 2 · all arms · ⊘ no verdict — n=16 against the registered 393 floor · descriptive only
armranvoidmed tokensimmutable / total URLsnames vertraps
web3ctxOURS1603,071316 / 32410 / 161 / 6 (4p)
web3ctx (intent=integrate)OURS1607,454424 / 51616 / 166 / 6
context71605580 / 06 / 160 / 6 (2p)
exa (code context)1601,0331 / 17111 / 162 / 6 (3p)
exa (web search)1608,98611 / 26313 / 165 / 6 (1p)
claude web search16017,1636 / 40411 / 166 / 6
firecrawl (developer search)1153,78126 / 2339 / 114 / 6 (2p)
ethskills (static bundle)185,0680 / 2751 / 1

The prediction we registered was wrong, and it goes first. We wrote “we expect to be the most expensive arm.” Measured with one tokenizer over full payloads, our default row is 3,071 median tokens — under three of the arms here. The earlier adversarial run’s “most expensive of four” priced a different unit and read our content-only figure against everyone else’s full payload. It licenses no cost claim in either direction.

The two web3ctx rows are ONE finding and are never shown apart. Called the way a client that does not know our argument names calls us, we catch 1 of 6 traps. With intent: "integrate" — the value that reaches the human-validated recipe — 6 of 6. The gap is the argument, and the weaker row is ours to publish.

context7’s 0/0 is its own sentence: its payloads contain no URLs at all. That is a different fact from citations that fail to pin, and collapsing the two would misreport it.

T6 carries a permanent disclosure. Our recipe gained the ERC-4626 rounding matrix on 2026-08-20 after losing this trap in the earlier adversarial run. A benchmark that quietly re-runs a question after fixing its answer is measuring its own patch — so this sentence travels with the row for as long as the row exists.

VOID — 5 row(s), listed: firecrawl-dev/Q06, firecrawl-dev/Q07, firecrawl-dev/Q08, firecrawl-dev/Q09, firecrawl-dev/Q10. All 5 are a real provider quota reached during the run. Excluded from numerator and denominator; that arm publishes as 11 ran / 5 void, never as a fraction of 16 and never as a zero.

claude web search caught 6 of 6 — matching our best row and beating our default one — at 17,163 median tokens against our 3,071, with 6/404 immutable citations against our 316/324. Both compositions belong to that comparison; either number alone tells a different story than the pair does.

NO VERDICT. n=16 against the registered floor of 393 (D17). Bootstrap CIs are in the result document; a confidence interval does not repair a small n, it states it.

Figure 2b · median tokens per answer · log scale · cost licenses no claim
web3ctx
3,071
web3ctx (intent=integrate)
7,454
context7
558
exa (code context)
1,033
exa (web search)
8,986
claude web search
17,163
firecrawl (developer search)
3,781
ethskills (static bundle)
85,068
Read with Figure 2, never alone. Cost and correctness are one trade, shown together.
Figure 2c · trap-by-arm scorecard · mint = caught · half = partial · hollow = missed
trapweb3ctx*web3ctx*context7exaexaclaude web searchfirecrawl
T1 deadline is NOT in the params struct on SwapRouter02
T2 amountIn: 0 spends the contract's WHOLE balance
T3 outputAmount IS the fee — the difference is what the relayer keeps
T4 wagmi v3 uses useConnection, not useAccount
T5 RainbowKit 2.x declares wagmi ^2 as its peer range
T6 BOTH halves — down when issuing shares/paying assets, UP when charging
Earlier adversarial run, superseded. The same-day stranger-session benchmark (agent-transcribed payloads, asymmetric token counts) is retained in the archive under its own banner — “directionally sound, not publication-grade,” its own words. The harness above is the repair it asked for; its pre-stated prediction is quoted in the archive, including the clause it got wrong.
Receipts

Verify a recipe yourself.

Every recipe is validated by a person against live chains and served with its transaction hashes. These are the flagship's — check them with any RPC, right now.

Figure 3 · flagship recipe receipt chain · re-verified by eth_getTransactionReceipt
stepchaintransactionwhat it proves
approvebase-sepolia0xd4b4e6395f…a3d0f5the router was allowed to move exactly this much USDC
approvebase-sepolia0xb0db8b5854…2421e6the router was allowed to move exactly this much USDC
burn (depositForBurn)base-sepolia0xdba9da61c7…1a0f74the USDC left the source chain
mint (receiveMessage)arbitrum-sepolia0xcd869bb6b1…bd2dabit arrived — the destination leg, on a different chain
attestation probebase-sepolia0x1460054173…6c3080the attestation was fetched, not assumed
What the receipts do not prove travels with them in the payload — a claim never exceeds its evidence. Full hashes, blocks and event decodes are served by the product and listed in the evidence repo.
Scope precision

Does the right project answer?

Share of delivered units whose project@version matches the query's gold scope, over the locked 200-question set at the current corpus.

Figure 4 · scope-precision@10, macro · wildcard-version · query-level bootstrap CIs
armscope-precision@10nabstained
dictionary-word floor OFF81.0% [95% CI 74.3 – 87.2]12720 / 147
floor ON — as shipped80.4% [95% CI 73.5 – 86.7]12720 / 147

The abstention count rides with the number, always — 20 / 147 in both arms. Precision computed only over answered questions would rise every time the server declined to answer, which would reward silence.

Not comparable to a published code-retrieval precision. CodeGrep and its kin retrieve within one repository: project selection never enters their metric, so under our definition their BM25 would read 1.000 — including the configuration measured at 0.375. The two numbers measure different stages, and ours was pre-registered as an upper bound where a pass is “suggestive, never proof”.

Composition: measured live against the deployed surface, macro-averaged, wildcard-version, both arms on the same questions. The two rows differ by one shipped mechanism and by nothing else — the identical abstention counts are the evidence for that.

Dead experiments

What we measured and rejected.

Every candidate below was pre-registered before it ran, failed its gate, and was not shipped. A candidate that fails its gate publishes anyway — that is the point of the gate.

candidateits pre-registered gatemeasuredwhy it died
Doc-side sparse expansionprose recall +2.00pp+0.00pp against a +2.00pp barFailed on inertness, not intrusion. The model ran and moved nothing, at cost. The intrusion clause passed — descriptive only.
Unit-class demotionpayload quality on affected queries7 of 7 evicted, on 2.6% of queriesThe mechanism worked exactly as specified and could never reach its own trigger — the anchor it was written for holds zero inert units.
Section-select for recipe bodiesno anchor regressesflipped both target anchors — and lost a thirdIt improved the two anchors it was aimed at and silently broke the one it was not watching. A gate registered before a run is not editable after it.
Inline bodies in search resultscompletion rate and tokens per completed taskcompletion 8/12 → 7/12Completion fell and tokens per completed task rose. It improved the linear term and multiplied the quadratic one.
Content-hash dedupe≥15% cumulative token saving4.7% cumulativeEvery structural clause held and the saving was a third of its bar — and it could not have fired on the deployed transport at any bar.
Corpus & receipts
Corpus
104 integration projects + 1,193 specification ids (609 ERC + 583 core EIP + 1 in both repos) = 1,297 total
the pair, always — a bare total does not publish
Units
585,017
each labelled project@version with a pinned, re-fetchable source
Recipes
24
validated on-chain by a person, served with checkable receipts