What WORKS today (proven, live)
Comprehension- Any OpenAPI or docs surface → question-shaped tools; auth invisible to the agent, injected at call time. 14+ real specs validated; 0 expose an auth header.
- A draft spec recovered from a human docs page, with claims marked
VERIFIED/REFUTEDagainst reality rather than trusted. - PDA seed recovery from Anchor IDLs and from program source (proven on a no-IDL/Steel program) — including recipes the IDL structurally drops or hides in accounts that travel only as remaining accounts.
- Auto-comprehend-on-pick: point at a project → a generated program config, differential-proven equal to hand-authored ground truth on 4 programs, plus an explicit measured overlay of what could not be derived from any public surface.
- Provenance on every edge, one canonical vocabulary. Surface graph:
EXTRACTED>DECLARED>INFERRED>CLAIMED→VERIFIED/REFUTED. Program graph:EXTRACTED/RECOVERED/FLAGGED. - Real recovered facts a coding agent cannot get from the surface: Meteora’s
base_factor4th seed (the deprecated 3-seed scheme silently derives the wrong pool), Pump.fun’sbonding_curve_v2(required, invisible in the IDL), a fee-recipient field resolved empirically by a refuting Receipt, the liquidity-bitmap bin-array walk (the naive heuristic fabricates dead accounts). - Cross-API correlation on declared value-domain joins, proven across three real APIs.
- Hosted and local MCP; scale-adaptive tool listing (full defs withheld above scale, recovered per tool on demand).
- Measured context cuts of −77% / −89% on two real specs, with first-call-correctness held. Bytes measured; tokens estimated.
find_start— intent → the right (program, instruction) start point, ranked, with a dependency-ordered derive plan, provenance on every account, declared preludes, and an honest no-start below the retrieval floor. See find_start.- The Scorecard (
gecko report) and the Playground.
-
The simulate→Receipt engine, live-proven twice on a surfpool mainnet fork (a
mainnet-backed snapshot — not mainnet), simulation only, $0, nothing signed or
broadcast:
CU numbers are measured per run and vary slightly with on-chain state; the stable claim is the side-by-side verdict. See The Receipt.
- A categorical corpus (
observed/reported/synthetic/simulatedtiers) plus an N-confirmed drift detector. Values-free by construction, and audited. - The
simulatedtier is wired end to end: the landing orchestrators and thesimulatetool take an explicitrecord_toopt-in (default: record nothing), andgecko driftreads the series back — categorical rows only, never a pubkey, amount, or log.
- Seven fail-closed layers: spec sanitizer · per-tool quarantine · image Skill Guard (rendered-pixel payloads, encoded-content rescan) · SSRF netguard · out-of-band auth-host anchoring · verdict signing gate · an AST-enforced boundary that proves the landing layer contains no sign or send path.
NOT built yet (honest)
Nothing below is claimed anywhere else on this site.- The drift scheduler — re-simulation on a cadence. Today the series accrues only when runs happen.
- Hosted point-and-simulate, and hosted program surfaces beyond the first.
- Signing-gate binding to the exact simulated message hash (
evaluate_tx), and the TEE credential backend. Today’s gate is verdict-based; the diagrams label it that way. - Catalog breadth — 4,500 projects listed, 5 wired deep. Non-Anchor generalization is proven once; wider coverage is open.
- Pump.fun sell round-trip, ORE claim, MetaDAO fund executable intents.
- The semantic/vector retrieval tier — deliberately OFF behind an evidence gate. It flips only on measured lexical recall failure, not fashion. One measured negative result on embeddings is on record.
- Cross-customer episodic pooling — tenancy is local-only until a consent/egress layer exists.
- Live x402 billing — stub by design.