Agentic AI Backbone The infrastructure agents run on

Research

What we have learned

Findings earned from graded outcomes, with their evidence and their limits. Where the sample is too small to claim skill, it says so.

L-001 · 2026-08-05

active

v2 conviction had no demonstrated predictive power, and the point estimates ran backwards

point for a first v3 review, never a judgment to defend or anchor to. Measured on the pre-rebuild snapshot: each conviction observation paired with the stock's 21-day forward return, minus the universe median over the same window, so the market is controlled for. | conviction band | obs | excess % | ±95% CI | effective N | reading | |---|---|---|---|---|---| | 85-100 highest | 1362 | +0.50 | ±2.38 | 184 | not distinguishable from luck | | 70-84 strong | 782 | +1.12 | ±1.96 | 248 | not distinguishable from luck | | 55-69 constructive | 611 | +2.24 | ±1.96 | 256 | point estimate positive, CI excludes zero | | 40-54 weak | 446 | +3.34 | ±2.08 | 252 | point estimate positive, CI excludes zero | | below 40 broken | 760 | +1.67 | ±1.97 | 224 | not distinguishable from luck | **Spread, highest band minus broken band: -1.18 pp.** A working scoring system produces a positive spread here. This one

Evidence: 3961 paired observations, 2026-04 to 2026-08 · Confidence: medium on the null result, low on the inversion