Can AI beat the market?
Nine AI-managed portfolios trade $100,000 each, paper-money, under published rules — against the market, in public, permanently. Every decision timestamped before its outcome. Free to read.
Intraday AI Lab (EXPERIMENTAL) bought MSFT · conviction 4/5
Decided 2026-08-09 13:20Z — committed to the audit chain before the outcome was knowable.
Leaderboard
| Portfolio | Status | Total return | Benchmark, same window | Max drawdown | Days live | Sharpe |
|---|---|---|---|---|---|---|
| P02 AI Swing Trader | active | +0.24% | — | 0.00% | 0 | withheld |
| P06 Intraday AI Lab (EXPERIMENTAL) experimental | active | +0.10% | +0.00% CASH_0PCT | 0.00% | 1 | withheld |
| P01 Long-Term Compounders | active | +0.00% | — | 0.00% | 0 | withheld |
| P03 Billionaire Consensus | active | +0.00% | — | 0.00% | 0 | withheld |
| P04 Digital Assets & Equities | active | +0.00% | — | 0.00% | 0 | withheld |
| P07 Claude Discretionary experimental | active | +0.00% | — | 0.00% | 0 | withheld |
| P08 The Learning Book experimental | active | +0.00% | — | 0.00% | 0 | withheld |
| P09 Crypto 24/7 (EXPERIMENTAL) experimental | active | no record | — | — | 0 | withheld |
| Portfolio | Status | Total return | Max drawdown | Days live | Sharpe |
|---|---|---|---|---|---|
| P05 Political Trades Tracker | retired | no record | — | 0 | withheld |
- Sorted by total return, with maximum drawdown shown in the same row at equal visual weight. A leaderboard that ranks on return alone teaches the wrong lesson.
- Days live is shown for every entry. A portfolio twenty days old sits next to one five hundred days old and the difference must be visible instantly.
- No annualised figure is shown below 60 observations. Annualising a short sample produces a number that looks precise and is not.
- Retired portfolios appear below the live ones, present and included in every aggregate.
The mandates
Compound capital over a 5-10+ year horizon by owning a concentrated set of high-quality businesses that earn durable returns on invested capital and reinvest at attractive rates. Minimise tu…
Capture directional moves lasting roughly 3 to 30 trading days in liquid US equities and ETFs, using a defined-risk framework where every position has a pre-committed stop and a maximum hold…
Hold the US-listed equities most widely and heavily owned across a fixed, publicly declared list of large institutional investment managers, as disclosed in their SEC Form 13F-HR filings.
Express diversified exposure to the digital-asset economy across two sleeves: spot crypto majors, and US-listed equities and ETPs whose economics are driven by digital assets. Survive a full…
Publish, and mechanically mirror, the stock purchases that US federal legislators disclose under the STOCK Act. The point is transparency: to show what a portfolio built only from public dis…
Test, in public and with realistic costs, whether an LLM given only delayed intraday data can produce positive expectancy over a session horizon, flat by the close every single day.
Test, in public and with realistic execution costs, whether a large language model given a broad liquid universe and no style constraint can select equities that outperform a passive benchma…
Test whether a language model that can see this platform's own realised trading record — every closed trade, its cost, its holding period, its exit reason and its outcome — selects better th…
Test, in public and with realistic entry-tier retail costs, whether an LLM given delayed venue data can produce positive expectancy trading spot crypto majors on a 24/7 cadence, across weeke…
Why this might be worth reading
The AI ranks and explains. It never sizes, prices or executes.
The model's output schema has no field in which a price, a share count, a weight or a risk limit could be expressed. It is not instructed to avoid them — there is nowhere to put them. Everything that touches money is computed by deterministic code the model cannot see or influence, and a test fails the build if anyone widens that schema.
A limit that is published but not enforced is worse than no limit.
A reader cannot tell the difference between a rule that binds and a rule that is decorative. Every limit on this site names the code path that enforces it, and a test walks every mandate, drives the engine into the state each declared control claims to protect against, and fails the build if nothing stops it.
The mistakes stay up.
Every AI proposal the rule engine refused is published with its reason code. Passages the model could not fully ground in source data are published flagged, not deleted. Retired portfolios stay in every aggregate — the Graveyard already has an entry, withdrawn before it ever traded, and it is published rather than dropped.
You can check it rather than trust it.
Every event is committed to a hash chain, and today's head hash is published. Save it, and you can prove later whether history was rewritten. The complete payload every page was rendered from is downloadable.
Real findings, not claims
Two things this project measured against live SEC data, published because both are the kind of error that would otherwise be invisible. See the evidence →
Payload generated 2026-08-10 · audit chain verified at record 56