01Autonomous Market Intelligence
Forward-collected paper strategy · not real-money live
Daily Russell 1000 predictions have been generated out of sample since April 2025 with autonomous web research. The reported top-20 long portfolio produced positive factor-adjusted returns, but the signal was concentrated in the very top names and remains a research portfolio rather than deployed capital.
Inspect the project ↗02KTD-Fin
Leakage-controlled historical benchmark · not live
A masked CSI 300 evaluation with Barra attribution found that apparent gains were largely market and style exposure, with limited persistent stock-selection alpha. It is a useful warning that fluent rationales do not establish security-selection skill.
Inspect the project ↗03NextFund
Live-market evaluation infrastructure · no verified returns
Time-consistent market access and persistent logs connect each observation, justification and trade across U.S., Hong Kong and China A-share markets. That improves auditability, but the paper reports a platform demonstration—not evidence of excess returns.
Inspect the project ↗04DX production fleets
Forward live · real-money + live-price paper fleets
A six-month record covers 3,505 real-ETH agents and 500–599 mostly paper perpetuals agents. Neither fleet showed a directional edge: only 16.2% of the real-money vaults finished profitable, while the paper fleet lost money and trailed matched retail. Controls and order-path mechanics mattered more than prompt prose.
Inspect the project ↗05FinSkillBench
Point-in-time task benchmark · not live trading
Across 9 models and 2,603 investment-management episodes, curated procedural skills lifted mean scores from 0.366 to 0.528; agents writing their own skills gained little. Reliable process scaffolding mattered more than improvised prompt complexity.
Inspect the project ↗06LLM Trading Lab
Forward-only, public microcap experiment
Excellent transparency: daily logs, trades, benchmarks and risk rules. Early gains are interesting, but one tiny portfolio is not statistical proof.
Inspect the project ↗071rok
Forward test · public paper accounts
Seven models receive the same tools, prompts and $100K paper accounts on a weekly clock. The open harness is useful because research and order execution are separate; its leaderboard is evidence about one live paper protocol, not real-money alpha.
Inspect the project ↗08InvestLogicBench2026
Forward test · simulated capital
Across seven weeks, most agents trailed SPY; the best general reasoning scores did not predict returns. The durable lesson is to score event selection and thesis continuity, not model prestige.
Inspect the project ↗09AQuA
Causal walk-forward backtest · not live
Self-improving research agents carried empirical beliefs across runs inside sealed evaluation sandboxes. Promising process evidence, but still historical and not a forward portfolio result.
Inspect the project ↗010CLQT benchmark
Backtest + four-week live paper track
The best Sharpe ratio was not the most capable agent. Models often traded in the direction of their analysis but sized positions inconsistently with their own reasoning.
Inspect the project ↗011FIDES
Out-of-sample backtest protocol · not live
FIDES checks whether an agent's prose, code and realized track record describe the same strategy. Only 2 of 40 generated strategies beat buy-and-hold, and claimed outperformance was badly calibrated.
Inspect the project ↗012builderr Trading Agent
Forward test · shared paper sandbox
Its July 7–September 4 round used common fills, fixed risk limits and a public agent brief. The repository describes the protocol but does not yet publish a final leaderboard, so no performance claim is credited.
Inspect the project ↗013Vibe-Trading
Open-source prompts and tooling · no verified returns
A useful library of research, backtest and bull/bear committee prompts with explicit paper/live connector controls. It is implementation evidence—not evidence that the generated strategies outperform.
Inspect the project ↗