Skip to main content
Recorded Autoresearch Archive

Recorded research. Mutations, receipts and rejections — without an autonomy claim.

This archive records how research engines proposed, tested and rejected strategy mutations. A retained train candidate is still only a hypothesis until an independent proof opens its sealed holdout.
Inspired by Andrej Karpathy
×

Recorded campaigns compare bounded mutations on historical data. The page exposes their run log; it is not a promise of continuous execution or seven agents running now.

Mutate
→
Test
→
Keep/Discard
→
Record
↺
-
Experiments
-
Improvements
10
Documented Engines
-
Best Score /100
ILLUSTRATION — SCRIPTED EXAMPLE LINES, NOT A LIVE FEED
These lines are fixed examples of the log format. They are not current results. Dated records are in the Recorded Run Log below.

Recorded Schedule

When scheduled: mutate → test → keep/discard → record

Hall of Fame

Loading...

Cross-Engine Fitness

Loading...

Recorded Run Log

Loading...

Living Wiki — Lessons

Loading...

The Karpathy Method

1

Mutate

Take the best parameters. Apply small random mutations. Like DNA — most are neutral, rarely one is beneficial.

2

Test

Replay the mutation on the campaign dataset. Historical jobs used a bounded per-experiment budget; the receipt records what actually ran.

3

Keep or Discard

If <code>universal fitness</code> (0-100) improves, keep. Otherwise discard. Only improvements survive.

4

Record

Each scheduled track preserves its hypothesis, result and rejection reason. No convergence or continuous-autonomy claim is inferred.

5

Meta-Harness

The agent that optimizes Darwin itself. Mutation rate, crossover ratio, fitness weights — all auto-tuned. <strong style='color:#ef4444'>Darwin evolves strategies. Meta-Harness evolves Darwin.</strong>

Explore Knowledge Graph Nerve Center

How AI Strategy Evolution Works

Strategy Arena's Evolution Lab is a read-only archive of research campaigns inspired by autoresearch. Recorded engines such as Darwin, Leviathan, Chimera, Invictus, Hydra, Portfolio and PromptForge mutated parameters and replayed them on historical BTC data. Their outputs are hypotheses and run records, not proof of profitability or a promise of continuous execution. Promotion requires a separate immutable Lab proof, Rust CPU authority and a sealed holdout verdict.

📖 How it works
The terminal is an illustration with fixed example lines. Dated records are in the Recorded Run Log. ×