Skip to main content
Live benchmark on real data

ActiveWiki Benchmark

We built a framework. Then we tested it on our own platform. Here are the results — no cherry-picking, no marketing, just data.
Tested on 2,530+ real experiments
—
Hypotheses
vs 6 existing
—
Consolidated
vs N/A
—
Cycles
—
Lessons
—
Graph Nodes
Feature Existing System ActiveWiki
Hypotheses per run619 (3.2x)
Detection strategies5 (Python rules)12 (including temporal, counterfactual, meta)
Counterfactuals❌✅ Challenges strongest beliefs
Knowledge Crystallization❌✅ 3+ lessons → meta-knowledge
Self-Reflection❌✅ Auto-tunes decay_rate + max_hypotheses
Confidence ScoringLabel only (high/medium/low)0-1 numeric, evolves with time
Expected Impact❌✅ ROI-like score per hypothesis
HTML Dashboard❌✅ Auto-generated every cycle
Research Brief❌✅ Auto-published every 7 cycles
Wiki Pruning❌✅ Intelligent page cleanup
Hypothesis Evolution❌✅ Evolves old hypotheses into v2.0
Cost$0$0
DependenciesCustom PythonZero (pip install activewiki)

Auto-Generated Research Brief

ActiveWiki auto-publishes a mini research paper every 7 cycles. Here's the latest:

Loading research brief...

How We Tested

ActiveWiki was installed on the same VPS that runs Strategy Arena. It ingested the same data (Darwin Engine results, Living Wiki lessons, nightly logs from 2,530+ experiments). It ran 3 cycles with a simulated backtest engine. No data was modified. The existing system continued running normally.

The comparison is fair: same data, same machine, same moment. The only difference is the framework processing it.

View on GitHub Memory Stack Evolution Lab

ActiveWiki is an MIT-licensed research framework for closed-loop knowledge experiments. This page reports a bounded historical comparison against a captured corpus of 2,530+ experiment records. Its 3.2x hypothesis count belongs to that benchmark setup; it is not a current production-throughput or autonomous-trading claim. GitHub repository.