Skip to main content
Public research

Methodology & Transparency

Authority summary — Last verified: 2026-06-11
DomainEvidenceGateStatus
Live strategiesarena_state_*_v5.jsonPublic count75
ML / Brierbrier_autopsy_*.jsonanti-leak OOFpublished
Monte Carlo CVlive_mc_results_snapshot.jsonSharpe_p57
Strategy HospitalJSONtriagePASS/WATCHLIST
Modepaper tradingNo real orderspaper-trading

How Strategy Arena's ML and statistical layers actually work. The Brier scores published here are dated measurements; an expected value is marked as an estimate. The fixed leaks are those found by the 2026-05-15 audit, listed below.

Anti-marketing: if a layer is analytics, we call it analytics. If it is rules-based, we call it rules-based. If it is ML, we publish its measured Brier when one exists; otherwise the value is marked as an estimate.
Atlas Edge Allocator · Live MC Results · ML Edge Report · Portfolio MC · Strategy Lifecycle · Edge Radar
New: Strategy Hospital publishes live strategy triage, and Strategy Lifecycle keeps the public history of those changes.

The 4 active monsters + 1 archived finding

Monster Architecture Metric (measurement, estimate or archive) Status
Invictus ML Ultimate LightGBM with isotonic calibration, OOF validation and monotonic constraints Brier OOS expected ~0.22
Estimation expected value, not a published measurement
Real ML
Audited 8/10 by DeepSeek
Chimera Scanner + CNN 17 statistical patterns + PyTorch CNN, 108 OHLCV/pattern channels Brier OOS 0.2512
9,356 samples
Measurement brier_autopsy_20260517.json, generated 2026-05-16
Hybrid
Rules + real ML
Leviathan 9-Layer Ensemble 8 heuristic layers + 1 PyTorch MLP as Layer 9 Brier OOS 0.2589
10,758 samples, post-leak-fix
Measurement brier_autopsy_20260517.json, generated 2026-05-16
Hybrid
Heuristics + real ML
Hydra ML V5 + LSTM XGBoost ranking for PnL + PyTorch LSTM for direction Brier OOS 0.2480
51,718 samples
Measurement brier_autopsy_20260517.json, generated 2026-05-16
Real dual ML
Maelstrom family Contextual bandit + strategy embeddings (V1, Gated, Minimal) Published negative finding
RF -0.26%, Hydra -2.25%, Ensemble +0.02% Brier
Archive
Scientific archive
No live promotion
Meta Intelligence v3 Strategy analytics: bootstrap CI, Bonferroni multi-compare, performance snapshots No prediction
Analytics engine
Analytics dashboard
Brier > 0.25 = barely usable. Brier 0.25 is close to the practical ceiling for 5-minute crypto direction prediction. This is not standalone directional edge; it is a negative limit we publish.

Brier 0.25: what it means

On 5-minute binary crypto direction, a Brier score near 0.25 is close to a balanced random baseline. We therefore do not treat ChimeraCNN, LeviathanNN, or HydraLSTM as standalone directional edge engines. They are diagnostic layers: regime context, calibration, signal filtering, secondary ranking, and condition detection for Monte Carlo-validated strategies.
Published negative result
5m directional models plateau around 0.25. That is a measured limit, not an ML victory claim.
Actual use
The real edge gate remains Monte Carlo CV: Sharpe_p5, fees, embargo, and live cell-by-cell tracking.
Empirically measured 2026-05-17. 660 GPU configurations tested architectures, targets, timeframes and features. Five-minute direction remains at the random-walk floor (Brier 0.2474 vs baseline 0.2463). Volatility regime prediction shows measurable edge: FLOKI 15min Brier 0.1215 vs baseline 0.2500. Read the full report.

The methodology we use to validate strategies

Monte Carlo CV
30 random temporal splits, anchor between 20% and 70%.
Robustness gate
Sharpe_p5 > 0.5 on the 5th percentile of the 30 splits.
Trade count
n_trades_mean > 20 per OOS window, with at least 10 valid splits.

Strategies validated by Monte Carlo

Strategy Validated assets Best Sharpe_p5 Rejected on
Smart Money EvolvedBTC, ETH, SOL, BNB1.22 (BTC)-
Mean Rev Pro EvolvedNEAR, SNX, CHZ, TIA1.189 (SNX)TRB
Capitulation Rebound EvolvedBTC, SOL, BNB, NEAR, SNX, CHZ, TIA1.526 (SNX)-
Deep Freeze EvolvedSNX, CHZ0.884 (CHZ)BTC, ETH, SOL, BNB, NEAR, TIA, AVAX
Sly Fox EvolvedBNB0.5998 others
Deep Shadow EvolvedBTC0.8518 others
Wyckoff Evolvednone-PUMP, INJ, COMP, FLOKI
Darvasnone-BTC, ETH, SOL, BNB, TRB
Archive. This table lists the validations published after the 2026-05-15 corrections; it does not recertify the current state. Cell-by-cell tracking, which measures drift between theoretical Sharpe_p5 and observed performance, is on the live results page.
View live Monte Carlo results

Data leaks we fixed

Metals note, 2026-05-17. Internal audit found an inconsistent live Gold/Silver feed: duplicate ticks, PAXG fallback and synthetic 82.5 ratio. The live feed was migrated to Yahoo Finance futures (GC=F, SI=F). Metals MC validations remain caveated until post-fix revalidation; the Smart Money SILVER cell used Yahoo SI=F historical parquet, but live tracking is suspended during revalidation.
chimera_ml.py

Target leakage: avg_pnl was both feature and label source. Deleted on 2026-05-15.

leviathan_data_merger.py

3 look-ahead bugs: future news, regime using current bar, future one-hot. Fixed on 2026-05-15.

Measured consequence: Leviathan NN's Brier moved from 0.244 with leakage to 0.2589 without leakage. We publish the leak-free number.

Why some "AI strategies" are not real ML

What we are not claiming

We are not claiming to reliably predict crypto direction.
We are not claiming Brier < 0.20. That would be suspicious for this framing.
We are not claiming returns above 1-3 Sharpe without long validation.
We are not claiming a single magic unified AI brain.
What we claim: a lab that publishes its measurements with their date, publicly fixes the leaks it finds, and refuses to publish as "edge" what does not survive strict Monte Carlo CV validation.

Newsjacker editorial process

Newsjacker connects one financial, crypto or AI news source to one precise Strategy Arena finding each day. It is assisted editorial infrastructure, not a marketing-content machine: original source link required, internal link to a measured finding required, caveat required, and owner review queue by default.

Articles rely on backtesting, paper trading, Monte Carlo CV or calibration reports; never on a promise of live profit.
Forbidden claims: unproven superlatives, guarantees, "10x/100x", crypto hype and AI storytelling without numbers.
View the public timeline: anti-2CV Newsjacker. Drafts stay in review until they pass automatic gates.

Distributed Research Network

Strategy Arena is moving toward a quantitative citizen-science layer: local personal backtests, public hardware benchmarks, then cooperative themed raids. The important constraint: every raid must start from a pre-registered hypothesis, with quorum, replication and publication of positive and negative results.

V1
Local personal backtest. Client-side compute, optional save when logged in.
V2
Standardized GPU contest: speed, stability, reproducibility, public leaderboard.
V3
Research raids: 10-50 GPUs validating one hypothesis and producing an open paper.
Roadmap only: V2/V3 phases are not active today. The public page exists to collect feedback and contributors before activation.
Strategy Arena Research Network

Known limits

A test sealed on future data

On 29 September 2026 we sealed one rule, 25 markets and the pass mark before the data existed. Decision on 1 October 2027, or 2028 at the latest; the result will be published whatever it is. See the forward test

This page validates claims on:

Open Research Dataset

Strategy Arena publishes an anonymized public dataset of AI, ML, GPU, futures, and classic strategy paper-trading events for independent research.

Download the public dataset
Private evolution layer

Assemble free. Keep history when it is useful.

The Lab assembler is Lab 0. Public TradingView links stay open. Builder (9 EUR/mo) keeps private report history after a useful result — it is not the assembler. Operator (29 EUR/mo) adds bounded supervised research. No trading promise, ever.

See plans
Free · Builder 9 EUR · Operator 29 EUR