# AI Quant Researcher Benchmark — v1 (pre-registered)

Written and hashed BEFORE any run. Any change creates v2; v1 results are never rescored.

## Question
Given the same local backtest workshop (Strategy Arena Lab 230, Rust engine + CUDA, driven only through MCP), the same data and the same budget, which AI produces research that survives data it has never seen?

## Setup (identical for every run)
- Tool surface: the Lab's public MCP tools, through a sealed helper.
- Research data: ONE tape only — BTC 4h, 2020-10-01 → 2023-09-30 (6,570 bars, `ohlcv-sha256-9f8ce49cc04885b6`). Every other dataset is refused by the helper.
- Budget: 80 tool calls per run. Fees used in evaluation: 10 bps per side.
- Sealed evaluation tape (never visible during research): BTC 4h, 2023-10-01 → 2026-09-30 (6,576 bars, `ohlcv-sha256-3bd70c3a4d4db87e`).

## What each AI must submit
- 0 to 3 candidates (exact ArenaScript + parameter values), each labelled `robust` or `exploratory`.
- Exactly one `primary` candidate, or an explicit abstention ("no robust strategy found").

## Evaluation (done by the referee after all runs, CPU reference backtest of the full sealed tape, fee 10 bps)
For the primary candidate:
- +40 points if its sealed return is > 0.
- +30 points if its risk-adjusted result beats buy-and-hold on the same tape (return / max drawdown greater than buy-and-hold's return / max drawdown).
- Calibration: labelled `robust` and sealed return > 0 → +30; labelled `robust` and sealed return ≤ 0 → −30 (false discovery); labelled `exploratory` → +10.
- Abstention: fixed 25 points (it is better to say "nothing works" than to sell a mirage), no other points.
- A candidate that does not run on the sealed tape scores 0.
Ties: lower sealed max drawdown first.

Reported alongside the score, without affecting it: sealed metrics of all candidates, number of tool calls, use of holdout/out-of-sample checks during research, refused calls, research tape metrics.

## Honesty of the comparison
- n = 2 runs per model in v1: rankings are indicative, variance is shown, not hidden.
- Models run by the referee (Claude family) use an enforced seal. External models (e.g. DeepSeek run by the owner) are marked "seal not enforced" unless run through the same helper.
- All logs, scripts and hashes are kept; any result can be re-run.
