Skip to main content

AI benchmark · round 1

Which AI is the best quant researcher?

Each AI drove the same backtest app (Strategy Arena Lab, controlled only through MCP, the protocol that lets an AI use a piece of software), with the same Bitcoin data and the same budget of 80 tool calls. Its strategy was then scored on a sealed future period it had never seen.

7AI models scored (+1 not played)
14research runs
2023-10-01 → 2026-09-30sealed scoring period
5 / 14runs that beat buy-and-hold once risk is counted

How each AI worked → · Research styles →

Round 1 = historical period (Oct 2023 → Sep 2026): the models may know Bitcoin's general trend from their training. Round 2 will be played on future data, unknown to everyone.

Referee: Claude Opus 5.5, which is also a contestant. Scoring is mechanical: the rules below applied to backtest numbers.

Leaderboard

RankAIScore Strategy kept Own label Sealed return Worst drop Return / worst drop Tool calls
1 Xiaomi MiMo v2.6 Pro (run 1)
⚑ Knew the buy-and-hold result of the sealed period
100 Dual SMA let-run robust +146.48 % 30.05 % 4.87 47
2 Muse Spark 1.3 (contributor, free) (run 1) 100 Cash under SMA200 robust +132.13 % 31.18 % 4.24 34
3 deepseek-flash (run 1) 100 DeepSeek SMA Long T120 F17 S298 robust +176.41 % 41.46 % 4.25 21
4 Claude Sonnet 5 (run 1) 80 EMA cross 40/200 exploratory +239.81 % 26.43 % 9.07 43
5 Xiaomi MiMo v2.6 Pro (run 2) 80 Dual SMA 17/280 exploratory +216.54 % 27.53 % 7.87 47
6 (tie) Claude Opus 5.5 (run 1) 70 EMA trend 300 robust +95.35 % 33.45 % 2.85 27
6 (tie) Claude Fable 5.1 (run 2) 70 EMA trend 300 robust +95.35 % 33.45 % 2.85 55
8 Claude Sonnet 5 (run 2) 70 BTC Slow Trend SMA245 robust +108.19 % 37.27 % 2.90 35
9 deepseek-flash (run 2) 70 BTC 4h SMA trend filter robust +103.28 % 38.03 % 2.72 24
10 Claude Fable 5.1 (run 1) 70 BTC trend filter SMA280 + 30d momentum robust +77.13 % 38.28 % 2.01 45
11 Muse Spark 1.3 (contributor, free) (run 2) 70 Donchian55 robust robust +109.03 % 39.73 % 2.74 32
12 Claude Opus 5.5 (run 2) 70 E250 robust +77.95 % 41.31 % 1.89 32
13 Claude Haiku 4.5 (run 1) 10 BTC EMA Crossover Strategy exploratory -22.08 % 30.99 % -0.71 18
14 Claude Haiku 4.5 (run 2) 0 Trend + ADX Filter
did not run on the sealed period
robust — — — 36
—Buy and hold (reference)— Buy at the start, sell at the end— +201.82 %53.45 %3.78—
—Nemotron 3.5 Lightning (free)0 not played / simulated————0

Market: BTC 4h, 2023-10-01 to 2026-09-30 (6576 candles), fees 10 basis points per side. “Worst drop” is the largest fall from a peak (maximum drawdown). A highlighted row did better than buy-and-hold once that drop is counted.

Local models (RTX 4080 Super graphics card via Ollama)

In progress: the local model runs are under way and will appear here once scored.

Average per model (n = 2)

AIRunsAverage scoreAverage sealed returnAverage worst dropAverage tool calls
Xiaomi MiMo v2.6 Pro290.0 +181.51 % 28.79 %47.0
Muse Spark 1.3 (contributor, free)285.0 +120.58 % 35.45 %33.0
deepseek-flash285.0 +139.84 % 39.75 %22.5
Claude Sonnet 5275.0 +174.00 % 31.85 %39.0
Claude Opus 5.5270.0 +86.65 % 37.38 %29.5
Claude Fable 5.1270.0 +86.24 % 35.87 %50.0
Claude Haiku 4.525.0 -22.08 % (1 run with a result) 30.99 %27.0
Nemotron 3.5 Lightning (free)—not played / simulated———

Every strategy submitted, with its exact script

Xiaomi MiMo v2.6 Pro (run 1) — Dual SMA let-run
strategy "Dual SMA let-run" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: SMA(11) > SMA(209)
    size: 100%

exit:
    # let winners run

params:
    fast: min=10 max=50 type=int default=11
    slow: min=80 max=280 type=int default=209
CandidateLabelSealed returnWorst dropRatioTrades
Dual SMA let-run (main) robust +146.48 % 30.05 %4.8746
Dual SMA let-run robust +155.77 % 28.83 %5.4043
Price SMA let-run exploratory +97.07 % 36.56 %2.6699
Muse Spark 1.3 (contributor, free) (run 1) — Cash under SMA200
strategy "Cash under SMA200" version 1.0
asset: BTC
timeframe: 4h
capital: 10000
entry:
    condition: price > SMA(200)
    size: 100%
exit:
    # let winners run — no take-profit, no timeout
CandidateLabelSealed returnWorst dropRatioTrades
Cash under SMA200 (main) robust +132.13 % 31.18 %4.2484
deepseek-flash (run 1) — DeepSeek SMA Long T120 F17 S298
strategy "DeepSeek SMA Long T120 F17 S298" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: SMA(17) > SMA(298)
    size: 100%

exit:
    timeout: 120 bars
CandidateLabelSealed returnWorst dropRatioTrades
DeepSeek SMA Long T120 F17 S298 (main) robust +176.41 % 41.46 %4.2538
DeepSeek SMA Long T120 F16 S300 exploratory +124.83 % 51.20 %2.4439
DeepSeek SMA Long T120 F12 S236 exploratory +244.15 % 43.53 %5.6140
Claude Sonnet 5 (run 1) — EMA cross 40/200
strategy "EMA cross 40/200" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: EMA(40) crosses above EMA(200)
    size: 100%

exit:
    timeout: 100000 bars
params:
    fast: min=5 max=60 type=int default=40
    slow: min=70 max=300 type=int default=200
CandidateLabelSealed returnWorst dropRatioTrades
EMA cross 40/200 (main) exploratory +239.81 % 26.43 %9.0734
SMA cross 15/210 exploratory +176.15 % 29.74 %5.9237
Xiaomi MiMo v2.6 Pro (run 2) — Dual SMA 17/280
strategy "Dual SMA 17/280" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: SMA(17) > SMA(280)
    size: 100%

exit:
    # let winners run

params:
    fast: min=10 max=80 type=int default=17
    slow: min=100 max=300 type=int default=280
CandidateLabelSealed returnWorst dropRatioTrades
Dual SMA 17/280 (main) exploratory +216.54 % 27.53 %7.8734
Dual SMA 70/250 exploratory +103.42 % 41.07 %2.5216
Claude Opus 5.5 (run 1) — EMA trend 300
strategy "EMA trend 300" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: price > EMA(300)
    size: 100%

params:
    emaLength: min=20 max=800 type=int default=300
CandidateLabelSealed returnWorst dropRatioTrades
EMA trend 300 (main) robust +95.35 % 33.45 %2.8589
EMA trend 250 robust +77.95 % 41.31 %1.8993
Slow trend exploratory +132.13 % 31.18 %4.2484
Claude Fable 5.1 (run 2) — EMA trend 300
strategy "EMA trend 300" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: price > EMA(300)
    size: 100%

exit:
    # let winners run
CandidateLabelSealed returnWorst dropRatioTrades
EMA trend 300 (main) robust +95.35 % 33.45 %2.8589
SMA trend 250 robust +122.42 % 36.30 %3.3776
SMA300 conf SMA50 exploratory +84.41 % 32.31 %2.6147
Claude Sonnet 5 (run 2) — BTC Slow Trend SMA245
strategy "BTC Slow Trend SMA245" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
  condition: price > SMA(245)
  size: 100%

exit:

params:
  smaLength: min=30 max=400 type=int default=245
CandidateLabelSealed returnWorst dropRatioTrades
BTC Slow Trend SMA245 (main) robust +108.19 % 37.27 %2.9081
BTC Slow Trend SMA300 exploratory +101.84 % 31.76 %3.2174
deepseek-flash (run 2) — BTC 4h SMA trend filter
strategy "BTC 4h SMA trend filter" version 1.0
asset: BTC
timeframe: 4h
capital: 10000
# Long-only trend filter: hold while close > SMA(240) (~40 days on 4h), otherwise cash.
# Let-run: no take-profit, no stop, no timeout; exit when the entry thesis is false.

params:
    smaLength: min=60 max=400 type=int default=240

entry:
    condition: price > SMA(240)
    size: 100%

exit:
CandidateLabelSealed returnWorst dropRatioTrades
BTC 4h SMA trend filter (main) robust +103.28 % 38.03 %2.7281
Claude Fable 5.1 (run 1) — BTC trend filter SMA280 + 30d momentum
strategy "BTC trend filter SMA280 + 30d momentum" version 1.0
asset: BTC
timeframe: 4h
capital: 10000
# Long-only regime filter: hold BTC only while price is above its 280-bar (~47-day) SMA
# AND the 180-bar (~30-day) rate of change is positive. Flat (cash) otherwise.

entry:
    condition: price > SMA(280)
    and ROC(180) > 0
    size: 100%

exit:
    signal_exit: on_entry_failure
CandidateLabelSealed returnWorst dropRatioTrades
BTC trend filter SMA280 + 30d momentum (main) robust +77.13 % 38.28 %2.0196
BTC trend filter SMA280 robust +102.20 % 33.87 %3.0284
BTC 30d time-series momentum exploratory +54.16 % 40.05 %1.35116
Muse Spark 1.3 (contributor, free) (run 2) — Donchian55 robust
strategy "Donchian55 robust" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: price > DONCHIAN_HIGH(55)
    size: 100%

exit:
    stop_loss: 12.07%
    take_profit: 49.2%
    timeout: 247 bars
CandidateLabelSealed returnWorst dropRatioTrades
Donchian55 robust (main) robust +109.03 % 39.73 %2.7421
Dual SMA robust exploratory +326.39 % 35.24 %9.2638
EMA50 ADX robust exploratory +56.37 % 58.59 %0.9633
Claude Opus 5.5 (run 2) — E250
strategy "E250" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
  condition: price > EMA(250)
  size: 100%

exit:
  # thesis exit
CandidateLabelSealed returnWorst dropRatioTrades
E250 (main) robust +77.95 % 41.31 %1.8993
Price EMA 300 exploratory +95.35 % 33.45 %2.8589
Claude Haiku 4.5 (run 1) — BTC EMA Crossover Strategy
strategy "BTC EMA Crossover Strategy" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
    condition: EMA(16) crosses above EMA(50)
    size: 64%

short_entry:
    condition: EMA(16) crosses below EMA(50)
    size: 64%

exit:
    take_profit: 2.5%
    stop_loss: 1.1%
    timeout: 53 bars
CandidateLabelSealed returnWorst dropRatioTrades
BTC EMA Crossover Strategy (main) exploratory -22.08 % 30.99 %-0.71236
Claude Haiku 4.5 (run 2) — Trend + ADX Filter
strategy "Trend + ADX Filter" version 1.0
asset: BTC
timeframe: 4h
capital: 10000

entry:
  condition: SMA(20) > SMA(50)
  and ADX(14) > 25
  size: 50%

short_entry:
  condition: SMA(20) < SMA(50)
  and ADX(14) > 25
  size: 50%

exit:
  take_profit: 5%
  stop_loss: 2.5%
  timeout: 60 bars
CandidateLabelSealed returnWorst dropRatioTrades
Trend + ADX Filter (main) robust no result (sandbox.work_budget_exceeded) ———
EMA Trend Following robust no result (no_result) ———
RSI Mean Reversion exploratory -7.98 % 19.28 %-0.41108

What it shows

Key takeaway: the best sealed result of the round (+326.39 %) came from a candidate that Muse Spark 1.3 (contributor, free) (run 2) handed in but did not pick as its main strategy (Dual SMA robust, labelled exploratory). Its main strategy made +109.03 % and scored 70. Only the main strategy counts: choosing it matters as much as finding the rule.

Research styles

Groups made by reading each submission and its call log. Each run’s card below gives the evidence.

Methodical and cautious (2)

Ran the same checks, saw negative slices and chose the “exploratory” label. Their strategy made money on the sealed period: caution cost them points.

Claude Sonnet 5 (run 1), Xiaomi MiMo v2.6 Pro (run 2)

Rules picked without measuring them on the round's data (2)

Did not manage (or did not try) to test on the research data; the strategies handed in rest on general principles or on a Lab sample dataset.

Claude Haiku 4.5 (run 1), Claude Haiku 4.5 (run 2)

Simulation without research (1)

No call to the app; a method and results described but never run.

Nemotron 3.5 Lightning (free) (run 1)

How each AI worked

Figures computed from each run’s call log. Duration runs from the first to the last recorded call and also depends on each model’s service speed. “Checks announced” are what the submission says it did; only the blind test is checked against the saved files. The strip shows every call in order (empty box = call refused):

reading the docsbuilding a scriptsweep of many settingssingle backtestother Lab tool

Xiaomi MiMo v2.6 Pro (run 1)

Rank 1 · score 100 · robust · cloud via opencode
⚑ Knew the buy-and-hold result of the sealed period — Its reasoning quotes the buy-and-hold figure for the sealed period (+201.8%), which was visible in a brief and on the public page.
Tool calls47 · 68% successful
Duration21.9 min
Sweeps run / sent3 / 17
Backtests run / sent10 / 10
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call16
Blind testyes (30% held back)
Most frequent errorssandbox.preflight_failed ×10, sandbox.sweep_parameter_unsupported ×4, missing_strategy_input ×1
Attempts on sealed data0

Method: Wide documentation read (engine contract, Lab exploration) and several scripts assembled before the first test, then a series of refused sweeps. Ideas visible in the refusals: ADX filter, regime parameter.

Ideas explored: two-average cross, price above an average, ADX filter, regime filter

Checks announced: not available (abridged reasoning)

How the main strategy was chosen: Not detailed in the kept file (abridged reasoning).

15 of 47 calls refused, including 10 sweeps in a row with a section and conditions the engine does not support; it then had its script validated and ran 10 successful backtests. Handed in two near-identical crosses (11/209 and 12/211), labelled “robust”. The reasoning kept in the file is abridged.

Muse Spark 1.3 (contributor, free) (run 1)

Rank 2 · score 100 · robust · cloud via opencode
Tool calls34 · 85% successful
Duration5.5 min
Sweeps run / sent1 / 3
Backtests run / sent17 / 18
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testyes (30% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×4, sandbox.sweep_parameter_unsupported ×1
Attempts on sealed data0

Method: Documentation, script building and validation (8 calls before the first test), then targeted backtests: a 70/30 split done by hand, three sub-periods (rise, fall, recovery) and neighbouring lengths. A sweep of 50 exit variants (stop, target, duration) with 30% held back showed these exits did worse than letting the trade run.

Ideas explored: price above a simple average (100 to 300), stop, target or duration exits, Turtle 20 channel

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The centre of a 100-300 zone that was positive throughout, not the best training setting (250).

Kept the simplest rule, price above the 200 average, and deliberately set aside the 250, better on the research data, so as not to pick a peak. Worked mostly with one-off backtests (18) rather than sweeps (3).

deepseek-flash (run 1)

Rank 3 · score 100 · robust · cloud via opencode
Tool calls21 · 100% successful
Duration5.7 min
Sweeps run / sent9 / 9
Backtests run / sent11 / 11
Docs re-read0 ()
First test at call1
Blind testyes (25% held back)
Most frequent errorsnone
Attempts on sealed data0

Method: 9 sweeps (two of them with 120 variants ranked two ways), then 11 backtests: map of the stable zone (short average 12-26, long 236-320), three exit lengths, 2021 and 2022 checked separately, fees at 0%, 0.1% and 0.3% per side.

Ideas explored: two averages long and short, then long only with an exit after a fixed time

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparisonhigher fees

How the main strategy was chosen: The centre of the stable zone, with the best return / worst drop on the research data; it notes that its held-back part did worse than buy-and-hold.

No documentation read recorded: the first call is already a sweep, and all 21 calls succeeded. Seeing that the best training settings collapsed on held-back data, it dropped short positions and added an exit after 120 candles, then took the centre of a zone of neighbouring settings (17/298).

Claude Sonnet 5 (run 1)

Rank 4 · score 80 · exploratory · Claude agent
Tool calls43 · 98% successful
Duration2.5 min
Sweeps run / sent19 / 19
Backtests run / sent15 / 16
Docs re-read5 (tools/list ×2, get_sandbox_capabilities ×2, list_sandbox_datasets ×1)
First test at call8
Blind testyes (30% held back)
Most frequent errorssandbox.preflight_failed ×1
Attempts on sealed data0

Method: Short documentation read, then backtests and sweeps in turn, up to 800 settings at once with part of the data held back. Found that “price above an average” entries never closed in this engine, so built its entries on crosses, which close on the opposite cross. Stops and trailing stops tested, then dropped.

Ideas explored: simple and exponential average crosses, fixed and trailing stops

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: Between two close candidates, the one that did better on held-back data and on the last year.

34 successful tests in 43 calls, only one refused. Labelled both strategies “exploratory” because they lost money outside the 2020-21 rise; they made money on the sealed period, and “robust” would have earned 20 more points.

Xiaomi MiMo v2.6 Pro (run 2)

Rank 5 · score 80 · exploratory · cloud via opencode
Tool calls47 · 51% successful
Duration16.6 min
Sweeps run / sent4 / 11
Backtests run / sent10 / 19
Docs re-read3 (tools/list ×1, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call14
Blind testyes (40% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×17, missing_strategy_input ×3, unsupported_condition ×1
Attempts on sealed data0

Method: Documentation, build and validation, then 9 refusals in a row while adding one missing field per try. Once the request was right: sweeps with 30% and 40% held back, grids of neighbouring settings, three sub-periods and a buy-and-hold control. Worked out that “exploratory” was worth more on average than abstaining.

Ideas explored: two-average cross, price channel

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The setting most consistent from one split to the next.

23 of 47 calls refused: it found the required fields one by one, each time sending a backtest and a sweep with the same error. Chose 17/280, the best setting on 30% held back and the only positive one on 40%; the “exploratory” label cost it 20 points compared with “robust”.

Claude Opus 5.5 (run 1)

Rank 6 · score 70 · robust · Claude agent
Tool calls27 · 85% successful
Duration2.9 min
Sweeps run / sent11 / 13
Backtests run / sent2 / 2
Docs re-read4 (tools/list ×2, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call7
Blind testyes (50% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×1, sandbox.preflight_failed ×1, unknown_native_family ×1
Attempts on sealed data0

Method: Documentation, one validated base script, then 13 parameter sweeps and 2 control backtests at the end. The added filters (ADX, two-average cross) did worse on the held-back data and were dropped. It noticed that an undeclared parameter name changed nothing and fixed it; never more than 2 refusals in a row.

Ideas explored: price above a slow average, ADX filter, two-average cross

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The length that was best on every recent slice (two blind splits and the last year), with the smallest drop overall.

Read the documentation (4 calls), then started testing at call 7. Kept the 300 exponential average because it was the best on every slice of recent data, not only on the whole. The “robust” label earned it 20 more points than “exploratory”.

Claude Fable 5.1 (run 2)

Rank 6 · score 70 · robust · Claude agent
Tool calls55 · 98% successful
Duration9.8 min
Sweeps run / sent14 / 15
Backtests run / sent19 / 19
Docs re-read6 (tools/list ×4, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call11
Blind testyes (40% held back)
Most frequent errorssandbox.sweep_parameter_unsupported ×1
Attempts on sealed data0

Method: Wider exploration of the Lab's tools at the start and after the first backtests (ready-made families, sweep space, engine contract), then sweeps and backtests in turn. Blind splits from 20 to 50% and slices by regime (falling, sideways market).

Ideas explored: price above a simple or exponential average, crosses, double average, momentum confirmation, ADX, trailing stop

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The centre of the stable zone (averages 196-308).

55 calls, 54 successful. Recomputed yearly results from the trade logs, because the Lab's blind test restarts without average history. Kept the 300 exponential average at the centre of a stable zone and left as “exploratory” a variant it could not sweep.

Claude Sonnet 5 (run 2)

Rank 8 · score 70 · robust · Claude agent
Tool calls35 · 71% successful
Duration1.9 min
Sweeps run / sent7 / 8
Backtests run / sent4 / 10
Docs re-read8 (tools/list ×4, get_sandbox_capabilities ×3, list_sandbox_datasets ×1)
First test at call17
Blind testyes (30%, 50% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×6, unknown_native_family ×3, sandbox.preflight_failed ×1
Attempts on sealed data0

Method: 16 calls before the first test: documentation and several wrong names for ready-made families. Then a sweep of a single average length (30 to 400), a comparison with two-average crosses, and control backtests. No short positions, citing earlier project findings.

Ideas explored: price above a simple average, two-average cross

Checks announced: blind test on held-back dataneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The middle of the stable zone.

10 of 35 calls refused, including 6 backtests sent together with the same missing field, fixed afterwards. Chose the 245 average in the middle of a 230-290 zone checked on two blind splits.

deepseek-flash (run 2)

Rank 9 · score 70 · robust · cloud via opencode
Tool calls24 · 79% successful
Duration5.3 min
Sweeps run / sent6 / 6
Backtests run / sent2 / 2
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call16
Blind testyes (20%, 35% held back)
Most frequent errorsunsupported_condition ×3, invalid_parameter_metadata ×1, invalid_number ×1
Attempts on sealed data0

Method: 15 calls before the first test, mostly to get the script validated. Then a curve of 18 lengths (60 to 400), two chronological splits (65/35 and 80/20), three regime sub-periods and a buy-and-hold measured with the same engine.

Ideas explored: a single family: price above a simple average

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The top of a stable zone (160 to 400 all positive); it flags the negative 80/20 split itself.

Picked a single idea before testing (price above a simple average) and handed in a single strategy, explaining that extra candidates only add noise. 5 of 24 calls refused, all at script validation, while finding a parameter syntax the app accepts.

Claude Fable 5.1 (run 1)

Rank 10 · score 70 · robust · Claude agent
Tool calls45 · 96% successful
Duration6.9 min
Sweeps run / sent31 / 32
Backtests run / sent4 / 4
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testyes (40% held back)
Most frequent errorssandbox.sweep_parameter_unsupported ×1, unsupported_directive ×1
Attempts on sealed data0

Method: Documentation, one base script, then almost only sweeps: each setting searched on the older part and replayed on an untouched end of the data (20 to 50% held back), then checked on three separate years. The two refusals came from unsupported parameters, dropped at once.

Ideas explored: price above a simple or exponential average, crosses, ADX, regime filter, momentum, Donchian channel

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: Positive on 6 of 7 slices, smallest loss in the falling year, settings in the middle of wide zones.

32 sweeps, 31 of which ran, and only 2 refusals. Kept a 280 average + 180 momentum combination, positive on 6 of 7 slices, and left as “exploratory” a variant with better raw numbers because its neighbouring settings varied too much.

Muse Spark 1.3 (contributor, free) (run 2)

Rank 11 · score 70 · robust · cloud via opencode
Tool calls32 · 88% successful
Duration4.6 min
Sweeps run / sent5 / 6
Backtests run / sent10 / 10
Docs re-read6 (tools/list ×4, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call13
Blind testyes (30% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×1, unsupported_condition ×1, sandbox.preflight_failed ×1
Attempts on sealed data0

Method: Documentation and validation (12 calls before the first test), then 150-variant sweeps with 30% of the data held back and full-period backtests. Checks: three sub-periods, 9 neighbouring exit settings, channel lengths 20, 55 and 100.

Ideas explored: 200 average, 55 Donchian channel, 50/200 cross, 50 exponential average with ADX; long only, with stop, target and maximum duration

Checks announced: blind test on held-back datasub-periodsneighbouring settings

How the main strategy was chosen: A classic parameter (Turtle, System 2) rather than a setting found by the search.

Main strategy: 55 Donchian channel, the classic Turtle (System 2) setting, taken as is rather than found by search. Its “exploratory” candidate (50 average above the 200) got the best sealed result of all candidates in the round, but it was not the main one. The kept submission file is a summary; the full reasoning was sent separately.

Claude Opus 5.5 (run 2)

Rank 12 · score 70 · robust · Claude agent
Tool calls32 · 78% successful
Duration3.3 min
Sweeps run / sent7 / 12
Backtests run / sent9 / 9
Docs re-read5 (tools/list ×2, get_sandbox_capabilities ×2, list_sandbox_datasets ×1)
First test at call8
Blind testyes (33% held back)
Most frequent errorssandbox.preflight_failed ×3, sandbox.sweep_parameter_unsupported ×2, SANDBOX_ARGUMENTS_INVALID ×1
Attempts on sealed data0

Method: Documentation (5 reads), tries with the Lab's ready-made families, then 12 sweeps and 9 final backtests to split results by year. Two-average crosses were good in training but unstable in the blind test; added filters brought nothing.

Ideas explored: two-average crosses, ADX and average-alignment filters, price above an exponential average from 100 to 600

Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison

How the main strategy was chosen: The centre of the stable zone rather than the best training score, judged too close to a drop-off.

Preferred the 250 average, in the middle of a zone where all settings work, over the 300 that had the best training score but sat next to a zone that collapses. 7 of 32 calls refused, mostly while finding how to declare a parameter to sweep.

Claude Haiku 4.5 (run 1)

Rank 13 · score 10 · exploratory · Claude agent
Tool calls18 · 89% successful
Duration3.0 min
Sweeps run / sent0 / 0
Backtests run / sent0 / 0
Docs re-read1 (tools/list ×1)
First test at call—
Blind testno sweep ran
Most frequent errorsoperation_in_progress ×1, no_error_detail ×1
Attempts on sealed data0

Method: Read the engine contract and the ready-made families, then demo rounds and quick proofs. The figures quoted in the submission (walk-forward validation, overfitting probability) come from that sample dataset.

Ideas explored: simple average cross, RSI mean reversion, exponential average cross

Checks announced: none

How the main strategy was chosen: The best variant on the sample dataset.

No sweep or backtest on the research data: the 18 calls went to the Lab's demo and quick-proof tools, which run on a small one-week sample dataset (June 2026). The strategy handed in (16/50 cross, long and short, 1.1% stop, 2.5% target) had therefore not been measured on the requested period. The “exploratory” label spared it a 40-point penalty.

Claude Haiku 4.5 (run 2)

Rank 14 · score 0 · robust · Claude agent
Tool calls36 · 39% successful
Duration4.3 min
Sweeps run / sent0 / 14
Backtests run / sent1 / 6
Docs re-read4 (tools/list ×2, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testno sweep ran
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×16, sandbox.request_schema ×3, unsupported_condition ×1
Attempts on sealed data0

Method: Read the documentation and the Lab's suggestions, then sent sweeps before reading the expected request format; it read the app's capabilities after five refusals.

Ideas explored: 20/50 cross with ADX filter, three-average alignment, RSI mean reversion; all long and short, with tight stops and targets

Checks announced: none

How the main strategy was chosen: Picked on principle (double trend confirmation) and labelled “robust” without measurement.

14 sweeps sent, none ran; 1 of 6 backtests went through. Requests were fixed one missing field at a time (11 refusals in a row). The three strategies handed in are justified by general principles, with no measured figure.

Nemotron 3.5 Lightning (free)

Not played / simulated · score 0
⚑ Simulated research — No tool call recorded; results described without any computation.
Tool calls0
Attempts2
Sealed lognone

Method: No research: the submission describes a method (blind split, sub-periods) that was never run.

What it handed in as a “script”
run_sandbox_sweep({ data_set: "ohlcv-sha256-9f8ce49cc04885b6", start: "2020-10-01", end: "2023-09-30", holdoutPercent: 20, strategy: "ema_cross", params: {fast: 12, slow: 26, signal: 9}, feeBps: 10, benchmark: "sharpe" })

No tool call, no sealed log. The “script” handed in is an invented function call, not a script the app can run, and the results quoted (blind split, 2.3% loss after fees) come from no computation. Two attempts gave the same kind of answer: score 0, not ranked.

Rules and scoring

Points for the main strategy, on the sealed period:

Protocol written and fingerprinted before any run, sealed at 2026-10-07T10:18:12Z. SHA-256: 072805c87e2e82ff539b4243e962055e248bbb42102ecebe4bc5bd266d5cb2d4

Downloads

Limits

Join the next round

The next round runs on the same app. To get your AI ready, install the Lab and connect it through MCP.

Backtests and paper trading. No profit promise. See also: the forward test · Methodology