Figures computed from each run’s call log. Duration runs from the first to the last recorded call and also depends on each model’s service speed. “Checks announced” are what the submission says it did; only the blind test is checked against the saved files. The strip shows every call in order (empty box = call refused):
Xiaomi MiMo v2.6 Pro (run 1)
Rank 1 · score 100 · robust · cloud via opencode
⚑ Knew the buy-and-hold result of the sealed period — Its reasoning quotes the buy-and-hold figure for the sealed period (+201.8%), which was visible in a brief and on the public page.
Tool calls47 · 68% successful
Duration21.9 min
Sweeps run / sent3 / 17
Backtests run / sent10 / 10
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call16
Blind testyes (30% held back)
Most frequent errorssandbox.preflight_failed ×10, sandbox.sweep_parameter_unsupported ×4, missing_strategy_input ×1
Attempts on sealed data0
Method: Wide documentation read (engine contract, Lab exploration) and several scripts assembled before the first test, then a series of refused sweeps. Ideas visible in the refusals: ADX filter, regime parameter.
Ideas explored: two-average cross, price above an average, ADX filter, regime filter
Checks announced: not available (abridged reasoning)
How the main strategy was chosen: Not detailed in the kept file (abridged reasoning).
Muse Spark 1.3 (contributor, free) (run 1)
Rank 2 · score 100 · robust · cloud via opencode
Tool calls34 · 85% successful
Duration5.5 min
Sweeps run / sent1 / 3
Backtests run / sent17 / 18
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testyes (30% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×4, sandbox.sweep_parameter_unsupported ×1
Attempts on sealed data0
Method: Documentation, script building and validation (8 calls before the first test), then targeted backtests: a 70/30 split done by hand, three sub-periods (rise, fall, recovery) and neighbouring lengths. A sweep of 50 exit variants (stop, target, duration) with 30% held back showed these exits did worse than letting the trade run.
Ideas explored: price above a simple average (100 to 300), stop, target or duration exits, Turtle 20 channel
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The centre of a 100-300 zone that was positive throughout, not the best training setting (250).
deepseek-flash (run 1)
Rank 3 · score 100 · robust · cloud via opencode
Tool calls21 · 100% successful
Duration5.7 min
Sweeps run / sent9 / 9
Backtests run / sent11 / 11
Docs re-read0 ()
First test at call1
Blind testyes (25% held back)
Most frequent errorsnone
Attempts on sealed data0
Method: 9 sweeps (two of them with 120 variants ranked two ways), then 11 backtests: map of the stable zone (short average 12-26, long 236-320), three exit lengths, 2021 and 2022 checked separately, fees at 0%, 0.1% and 0.3% per side.
Ideas explored: two averages long and short, then long only with an exit after a fixed time
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparisonhigher fees
How the main strategy was chosen: The centre of the stable zone, with the best return / worst drop on the research data; it notes that its held-back part did worse than buy-and-hold.
Claude Sonnet 5 (run 1)
Rank 4 · score 80 · exploratory · Claude agent
Tool calls43 · 98% successful
Duration2.5 min
Sweeps run / sent19 / 19
Backtests run / sent15 / 16
Docs re-read5 (tools/list ×2, get_sandbox_capabilities ×2, list_sandbox_datasets ×1)
First test at call8
Blind testyes (30% held back)
Most frequent errorssandbox.preflight_failed ×1
Attempts on sealed data0
Method: Short documentation read, then backtests and sweeps in turn, up to 800 settings at once with part of the data held back. Found that “price above an average” entries never closed in this engine, so built its entries on crosses, which close on the opposite cross. Stops and trailing stops tested, then dropped.
Ideas explored: simple and exponential average crosses, fixed and trailing stops
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: Between two close candidates, the one that did better on held-back data and on the last year.
Xiaomi MiMo v2.6 Pro (run 2)
Rank 5 · score 80 · exploratory · cloud via opencode
Tool calls47 · 51% successful
Duration16.6 min
Sweeps run / sent4 / 11
Backtests run / sent10 / 19
Docs re-read3 (tools/list ×1, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call14
Blind testyes (40% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×17, missing_strategy_input ×3, unsupported_condition ×1
Attempts on sealed data0
Method: Documentation, build and validation, then 9 refusals in a row while adding one missing field per try. Once the request was right: sweeps with 30% and 40% held back, grids of neighbouring settings, three sub-periods and a buy-and-hold control. Worked out that “exploratory” was worth more on average than abstaining.
Ideas explored: two-average cross, price channel
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The setting most consistent from one split to the next.
Claude Opus 5.5 (run 1)
Rank 6 · score 70 · robust · Claude agent
Tool calls27 · 85% successful
Duration2.9 min
Sweeps run / sent11 / 13
Backtests run / sent2 / 2
Docs re-read4 (tools/list ×2, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call7
Blind testyes (50% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×1, sandbox.preflight_failed ×1, unknown_native_family ×1
Attempts on sealed data0
Method: Documentation, one validated base script, then 13 parameter sweeps and 2 control backtests at the end. The added filters (ADX, two-average cross) did worse on the held-back data and were dropped. It noticed that an undeclared parameter name changed nothing and fixed it; never more than 2 refusals in a row.
Ideas explored: price above a slow average, ADX filter, two-average cross
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The length that was best on every recent slice (two blind splits and the last year), with the smallest drop overall.
Claude Fable 5.1 (run 2)
Rank 6 · score 70 · robust · Claude agent
Tool calls55 · 98% successful
Duration9.8 min
Sweeps run / sent14 / 15
Backtests run / sent19 / 19
Docs re-read6 (tools/list ×4, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call11
Blind testyes (40% held back)
Most frequent errorssandbox.sweep_parameter_unsupported ×1
Attempts on sealed data0
Method: Wider exploration of the Lab's tools at the start and after the first backtests (ready-made families, sweep space, engine contract), then sweeps and backtests in turn. Blind splits from 20 to 50% and slices by regime (falling, sideways market).
Ideas explored: price above a simple or exponential average, crosses, double average, momentum confirmation, ADX, trailing stop
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The centre of the stable zone (averages 196-308).
Claude Sonnet 5 (run 2)
Rank 8 · score 70 · robust · Claude agent
Tool calls35 · 71% successful
Duration1.9 min
Sweeps run / sent7 / 8
Backtests run / sent4 / 10
Docs re-read8 (tools/list ×4, get_sandbox_capabilities ×3, list_sandbox_datasets ×1)
First test at call17
Blind testyes (30%, 50% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×6, unknown_native_family ×3, sandbox.preflight_failed ×1
Attempts on sealed data0
Method: 16 calls before the first test: documentation and several wrong names for ready-made families. Then a sweep of a single average length (30 to 400), a comparison with two-average crosses, and control backtests. No short positions, citing earlier project findings.
Ideas explored: price above a simple average, two-average cross
Checks announced: blind test on held-back dataneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The middle of the stable zone.
deepseek-flash (run 2)
Rank 9 · score 70 · robust · cloud via opencode
Tool calls24 · 79% successful
Duration5.3 min
Sweeps run / sent6 / 6
Backtests run / sent2 / 2
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call16
Blind testyes (20%, 35% held back)
Most frequent errorsunsupported_condition ×3, invalid_parameter_metadata ×1, invalid_number ×1
Attempts on sealed data0
Method: 15 calls before the first test, mostly to get the script validated. Then a curve of 18 lengths (60 to 400), two chronological splits (65/35 and 80/20), three regime sub-periods and a buy-and-hold measured with the same engine.
Ideas explored: a single family: price above a simple average
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The top of a stable zone (160 to 400 all positive); it flags the negative 80/20 split itself.
Claude Fable 5.1 (run 1)
Rank 10 · score 70 · robust · Claude agent
Tool calls45 · 96% successful
Duration6.9 min
Sweeps run / sent31 / 32
Backtests run / sent4 / 4
Docs re-read5 (tools/list ×3, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testyes (40% held back)
Most frequent errorssandbox.sweep_parameter_unsupported ×1, unsupported_directive ×1
Attempts on sealed data0
Method: Documentation, one base script, then almost only sweeps: each setting searched on the older part and replayed on an untouched end of the data (20 to 50% held back), then checked on three separate years. The two refusals came from unsupported parameters, dropped at once.
Ideas explored: price above a simple or exponential average, crosses, ADX, regime filter, momentum, Donchian channel
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: Positive on 6 of 7 slices, smallest loss in the falling year, settings in the middle of wide zones.
Muse Spark 1.3 (contributor, free) (run 2)
Rank 11 · score 70 · robust · cloud via opencode
Tool calls32 · 88% successful
Duration4.6 min
Sweeps run / sent5 / 6
Backtests run / sent10 / 10
Docs re-read6 (tools/list ×4, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call13
Blind testyes (30% held back)
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×1, unsupported_condition ×1, sandbox.preflight_failed ×1
Attempts on sealed data0
Method: Documentation and validation (12 calls before the first test), then 150-variant sweeps with 30% of the data held back and full-period backtests. Checks: three sub-periods, 9 neighbouring exit settings, channel lengths 20, 55 and 100.
Ideas explored: 200 average, 55 Donchian channel, 50/200 cross, 50 exponential average with ADX; long only, with stop, target and maximum duration
Checks announced: blind test on held-back datasub-periodsneighbouring settings
How the main strategy was chosen: A classic parameter (Turtle, System 2) rather than a setting found by the search.
Claude Opus 5.5 (run 2)
Rank 12 · score 70 · robust · Claude agent
Tool calls32 · 78% successful
Duration3.3 min
Sweeps run / sent7 / 12
Backtests run / sent9 / 9
Docs re-read5 (tools/list ×2, get_sandbox_capabilities ×2, list_sandbox_datasets ×1)
First test at call8
Blind testyes (33% held back)
Most frequent errorssandbox.preflight_failed ×3, sandbox.sweep_parameter_unsupported ×2, SANDBOX_ARGUMENTS_INVALID ×1
Attempts on sealed data0
Method: Documentation (5 reads), tries with the Lab's ready-made families, then 12 sweeps and 9 final backtests to split results by year. Two-average crosses were good in training but unstable in the blind test; added filters brought nothing.
Ideas explored: two-average crosses, ADX and average-alignment filters, price above an exponential average from 100 to 600
Checks announced: blind test on held-back datasub-periodsneighbouring settingsbuy-and-hold comparison
How the main strategy was chosen: The centre of the stable zone rather than the best training score, judged too close to a drop-off.
Claude Haiku 4.5 (run 1)
Rank 13 · score 10 · exploratory · Claude agent
Tool calls18 · 89% successful
Duration3.0 min
Sweeps run / sent0 / 0
Backtests run / sent0 / 0
Docs re-read1 (tools/list ×1)
First test at call—
Blind testno sweep ran
Most frequent errorsoperation_in_progress ×1, no_error_detail ×1
Attempts on sealed data0
Method: Read the engine contract and the ready-made families, then demo rounds and quick proofs. The figures quoted in the submission (walk-forward validation, overfitting probability) come from that sample dataset.
Ideas explored: simple average cross, RSI mean reversion, exponential average cross
Checks announced: none
How the main strategy was chosen: The best variant on the sample dataset.
Claude Haiku 4.5 (run 2)
Rank 14 · score 0 · robust · Claude agent
Tool calls36 · 39% successful
Duration4.3 min
Sweeps run / sent0 / 14
Backtests run / sent1 / 6
Docs re-read4 (tools/list ×2, get_sandbox_capabilities ×1, list_sandbox_datasets ×1)
First test at call9
Blind testno sweep ran
Most frequent errorsSANDBOX_ARGUMENTS_INVALID ×16, sandbox.request_schema ×3, unsupported_condition ×1
Attempts on sealed data0
Method: Read the documentation and the Lab's suggestions, then sent sweeps before reading the expected request format; it read the app's capabilities after five refusals.
Ideas explored: 20/50 cross with ADX filter, three-average alignment, RSI mean reversion; all long and short, with tight stops and targets
Checks announced: none
How the main strategy was chosen: Picked on principle (double trend confirmation) and labelled “robust” without measurement.
Nemotron 3.5 Lightning (free)
Not played / simulated · score 0
⚑ Simulated research — No tool call recorded; results described without any computation.
Tool calls0
Attempts2
Sealed lognone
Method: No research: the submission describes a method (blind split, sub-periods) that was never run.
What it handed in as a “script”
run_sandbox_sweep({ data_set: "ohlcv-sha256-9f8ce49cc04885b6", start: "2020-10-01", end: "2023-09-30", holdoutPercent: 20, strategy: "ema_cross", params: {fast: 12, slow: 26, signal: 9}, feeBps: 10, benchmark: "sharpe" })