Skip to main content
← Back to blog

Which AI for trading? Gemini, GPT, Grok and Claude scored

📅 2026-10-08
✍️ Strategy Arena
ai for trading which ai for trading gemini gpt astra grok claude deepseek ai benchmark ai leaderboard backtest mcp

Results refereed on 7 and 8 October 2026. The sealed figures come from the referee files, published on the AI researcher leaderboard.

The short answer

On 7 and 8 October 2026, we put Gemini 3.8 Flash, GPT 5.6 Sol, Grok 4.6, Grok 4.7 and Astra 6 through the same test as Claude, DeepSeek, MiMo and Muse: find a Bitcoin strategy with the same backtest app, then hand it to a market period no AI had seen.

  • Gemini 3.8 Flash scored 100 out of 100 in both of its sessions (results of 7 October 2026), with the same rule found twice: 12-candle simple average above the 210 average.
  • GPT 5.6 Sol built the most elaborate rules (channel breakouts, momentum, tight stops) and scores 70 points per session (7 October 2026), with a wide gap between what it reported on its research data and the sealed result.
  • Astra 6 (8 October 2026) handed in the same single-condition rule twice: price above the 240 simple average, otherwise cash. +103.28% on the sealed period, “exploratory” label, 50 points per session.
  • Grok 4.6 (a single session) and Grok 4.7 also chose the “exploratory” label; Grok 4.7 abstained in its second session after exposing its own mirages.
  • Of the 23 scored runs as of 8 October 2026, 7 beat buy-and-hold once risk is counted.

So “which AI for trading?” gets a measured answer, not an opinion. It also has limits, set out below.

Why this ranking does not look like the usual comparisons

Most AI trading comparisons describe a “personality”: one model would be cautious, another aggressive. Those are impressions. Here every AI did the same job under the same conditions, and the judge reads numbers only.

  • Same tool. Each AI drove Strategy Arena Lab through MCP, the protocol that lets an AI use a piece of software. No copy-pasted code, no human help during the session.
  • Same data. Bitcoin 4-hour candles, from 1 October 2020 to 30 September 2023 for research. The sealed helper refuses any other data.
  • Same budget. At most 80 tool calls per session.
  • Sealed period. The referee then replays the main strategy from 1 October 2023 to 30 September 2026, data the AI never touched.
  • The AI commits. It labels its strategy “robust” or “exploratory”. A “robust” strategy that loses money costs points. Saying “nothing works” (abstention) earns fixed points.
  • Everything is published. The protocol, its SHA-256 fingerprint computed before any run, every submission, every call log and every referee result can be downloaded from the leaderboard page. The methodology covers the rest.

This setup measures one precise skill: searching for a rule, checking it, and judging its strength yourself. That is what you want from an AI that helps you trade.

The sessions of 7 and 8 October 2026

Reference: buying Bitcoin at the start of the sealed period and holding it to the end returns +201.82%, with a worst drop of 53.45% (return / worst drop: 3.78), according to the referee on 7 October 2026.

AI Session Main strategy Label Sealed return Worst drop Return / drop Score Calls
Gemini 3.8 Flash 1 Dual SMA 12-210 robust +152.72% 30.20% 5.06 100 71
Gemini 3.8 Flash 2 Dual SMA 12-210 robust +152.72% 30.20% 5.06 100 38
GPT 5.6 Sol 1 BTC Filtered Breakout 60-250-40 robust +20.91% 29.00% 0.72 70 22
GPT 5.6 Sol 2 Event momentum robust robust +62.95% 36.99% 1.70 70 31
Grok 4.6 (n = 1) 1 sma200_cd12 exploratory +108.76% 30.99% 3.51 50 20
Grok 4.7 1 ema 170 long exploratory +99.92% 36.60% 2.73 50 40
Grok 4.7 2 none (abstention) — — — — 25 11
Astra 6 1 Astra6 SMA240 long cash exploratory +103.28% 38.03% 2.72 50 34
Astra 6 2 Astra6 S2 SMA240 trend exploratory +103.28% 38.03% 2.72 50 26
Buy and hold — reference — +201.82% 53.45% 3.78 — —

Sealed figures of 7 and 8 October 2026, fees of 0.1% on each buy and sell. “Calls” = tool calls recorded in the session's sealed log.

Average per model, all runs

AI Runs Average score Average sealed return
Gemini 3.8 Flash 2 100.0 +152.72%
Xiaomi MiMo v2.6 2 90.0 +181.51%
Muse Spark 1.3 2 85.0 +120.58%
deepseek-flash 2 85.0 +139.84%
Claude Sonnet 5 2 75.0 +174.00%
GPT 5.6 Sol 2 70.0 +41.93%
Claude Opus 5.5 2 70.0 +86.65%
Claude Fable 5.1 2 70.0 +86.24%
Grok 4.6 1 50.0 +108.76%
Astra 6 2 50.0 +103.28%
Grok 4.7 2 37.5 +99.92% (1 run with a result)
Claude Haiku 4.5 2 5.0 −22.08% (1 run with a result)

Averages as of 8 October 2026, computed from the referee files. Grok 4.6 played a single session: its row carries less weight than the others.

What each AI did

Gemini 3.8 Flash: test fast, keep a simple rule

Gemini tested from its first calls, without a long read of the documentation. In its first session it ran 57 backtests and 7 settings sweeps in 71 calls (log of 7 October 2026). Its conclusion: stop, target or timeout exits hurt the held-back part of the data, while a “let it run” rule held. It mapped a grid of 25 settings around its choice, checked three sub-periods (rise, fall, recovery) and tripled the fees.

In its second, independent session it landed on exactly the same rule: 12-candle simple average above the 210 one, long only. Same rule, same sealed result, 100 points both times (7 October 2026). Its secondary candidates in the second session even did better on the sealed period, but only the main strategy counts.

GPT 5.6 Sol: more complex rules, a wider gap

GPT 5.6 Sol dropped moving-average crosses, which it judged degraded out of sample, in favour of channel breakouts filtered by an exponential average, then a momentum rule with seven settings (stop, target, trailing stop, pause). Its method is serious: chronological splits, sub-periods, neighbouring settings.

The sealed result tells another story. Its first strategy reported +254.52% on the research data and makes +20.91% on the sealed period (refereed on 7 October 2026). The second reported +341.78% and makes +62.95%. Both stay positive, so the “robust” label costs no points, but neither beats buy-and-hold once risk is counted. The more settings a rule has, the more it can fit the research data.

Grok 4.6: a short, cautious session

Grok 4.6 played a single session, in 20 calls out of 80, all successful (log of 7 October 2026). Its rule: stay in Bitcoin while the price is above the 200 simple average, with a 12-candle pause after each exit. It labelled it “exploratory” because the edge came mostly from dodging the crash. Its second candidate, the same rule without the pause, did better on the sealed period. With n = 1, its score of 50 says little about its consistency.

Astra 6: a one-line rule, frozen before the check

Astra 6 played two independent sessions on 8 October 2026, in 34 then 26 calls out of 80, all successful, with no attempt refused by the sealed helper (logs of 8 October 2026). Both times it compared families of rules (two averages crossing, price above an average, momentum or RSI dips), froze its choice before looking at the end of the research data, then doubled the fees and tested next-open execution.

Both sessions handed in the same rule: stay in Bitcoin while the price is above the 240-candle simple average, otherwise go to cash. On the sealed period it makes +103.28% with a worst drop of 38.03%, a ratio of 2.72 against 3.78 for buy-and-hold. According to its submissions, the rule did not beat buy-and-hold once risk is counted on the held-back part: it labelled it “exploratory”, and the sealed result bears it out on that point.

Grok 4.7: exposing its own mirages

Grok 4.7's first session handed in a simple rule, price above the 170 exponential average, labelled “exploratory” because the second half of its research data made almost nothing.

The second session is the most instructive of the lot. Grok 4.7 fixed its selection rule before testing, then replayed its ideas with the Lab's indicator_lookback_v2 mode, where the second half keeps the first half's indicator history. The apparent edge of the 220 to 240 averages vanished, and the remaining positive cells were isolated spikes. It abstained: 25 points, the abstention score. On this rising sealed period, abstaining earns less than a winning trend strategy. It is still the right answer when the research finds nothing solid.

The findings that matter when choosing an AI

  1. The shape of the rule weighs more than the model's name. As of 8 October 2026, the 6 runs that handed in a slow cross of two averages (short of 11 to 40 candles, long of 200 to 298) all beat buy-and-hold once risk is counted. Of the 11 runs that kept “price above a long average” (both Astra 6 runs included), only one does.
  2. Picking the main strategy is a test in itself. Gemini, GPT and Grok all handed in secondary candidates that beat their main one in at least one session. Choosing without seeing the sealed period is part of the skill being measured.
  3. The label costs or pays. Saying “robust” on a winning strategy earns 30 points, saying “exploratory” earns 10. The Grok models and Astra 6 were cautious, and the scoring made them pay for it on this period. A clear case: Astra 6 and deepseek-flash's second session handed in the same rule (price above the 240 simple average), with the same sealed result; “robust” earned deepseek-flash 70 points, “exploratory” earned Astra 6 50.
  4. The number of settings is a warning sign. Among the sessions of 7 October 2026, the most elaborate rules (GPT 5.6 Sol) lost the most between the return reported on research and the sealed return.

For a trader, the lesson is practical: the AI is a research assistant, not an oracle. What makes it useful is the frame it works in: held-back data, realistic fees, a label it stands by, the right to say “nothing works”. Our earlier “which AI to use for trading” comparison describes the AIs of the live arena; this bench measures their work as researchers.

The limits of this ranking

  • n = 2 per model, n = 1 for Grok 4.6. The order between models can change with more sessions.
  • One historical period, one market. Bitcoin 4h, October 2023 to September 2026. A rising period favours long-only trend rules.
  • The models may know Bitcoin from their training. Round 1 is played on the past. Round 2 will be played on future data, unknown to everyone, like our forward test.
  • The referee is also a contestant. The referee session is Claude Opus 5.5. Scoring is mechanical (rules applied to backtest numbers) and the conflict is stated on the page.
  • Infrastructure of the 7 October sessions. They ran on Lab 235, whose default engine is bit-identical to the 230 used before (84 of 84 answers identical), with the sealed helper forced onto the “full” MCP profile (same 74 tools and same answers).
  • Infrastructure of the 8 October sessions. Astra 6 played on Lab 236, with the sealed helper again forced onto the “full” profile; proof and replay tools, whose built-in data overlaps the sealed period, were refused there. No attempt was refused.
  • grok-1 log cleaned and archived. 4 Grok 4.7 calls started by mistake under Grok 4.6's id were removed from its log and archived separately. The scored log keeps Grok 4.6's 20 calls, and the archive can be downloaded.

Put your own AI through the same bench

The Lab is free on Windows: download Strategy Arena Lab, then connect your AI through MCP. Each session's detail, with exact scripts, logs and method, is on the AI researcher leaderboard.

Already have a strategy, written by you or by an AI? Test it on data it has never seen: free Quick Check.

To look at one of these rules on a chart, a free TradingView account is enough: open TradingView. Partner link: Strategy Arena may earn a commission on a first subscription. It has no effect on the ranking, which is computed before and without this link. The leaderboard scripts are in ArenaScript, not Pine: the figures above come from our sealed backtest only.

Backtests on historical data, not investment advice. Past performance does not predict future results.

⚠️ Disclaimer — This article is for informational and educational purposes only. It does not constitute investment advice or a buy/sell recommendation. Past performance does not guarantee future results. Strategy Arena is an educational simulator with virtual capital. Always do your own research before making investment decisions.

Enjoyed this article? Share it

𝕏 Share on X ✈️ Telegram

Turn a local result into durable proof

The Lab runs free on your PC. Paid options only keep or expand the work — they never change the verdict.

Get the Lab (free) Builder / Operator Publish a Lab Report Lab Reports Hall
TradingView Starter
Choose the account path. Prove the strategy separately.
Your first chart and price alert do not require the Lab. Strategy Arena proof remains required before trusting or exporting a strategy.
Partner attribution is for first-time users and web purchases only. Strategy Arena may earn a commission on eligible purchases; starting with a free account needs no payment.
I have never had an account · Start free I already have an account Check a strategy
More TradingView paths
Evidence-gated Pine review