Rust Monte Carlo backtesting: 10,000 local simulations
Strategy Arena Lab resamples closed trades in blocks after chronological holdout to measure P5/P50/P95 tails, drawdown and observed ruin risk without uploading the data to a server.
What the Rust engine actually computes
The candidate is selected on training data first. The sealed chronological holdout is then opened once. Monte Carlo runs only after that decision to perturb the order of already observed trades; it never participates in candidate ranking.
Computation receipt
- 10000 deterministic paths by default
- ChaCha8 RNG with seed and SHA-256 input fingerprint
- Circular moving-block bootstrap, block size near √(trade count)
- P5/P50/P95 equity, return and maximum drawdown
- Observed profit, loss and 50% ruin-threshold frequencies
- Below 10 trades: zero simulations and no fake probability
Public server framework: 1000 simulations and verifiable snapshots
This historical public pipeline is separate from the local Rust engine. Each step produces a verifiable artifact, not just a Sharpe on one curve.
| Step | Component | Input | Output / gate | Proof |
|---|---|---|---|---|
| 1. Bootstrap | Trade / return resampling | Walk-forward backtest trade series | 1000 simulated PnL paths | monte-carlo.json |
| 2. Percentiles | p5 / p50 / p95 PnL & Sharpe | Bootstrap distribution | p5 PnL > 0 required | /facts/monte-carlo |
| 3. Robustness | Score 0–1 (subsample stability) | Inter-sim variance | Robustness > 0.6 | /live-results |
| 4. Calibration | Brier / reliability | Model probs vs outcomes | Published even when bad | /facts/ml-edge |
| 5. Drift | Walk-forward vs live paper | 5m paper equity | Alert if gap > threshold | /dashboard, /strategy-hospital |
Bootstrap assumes conditionally exchangeable trades — limit documented on /methodology (autocorrelation, regimes).
Five Monte Carlo pitfalls & StrategyArena fixes
- Too few trades — resampling 8 trades creates imaginary precision. Rust fix: below 10 closed trades, zero paths and an insufficient_trades status. Ten is a computation minimum, not sufficient statistical evidence.
- Ignoring the low percentile (p5) — great median, catastrophic left tail. Fix: published p5 PnL; gate p5 ≤ 0 → RECALIBRATE / BUG_SUSPECT.
- i.i.d. bootstrap on autocorrelated returns — overstates confidence. Fix: experimental block bootstrap + mandatory walk-forward in pipeline.
- MC without fees / slippage — inflated percentiles. Fix: same friction model as backtest (methodology).
- Single MC pass to checkbox — no drift monitoring. Fix: monthly re-MC + live paper comparison; snapshots in monte-carlo.json.
Public Monte Carlo tracker stats (updated: 2026-06-11)
Counts synced with strategy-arena.json when available; per-strategy MC detail in monte-carlo.json.
Researcher workflow
Local Lab: train → candidate lock → sealed holdout → Rust bootstrap (10000) → Lab Report
Public server: backtest → bootstrap (1000) → percentiles → calibration → drift check → Hospital
In the Lab, the seed, input fingerprint, method, trade count, path count and histogram are attached to the same Lab Report. On the site, public snapshots remain available in the Facts JSON. Both outputs are educational and trigger no live order.
Monte Carlo FAQ
- Why 10000 simulations in the Lab and 1000 on the public server?
- They are separate protocols. The server keeps its historical snapshots at 1000 simulations. The local Rust engine uses 10000 paths to reduce numerical noise in the current diagnostic. Raising N reduces neither market risk nor dataset bias.
- Why is there no result below 10 trades?
- The engine refuses false precision: it returns insufficient_trades, runs zero paths and shows no profit, loss or ruin probability.
- Does MC replace holdout or paper trading?
- No. Holdout measures a chronological period kept unseen; Monte Carlo reorders observations already known; paper trading then tests execution and drift. These evidence layers answer different questions.
- What does “risk of ruin” mean in version 1?
- It is the observed frequency of resampled paths crossing a 50% loss from initial equity. The threshold is explicit, but it is neither a prediction nor a universal risk definition.
Quick MC glossary
| Term | Role | Link |
|---|---|---|
| Bootstrap | Resampling trades with replacement | /facts/monte-carlo |
| p5 / p95 | Simulated PnL distribution tails | monte-carlo.json |
| Robustness score | Stability under perturbations | /live-results |
| Walk-forward | Temporal split anti look-ahead | /backtest |
| Drift | Backtest vs paper gap | /dashboard |
Explicit limits
- Crypto / perps: fat tails — classical bootstrap may understate extremes.
- Regime shifts: a historical MC pass does not guarantee the future regime.
- Block bootstrap never creates a regime, liquidity condition or shock absent from the measured history.
- Monte Carlo remains a post-selection diagnostic; it neither selects nor promotes the candidate.
- Educational content; no profit promise.
Run the diagnostic on your strategy
Import a compatible Pine strategy, replay a World Arena contract or generate an ArenaScript candidate. The Lab produces holdout evidence first, then attaches the Rust Monte Carlo receipt to the local report when at least 10 trades are available.