Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Monte Carlo trade-order shuffling shows how the same completed trades could have produced different equity paths in a different sequence. It can expose a historically fortunate order that understated drawdown or time under water—but it does not prove a strategy is false, predict its future, or by itself test whether the strategy has an edge.

What does trade-order shuffling test?

Start with a chronological list of completed trades and randomly rearrange that list many times. For each arrangement, rebuild equity from the same initial balance and recalculate path-dependent measures such as maximum drawdown, recovery duration, or whether equity crosses a specified threshold. The original backtest is one historical trajectory; the shuffled paths show how sensitive that trajectory was to the order of its observed outcomes.

In the simplest case, each trade has a fixed net profit or loss, position sizes and costs do not change with equity, and trades do not overlap. Every shuffle then contains exactly the same outcomes. Total additive P&L is unchanged, but the path can change substantially: a run of losses may arrive before gains rather than after them. Jesse’s documentation describes this sequence-shuffling workflow and recommends at least 1,000 scenarios; that is a software recommendation, not a universal statistical threshold or proof of adequacy. Jesse: Trade-Order Shuffling

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Here, “falsifies your equity curve” means stress-testing whether its favorable historical shape depends on its particular ordering. It does not mean disproving the strategy, showing that the recorded trades are invalid, or forecasting a future loss.

#1 Best Overall
Sale
How to Day Trade for a Living: A Beginner’s Guide to Trading Tools and Tactics, Money Management, Discipline and Trading Psychology (Stock Market Trading and Investing)
  • As a day trader, you can live and work anywhere in the world. You can decide when to work and when not to work.
  • You only answer to yourself. That is the life of the successful day trader. Many people aspire to it, but very few succeed. Day trading is not gambling or an online poker game.
  • To be successful at day trading you need the right tools and you need to be motivated, to work hard, and to persevere.

What must be fixed before you shuffle?

Define the question and the trade data

For a sequencing diagnostic, use a clean, chronological vector of completed trade-level net P&L. Decide in advance how fees and slippage are treated, and use that treatment consistently. If trade size or another attribute is integral to an outcome, keep the outcome paired with that attribute; do not independently shuffle fields and create combinations that never occurred.

Then specify what the simulation holds fixed: initial balance, the set of trades, sizing rules, costs, and any threshold used to represent a margin call or ruin. If the question is “does my equity curve depend on trade order?”, this conditional setup answers it for the observed set of trades—not for trades the strategy did not take.

Check whether the additive-trade model fits

Adding fixed trade P&Ls to an initial balance is a useful, explicit simplification. It may not reproduce portfolio behavior when position sizes are a fraction of current equity, positions overlap, trade durations affect exposure, or margin, stops, and liquidation depend on the path. In those cases, a list of independent closed-trade P&Ls may discard relevant mechanics. Model the trade- or bar-level state that drives those mechanics, or label the simpler calculation as a sequence stress test rather than a full portfolio simulation. The equity-reconstruction and peak-to-valley approach is also illustrated in the MQL5 discussion of Monte Carlo testing. MQL5: Statistical Robustness Testing on MQL5 CSV Exports

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you shuffle trades and rebuild equity in Python?

This NumPy example calculates absolute and percentage maximum drawdown and the longest run of consecutive trade steps below the previous equity peak. It assumes fixed additive P&L and positive equity peaks. Recovery duration is measured in trade steps, not calendar time. A run still underwater at the end is counted only through the end of the sample; it is not evidence of how long recovery would ultimately take.

import numpy as np


def path_metrics(trade_pnl, initial_equity=10_000.0, threshold=None):
    equity = initial_equity + np.r_[0.0, np.cumsum(trade_pnl)]
    peaks = np.maximum.accumulate(equity)
    drawdown = peaks - equity

    if np.any(peaks <= 0):
        raise ValueError("Percentage drawdown requires positive equity peaks")

    underwater = 0
    max_underwater = 0
    for current, peak in zip(equity[1:], peaks[1:]):
        if current < peak:
            underwater += 1
            max_underwater = max(max_underwater, underwater)
        else:
            underwater = 0

    return {
        "ending_equity": equity[-1],
        "max_drawdown": float(np.max(drawdown)),
        "max_drawdown_pct": float(np.max(drawdown / peaks)),
        "max_underwater_steps": max_underwater,
        "threshold_breached": (
            bool(np.any(equity <= threshold)) if threshold is not None else None
        ),
    }


def shuffled_scenarios(trade_pnl, initial_equity=10_000.0,
                       n_sims=10_000, seed=7, threshold=None):
    trade_pnl = np.asarray(trade_pnl, dtype=float)
    if trade_pnl.ndim != 1 or trade_pnl.size == 0:
        raise ValueError("trade_pnl must be a non-empty one-dimensional vector")
    if n_sims < 1:
        raise ValueError("n_sims must be at least 1")

    rng = np.random.default_rng(seed)
    observed = path_metrics(trade_pnl, initial_equity, threshold)
    scenarios = [
        path_metrics(rng.permutation(trade_pnl), initial_equity, threshold)
        for _ in range(n_sims)
    ]
    return observed, scenarios


# Supply chronological, net trade P&L values in account currency.
# observed, scenarios = shuffled_scenarios(trade_pnl, n_sims=10_000, seed=7)
# mdd = np.array([s["max_drawdown"] for s in scenarios])
# print(observed["max_drawdown"], np.quantile(mdd, [0.50, 0.90, 0.95]))

The initial balance is explicitly included as the first equity point, so drawdown is measured from the initial peak as well as from later peaks. The optional threshold is only a level-crossing flag; it is not a margin model. To estimate a margin-call probability, the simulated mechanics must actually represent the account’s margin rules, open exposure, and liquidation behavior.

Jesse’s research API provides another Python-oriented example: its documentation shows num_scenarios=1000 and reports original and scenario metrics such as return, drawdown, volatility, Sharpe, and Calmar. Those are API details, not a claim that every listed metric changes under an order shuffle. Jesse: Monte Carlo Analysis

Rank #3
Trading: Technical Analysis Masterclass: Master the financial markets
  • Language: english
  • Book - trading: technical analysis masterclass: master the financial markets
  • It is made up of premium quality material.

How should you read the simulated drawdowns?

For a metric where larger values are worse, compare the observed statistic with the shuffled distribution. Report the observed value, the simulated median and selected tail percentiles, plus the number of simulations and random seed. For example, an observed drawdown near the low end of the shuffled distribution means the historical order was unusually kind relative to those rearrangements. A high upper tail indicates that the same set of trades can form materially harsher paths under other orders.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These percentiles are conditional on the observed trade sample. Shuffling cannot invent a larger loss, add a new market regime, or create future slippage that is absent from the inputs. More simulations can reduce Monte Carlo noise in estimated percentiles, but they cannot fix an unrepresentative sample, an invalid model, or strategy overfitting.

At 1,000 scenarios, tail estimates are coarse: only about 50 sampled paths, on average, fall above the 95th percentile threshold. Jesse recommends at least 1,000 scenarios in its documentation, but that count should not be treated as a universal adequacy test. Choose the number in light of the tail you want to estimate, and report it transparently.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why isn’t this automatically a permutation test for edge?

A trade-order shuffle changes sequence while retaining the observed results. It is therefore a sequencing diagnostic, not a general null model for whether a strategy has predictive edge. A valid edge test needs a randomization that removes the directional effect under a defensible null, followed by recalculation of the chosen statistic. Random sign flips are one possible construction for some settings, not a universally correct null for every strategy.

Metric choice matters. If a statistic is order-independent for a fixed return vector, shuffling the order cannot create a meaningful distribution for that statistic. For instance, the Sortino ratio based on the same fixed return observations remains unchanged by reordering. As Ushana Kevin Iorkumbul puts it, “This means that shuffling the order of the returns does not change the Sortino Ratio at all, and a permutation test built on order-shuffling would produce a constant null distribution that tests nothing.” MQL5 article, 30 July 2026

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep these methods distinct when choosing what to run:

Method What is randomized Question it can address What it does not establish
Trade-order shuffle Order of the observed trades; the set of outcomes is held fixed. How sensitive are sequence-dependent equity paths and drawdowns to ordering? Whether trade selection has edge or whether future market conditions will resemble the sample.
Sign or label randomization Trade signs or labels according to a stated null. Whether a chosen statistic is unusual under that specific no-effect model. A valid null for every strategy; the randomization must match the hypothesis.
Bootstrap Trades or observations sampled with replacement. How sample composition affects uncertainty or metric stability. Pure sequencing sensitivity, since the resampled set can differ from the original.
Market-data or candle perturbation Market paths or input data, followed by rerunning the strategy. Sensitivity to altered market conditions. The same conditional question as merely rearranging completed trades.

When can you report a Monte Carlo p-value?

A descriptive trade-order stress distribution does not automatically produce a hypothesis-test p-value. Use a p-value only when the randomization corresponds to a justified null and the test statistic, tail direction, and comparison rule are specified. State whether the test is one-sided or two-sided, what counts as “extreme,” whether the original arrangement is included, and how ties are handled.

When estimating a permutation p-value from b exceedances among m randomly drawn permutations, do not report zero just because no sampled result was as extreme as the observed result. The finite-sample correction commonly written as (b + 1) / (m + 1) avoids that zero estimate under the usual sampled-permutation convention. Phipson and Smyth explain why random-permutation p-values should not be understated this way. Phipson and Smyth, “Permutation P-values Should Never Be Zero”

What should you pair with a trade-order stress test?

Use the shuffle as one diagnostic, not as a substitute for evaluating the strategy’s data and assumptions. Check performance on untouched out-of-sample data or through walk-forward evaluation; include realistic transaction costs; account for survivorship in the data; and consider how many strategy variants were tried before selecting this one. The cited implementations document shuffling and scenario comparison, not a guarantee of future profitability.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.