Algorithm evaluation7 min readPublished Jul 8, 2026Updated Aug 30, 2026

How to Evaluate Live Algorithm Performance

A due-diligence framework for labels, samples, drawdowns, trade distributions, regime exposure, and execution quality.

By HexTrade1,386 wordsSources reviewed Aug 30, 2026
A live algorithm record decomposed into provenance, trades, drawdowns, regimes, and execution differences

Verify what the word live actually means

A live label should identify the source, account scope, observation period, update process, and whether values reflect actual broker fills. It should not be inferred from real-time signals, a forward-test chart, or a frequently updated backtest. Ask how commissions, fees, slippage, deposits, withdrawals, and strategy changes are handled before comparing returns.

Create a provenance record from the published methodology. Determine whether the page reports a model, simulated account, vendor account, customer accounts, or aggregated signals. Note who controls the account and whether records can be reconciled to order or broker data. If those facts are unavailable, downgrade confidence rather than filling gaps with assumptions.

Mixed histories need clear segmentation. A strategy can have a long hypothetical backtest and a short live period, but the two should not be joined without visible labels. Use the live segment to study observed execution and behavior; use the backtest to examine broader historical hypotheses. Neither segment guarantees future performance.

Evidence labels answer different questions
LabelWhat it may showWhat still needs verification
Hypothetical/backtestRules applied to historical dataBiases, costs, fill model, and future robustness
Simulation/forward testSignals under ongoing market dataReal capital constraints and achievable fills
Live broker resultObserved orders and fills in an accountCompleteness, changes, capital, and representativeness
The evidence hierarchyProvenanceIdentify result type, account scope, source, an…Comparable returnsNormalize capital, contracts, fees, and withdra…Trade distributionInspect frequency, skew, concentration, and out…Drawdown pathMeasure depth, duration, recovery, and overlapp…Regime behaviorTest performance across periods and market cond…Execution realityCompare intended orders with actual acknowledge…
The evidence hierarchy. Performance interpretation becomes weaker when provenance, costs, or strategy changes cannot be verified.

Make capital and returns comparable

Total profit is not comparable without capital, contract size, deposits, withdrawals, and exposure. Reconstruct a net return series using a declared method and current contract specifications. Separate strategy return from cash flows and avoid annualizing a short or unusually favorable sample. Compare both absolute loss and return on the capital actually at risk.

Futures leverage complicates denominator choices. Account balance, required margin, notional exposure, and intended risk are different quantities. Margin is a performance bond and can change; it is not maximum loss. State which capital measure a percentage uses, and do not compare percentages built from different denominators.

Costs must follow each trade. Include commission, exchange and regulatory fees where applicable, and actual or conservatively modeled slippage. If a public result is gross, do not silently compare it with a net alternative. Estimate a range and retain the calculation. A high-turnover strategy can be especially sensitive to small cost changes.

  • Reconcile deposits, withdrawals, and transfers.
  • Use correct contract multipliers and quantities.
  • State whether returns use starting, average, or marked capital.
  • Include recurring and transaction costs consistently.
  • Avoid annualized claims from an inadequate sample.

Inspect the full trade distribution

Trade count alone does not determine sample quality. Inspect wins, losses, tails, holding periods, market exposure, clustering, and contribution by a few trades. Calculate average and median outcomes, payoff asymmetry, profit factor where meaningful, and the result after removing the largest winners. A stable-looking equity curve can hide concentrated risk.

Win rate needs average win, average loss, and exit behavior. A high win rate can coexist with rare severe losses; a lower win rate can accompany positive asymmetry. Review maximum favorable and adverse movement when available, but remember those are sample statistics rather than future bounds.

Plot outcomes by time, market, direction, session, and strategy version. Look for serial dependence and bursts of activity that reduce the effective sample. If many trades are near-identical or occur in one regime, the nominal count overstates independent evidence. Confidence should reflect diversity of observations, not a round number.

Metrics that need context
MetricPair it withCommon error
Win rateAverage win/loss and tail lossesEquating frequency with edge
Profit factorTrade count, costs, and stabilityIgnoring a few dominant trades
Average tradeMedian, distribution, and slippageAssuming the mean is typical
Trade countRegimes and independenceTreating clustered trades as independent

Analyze drawdown depth, duration, and mechanism

Maximum drawdown is one realized path, not a guaranteed worst case. Review drawdown depth, duration, recovery time, frequency, and the trades or exposures that caused it. Compare closed and open equity where available. Stress larger and longer losses, because future sequencing, gaps, dependence, and execution can exceed the observed record.

Plot the underwater curve rather than reading only the maximum value. Two strategies with the same maximum drawdown can differ materially if one recovered quickly while the other remained below its peak for months. Long stagnation affects capital opportunity, operator behavior, and whether a strategy remains within external account rules.

Identify the mechanism of severe losses. Were they concentrated in a volatility shock, repeated mean-reversion entries, an overnight gap, a contract roll, or execution failure? A causal hypothesis helps design a relevant stress. It does not justify deleting the period or assuming the next event will look identical.

Observed drawdown is not a loss limit

Historical maximum drawdown cannot cap future loss. Position sizing, independent account limits, and operational controls remain necessary.

Test regimes, parameters, and selection effects

Split results across meaningful periods without repeatedly searching for flattering partitions. Examine volatility, trend, session, direction, and market environments relevant to the strategy. Shift start dates, vary nearby parameters, widen costs, and reserve untouched data. Robustness means the thesis remains plausible across reasonable changes, not that every period is profitable.

Ask when the strategy should struggle. A credible mandate includes unfavorable conditions and an explanation of how risk is controlled. If no losing regime can be identified, the description may be too vague to test. Compare that expectation with actual drawdowns and avoid rewriting the mandate after observing results.

Selection is a major source of overconfidence. A catalog may show surviving or currently promoted systems while failed candidates disappear. Ask how additions, retirements, and revisions are recorded. Evaluate the selection process and preserve the full candidate set where possible, not just the best result chosen with hindsight.

  1. 1

    State the thesis

    Describe expected favorable and unfavorable conditions before splitting the record.

  2. 2

    Partition once

    Use economically relevant periods and retain all results, including failures.

  3. 3

    Perturb assumptions

    Vary dates, costs, fills, and nearby parameters without searching endlessly.

  4. 4

    Hold out evidence

    Reserve data or future observation that was not used to select and tune the system.

Turn due diligence into a sized, reviewable decision

Conclude with a confidence grade, unresolved risks, allocation ceiling, pilot plan, monitoring metrics, and pause criteria. Rejection or continued observation is a valid decision. If deploying, use the smallest exposure consistent with the plan and reconcile actual signals, orders, fills, costs, and positions before considering any increase.

Do not convert historical metrics directly into expected income. The CFTC cautions against claims that trading bots or systems can guarantee returns, and NFA rules explain limitations of hypothetical performance. Treat every historical record as uncertain evidence. Size from loss tolerance and operational capacity, not from a target payout.

Set review triggers for strategy revisions, unexplained divergence, cost drift, drawdown, regime change, and data-quality problems. Compare live behavior with the frozen thesis, but allow for statistical noise. Changes should be versioned and re-evaluated; quietly replacing settings destroys the continuity that made the record useful.

  • Evidence type and provenance are documented.
  • Returns are net and capital assumptions are explicit.
  • Trade and drawdown distributions have been inspected.
  • Out-of-sample and sensitivity results are retained.
  • Pilot, monitoring, pause, and retirement rules are approved.

Sources and methodology

HexTrade Research uses official product, exchange, regulator, and vendor documentation. Policies and platform behavior can change; follow the linked source and verify current terms before trading.

  1. 1.AI trading bots advisory CFTC, accessed Aug 30, 2026
  2. 2.Commodity trading systems sold on the internet CFTC, accessed Aug 30, 2026
  3. 3.The economic purpose of futures markets and how they work CFTC, accessed Aug 30, 2026
  4. 4.NFA hypothetical performance results requirements National Futures Association, accessed Aug 30, 2026
  5. 5.Position and risk management CME Group, accessed Aug 30, 2026
  6. 6.Margin: know what is needed CME Group, accessed Aug 30, 2026
  7. 7.HexTrade algorithms HexTrade Docs, accessed Aug 30, 2026
  8. 8.Drawdown HexTrade Docs, accessed Aug 30, 2026

Frequently asked questions

How long must a live track record be?

There is no universal duration. Adequacy depends on trade frequency, independence, regimes, strategy changes, and intended use. A calendar year with clustered trades may contain less evidence than a shorter but more varied sample.

What is the most important performance metric?

No single metric is sufficient. Start with evidence provenance, then evaluate net returns, trade distribution, drawdown, exposure, costs, regime behavior, and execution together against the strategy mandate.

Is live performance always better evidence than a backtest?

It provides actual account and execution observations when properly documented, but it may be short, selected, or affected by strategy changes. Backtests cover more history but introduce hypothetical limitations. Use each for its appropriate question.

Can the maximum historical drawdown guide sizing?

It can inform stress design but should not be used as a guaranteed ceiling. Apply larger and longer adverse scenarios, account limits, leverage constraints, and an independent loss budget.

Should I stop an algorithm after a losing streak?

Follow predeclared pause criteria. Investigate control failures, thesis violations, and statistically unusual behavior without reacting mechanically to every loss. If orders or positions cannot be reconciled, pause regardless of performance.

Next step

Put the research into a controlled workflow

Start small, verify the broker and account rules, and keep risk controls between every signal and live order.

Review live algorithm pages

Continue reading

Educational content only. Futures are leveraged products and can produce losses greater than the amount you expected to risk. This article is not financial, legal, or prop-firm compliance advice.