How to Evaluate Live Algorithm Performance
A due-diligence framework for labels, samples, drawdowns, trade distributions, regime exposure, and execution quality.
Verify what the word live actually means
A live label should identify the source, account scope, observation period, update process, and whether values reflect actual broker fills. It should not be inferred from real-time signals, a forward-test chart, or a frequently updated backtest. Ask how commissions, fees, slippage, deposits, withdrawals, and strategy changes are handled before comparing returns.
Create a provenance record from the published methodology. Determine whether the page reports a model, simulated account, vendor account, customer accounts, or aggregated signals. Note who controls the account and whether records can be reconciled to order or broker data. If those facts are unavailable, downgrade confidence rather than filling gaps with assumptions.
Mixed histories need clear segmentation. A strategy can have a long hypothetical backtest and a short live period, but the two should not be joined without visible labels. Use the live segment to study observed execution and behavior; use the backtest to examine broader historical hypotheses. Neither segment guarantees future performance.
| Label | What it may show | What still needs verification |
|---|---|---|
| Hypothetical/backtest | Rules applied to historical data | Biases, costs, fill model, and future robustness |
| Simulation/forward test | Signals under ongoing market data | Real capital constraints and achievable fills |
| Live broker result | Observed orders and fills in an account | Completeness, changes, capital, and representativeness |
Make capital and returns comparable
Total profit is not comparable without capital, contract size, deposits, withdrawals, and exposure. Reconstruct a net return series using a declared method and current contract specifications. Separate strategy return from cash flows and avoid annualizing a short or unusually favorable sample. Compare both absolute loss and return on the capital actually at risk.
Futures leverage complicates denominator choices. Account balance, required margin, notional exposure, and intended risk are different quantities. Margin is a performance bond and can change; it is not maximum loss. State which capital measure a percentage uses, and do not compare percentages built from different denominators.
Costs must follow each trade. Include commission, exchange and regulatory fees where applicable, and actual or conservatively modeled slippage. If a public result is gross, do not silently compare it with a net alternative. Estimate a range and retain the calculation. A high-turnover strategy can be especially sensitive to small cost changes.
- Reconcile deposits, withdrawals, and transfers.
- Use correct contract multipliers and quantities.
- State whether returns use starting, average, or marked capital.
- Include recurring and transaction costs consistently.
- Avoid annualized claims from an inadequate sample.
Inspect the full trade distribution
Trade count alone does not determine sample quality. Inspect wins, losses, tails, holding periods, market exposure, clustering, and contribution by a few trades. Calculate average and median outcomes, payoff asymmetry, profit factor where meaningful, and the result after removing the largest winners. A stable-looking equity curve can hide concentrated risk.
Win rate needs average win, average loss, and exit behavior. A high win rate can coexist with rare severe losses; a lower win rate can accompany positive asymmetry. Review maximum favorable and adverse movement when available, but remember those are sample statistics rather than future bounds.
Plot outcomes by time, market, direction, session, and strategy version. Look for serial dependence and bursts of activity that reduce the effective sample. If many trades are near-identical or occur in one regime, the nominal count overstates independent evidence. Confidence should reflect diversity of observations, not a round number.
| Metric | Pair it with | Common error |
|---|---|---|
| Win rate | Average win/loss and tail losses | Equating frequency with edge |
| Profit factor | Trade count, costs, and stability | Ignoring a few dominant trades |
| Average trade | Median, distribution, and slippage | Assuming the mean is typical |
| Trade count | Regimes and independence | Treating clustered trades as independent |
Analyze drawdown depth, duration, and mechanism
Maximum drawdown is one realized path, not a guaranteed worst case. Review drawdown depth, duration, recovery time, frequency, and the trades or exposures that caused it. Compare closed and open equity where available. Stress larger and longer losses, because future sequencing, gaps, dependence, and execution can exceed the observed record.
Plot the underwater curve rather than reading only the maximum value. Two strategies with the same maximum drawdown can differ materially if one recovered quickly while the other remained below its peak for months. Long stagnation affects capital opportunity, operator behavior, and whether a strategy remains within external account rules.
Identify the mechanism of severe losses. Were they concentrated in a volatility shock, repeated mean-reversion entries, an overnight gap, a contract roll, or execution failure? A causal hypothesis helps design a relevant stress. It does not justify deleting the period or assuming the next event will look identical.
Observed drawdown is not a loss limit
Historical maximum drawdown cannot cap future loss. Position sizing, independent account limits, and operational controls remain necessary.
Test regimes, parameters, and selection effects
Split results across meaningful periods without repeatedly searching for flattering partitions. Examine volatility, trend, session, direction, and market environments relevant to the strategy. Shift start dates, vary nearby parameters, widen costs, and reserve untouched data. Robustness means the thesis remains plausible across reasonable changes, not that every period is profitable.
Ask when the strategy should struggle. A credible mandate includes unfavorable conditions and an explanation of how risk is controlled. If no losing regime can be identified, the description may be too vague to test. Compare that expectation with actual drawdowns and avoid rewriting the mandate after observing results.
Selection is a major source of overconfidence. A catalog may show surviving or currently promoted systems while failed candidates disappear. Ask how additions, retirements, and revisions are recorded. Evaluate the selection process and preserve the full candidate set where possible, not just the best result chosen with hindsight.
- 1
State the thesis
Describe expected favorable and unfavorable conditions before splitting the record.
- 2
Partition once
Use economically relevant periods and retain all results, including failures.
- 3
Perturb assumptions
Vary dates, costs, fills, and nearby parameters without searching endlessly.
- 4
Hold out evidence
Reserve data or future observation that was not used to select and tune the system.
Turn due diligence into a sized, reviewable decision
Conclude with a confidence grade, unresolved risks, allocation ceiling, pilot plan, monitoring metrics, and pause criteria. Rejection or continued observation is a valid decision. If deploying, use the smallest exposure consistent with the plan and reconcile actual signals, orders, fills, costs, and positions before considering any increase.
Do not convert historical metrics directly into expected income. The CFTC cautions against claims that trading bots or systems can guarantee returns, and NFA rules explain limitations of hypothetical performance. Treat every historical record as uncertain evidence. Size from loss tolerance and operational capacity, not from a target payout.
Set review triggers for strategy revisions, unexplained divergence, cost drift, drawdown, regime change, and data-quality problems. Compare live behavior with the frozen thesis, but allow for statistical noise. Changes should be versioned and re-evaluated; quietly replacing settings destroys the continuity that made the record useful.
- Evidence type and provenance are documented.
- Returns are net and capital assumptions are explicit.
- Trade and drawdown distributions have been inspected.
- Out-of-sample and sensitivity results are retained.
- Pilot, monitoring, pause, and retirement rules are approved.
Sources and methodology
HexTrade Research uses official product, exchange, regulator, and vendor documentation. Policies and platform behavior can change; follow the linked source and verify current terms before trading.
- 1.AI trading bots advisory — CFTC, accessed Aug 30, 2026
- 2.Commodity trading systems sold on the internet — CFTC, accessed Aug 30, 2026
- 3.The economic purpose of futures markets and how they work — CFTC, accessed Aug 30, 2026
- 4.NFA hypothetical performance results requirements — National Futures Association, accessed Aug 30, 2026
- 5.Position and risk management — CME Group, accessed Aug 30, 2026
- 6.Margin: know what is needed — CME Group, accessed Aug 30, 2026
- 7.HexTrade algorithms — HexTrade Docs, accessed Aug 30, 2026
- 8.Drawdown — HexTrade Docs, accessed Aug 30, 2026
Frequently asked questions
How long must a live track record be?
There is no universal duration. Adequacy depends on trade frequency, independence, regimes, strategy changes, and intended use. A calendar year with clustered trades may contain less evidence than a shorter but more varied sample.
What is the most important performance metric?
No single metric is sufficient. Start with evidence provenance, then evaluate net returns, trade distribution, drawdown, exposure, costs, regime behavior, and execution together against the strategy mandate.
Is live performance always better evidence than a backtest?
It provides actual account and execution observations when properly documented, but it may be short, selected, or affected by strategy changes. Backtests cover more history but introduce hypothetical limitations. Use each for its appropriate question.
Can the maximum historical drawdown guide sizing?
It can inform stress design but should not be used as a guaranteed ceiling. Apply larger and longer adverse scenarios, account limits, leverage constraints, and an independent loss budget.
Should I stop an algorithm after a losing streak?
Follow predeclared pause criteria. Investigate control failures, thesis violations, and statistically unusual behavior without reacting mechanically to every loss. If orders or positions cannot be reconciled, pause regardless of performance.
Next step
Put the research into a controlled workflow
Start small, verify the broker and account rules, and keep risk controls between every signal and live order.
Review live algorithm pagesContinue reading
Educational content only. Futures are leveraged products and can produce losses greater than the amount you expected to risk. This article is not financial, legal, or prop-firm compliance advice.