How to evaluate a strategy
Win rate and profit factor are the two numbers this industry quotes most, and the two that tell you least. Either can be pushed wherever a seller wants it by choosing the stops, the exits, the test window and the start date, which is why they lead almost every sales page. The sections below explain what each common measure actually says and, more usefully, how each one gets inflated or misread. Nothing here is a pitch. The point is to make you harder to sell to, by us or by anyone else.
Contents
Win rate
- What it measures
- The fraction of closed trades that made money. Nothing else. It carries no information about how much a win pays or how much a loss costs, so a system can win nine trades out of ten and still lose money overall.
- What good looks like
- Anything from 30% to 70% can belong to a sound system. Trend-following approaches often win under 40% because they take many small losses while they wait for a large move. Mean-reversion approaches can win above 70% while carrying occasional large losses. The figure only means something when it sits next to the average win and the average loss.
- How it gets gamed
- This is the easiest number on the page to inflate. Widen the stop and cut winners early: the win rate climbs at once, because trades that would have been small losses get held until they recover, and gains are banked before they can reverse. The bill arrives later, as one rare and very large loss that gives back months of small wins. When a vendor leads with a high win rate, ask for the average loss and the largest single loss. Refusal to show those two numbers is the answer.
Profit factor
- What it measures
- Take every dollar the system won and divide it by every dollar it lost. A profit factor of 2 means two dollars gained for each dollar given back over the whole test. The entire history gets compressed into one ratio, and the compression is the weakness, because the ratio says nothing about when the losses arrived or how deep the equity curve sank on the way.
- What good looks like
- Above 1 the system made money over the period tested. Below 1 it did not. Live systems that pay real commissions and slippage tend to sit somewhere between about 1.2 and 2. Treat anything above 3 on a meaningful trade count with suspicion, because ratios like that in liquid futures markets usually mean a short favourable window or a fitted result.
- How it gets gamed
- The quickest inflation is the window. Start the test after a losing stretch and end it after a winning one, and the ratio improves without the strategy changing at all. Leaving out commissions and slippage helps too, since costs come straight out of the gross profit. A high profit factor over a small sample is noise, because thirty trades can produce almost any ratio by chance. Ask for the trade count and the exact date range, and confirm costs are in the figures.
Expectancy and payoff ratio
- What it measures
- What one trade is worth on average, wins and losses netted together. Multiply the win rate by the average win, subtract the loss rate multiplied by the average loss, and the result is expectancy. Payoff ratio is one of its two moving parts, the average win divided by the average loss. Between them they expose the trade-off that win rate hides.
- What good looks like
- Positive, after costs, is the floor. A useful second check is expectancy divided by the average loss, which tells you what fraction of a typical loss each trade earns back. On payoff ratio, high win rates normally pair with low payoffs and low win rates with high ones, and neither pairing is better in itself. What matters is that the combination nets clearly positive.
- How it gets gamed
- Expectancy is harder to fake than win rate, which is exactly why fewer vendors quote it. The remaining tricks are familiar. Quote it before commissions and slippage, so a thin per-trade edge that costs would erase looks real. Or quote it from a stretch that leaves out the losing regime. It is also an average, so one enormous win can carry a sample that otherwise loses. Ask whether the number still holds with the single best trade removed. A real edge does. A lucky one does not.
Maximum drawdown and drawdown duration
- What it measures
- Maximum drawdown is the deepest peak-to-trough fall in the equity curve, in currency or as a percentage. Drawdown duration is the time the curve spent below its previous high before making a new one. Depth is the worst moment. Duration is how long the worst moment lasted.
- What good looks like
- Judge depth against return, since a 10% worst drawdown means something different on a system earning 30% a year than on one earning 8%. Then look harder at duration, because it is the number that decides whether a human being keeps running the system. Most traders can absorb one sharp loss. Very few can watch an account go nowhere for months, and the ones who cannot usually switch the system off at the exact bottom.
- How it gets gamed
- A backtested drawdown is a floor, not a ceiling. History only contains the bad periods that happened to occur, and the worst drawdown of a live system is routinely deeper than anything in its test. Quoting drawdown from closed-trade equity is another softener, since it hides how far open positions sank before they were closed. Duration is simply left out, almost universally, because a months-long flat stretch sells worse than a tidy depth figure. If duration is missing, assume it is missing on purpose.
Alpha, beta and R squared
- What it measures
- Three numbers out of one regression, the system against the market it trades. Beta is the slope: how far the system moves when the market moves. Alpha is the intercept: the return left over once the market has been paid for its share, usually quoted annualised. R squared is how much of the variance in the returns the market explains at all. Between them they answer the thing almost nobody asks a vendor, which is whether this is doing something of its own or is long exposure with extra steps.
- What good looks like
- A strategy that claims an edge independent of market direction should show beta near zero, a low R squared, and an alpha that is positive after costs. Low R squared says the returns come from the strategy's own decisions rather than from the index it happens to trade on. Beta near 1 with a high R squared says you are paying a subscription for exposure an index future would give you for the cost of commissions, and whatever alpha is left has to cover the subscription before it has done anything for you. Read alpha last, because it only means something once the other two say it is measuring the strategy rather than the market.
- How it gets gamed
- The common abuse is silence. A system that is mostly repackaged market exposure looks like skill for as long as the market rises, so the regression is simply never run. You can run it yourself from a trade log in a spreadsheet in an afternoon. Where the numbers are published, the benchmark is the first thing to check, because regressing a stock index strategy against an unrelated market makes beta look flatteringly low and alpha correspondingly high. The second is frequency: alpha computed on daily returns, weekly and monthly should land in the same neighbourhood, and a figure that only survives at one of the three is an artefact of the window rather than a property of the system. Ask what the returns were regressed against, over what period, and at what frequency. A blank look is also an answer.
Correlation between components
- What it measures
- When a system is built from several models or signals, pairwise correlation measures how similarly the components behave. Near +1, they win and lose together. Near zero, each one contributes something the others do not, and the combined equity curve comes out smoother than any single part on its own.
- What good looks like
- The lower the better, and it has to hold over the full history, not a chosen window. Genuinely low correlation between components is one of the hardest numbers in a report to fake, because it has to survive thousands of trades through every kind of market. High correlation means the diversification is cosmetic. One idea wearing several hats, and every hat comes off in the same storm.
- How it gets gamed
- A convenient window is the first trick, so ask for the figure over the entire test. Padding is the second: components that barely trade drag the average down without adding anything real. The label itself can lie, since seven independent models that all read the same moving average with different settings will show their kinship the moment anyone measures it. If a vendor claims diversification, the pairwise correlation matrix is the receipt. Ask for it.
Out-of-sample and walk-forward testing
- What it measures
- Not a metric. A method, and the most important thing on this page. In-sample data is the history a strategy's rules were fitted to, and performance there is close to meaningless, because with enough parameters any rule set can be made to fit any past. Out-of-sample data is history the rules never saw during development. Walk-forward testing repeats the split in rolling windows: fit on one stretch, test on the untouched stretch that follows, step forward, do it again. The stitched-together out-of-sample results are the only backtest numbers worth reading.
- What good looks like
- Out-of-sample results should sit within shouting distance of the in-sample ones. Some decay is normal. A collapse is a verdict. Coverage matters as much as length, so a walk-forward that spans a crash, a grinding bear, a melt-up and a quiet drift says far more than a longer test of one kind of weather.
- How it gets gamed
- This is the concept every other number on the page hides behind, because any metric computed on fitted data inherits the fit. The subtler abuse is repetition: run twenty variants, keep the one with the prettiest out-of-sample stretch, and the out-of-sample has quietly become in-sample. Some vendors also call a test walk-forward when the parameters were tuned once over the entire dataset, which is an ordinary curve fit with a better name. Ask how many variants were tried before this one, and ask which data the final rules never touched. The honest answer to the first question is rarely one.
Confidence intervals
- What it measures
- A backtest produces one number from one sample of history. A confidence interval turns that point estimate into the range where the true value can reasonably sit. Bootstrap methods build the range by resampling the actual trades many thousands of times and recomputing the statistic each time.
- What good looks like
- Narrow beats wide, and wide beats absent. The bottom of the range is what to read first, because a Sharpe whose interval reaches below zero belongs to a system whose entire edge might be luck. More trades and more years tighten the range, which is one more reason a short glossy backtest deserves suspicion.
- How it gets gamed
- Omission is the main trick, since a single confident number sells better than an honest range. Where an interval does appear, check the sample behind it, because one built on fifty trades is wide enough to drive a truck through, and quoting only its optimistic end is the same sleight of hand as quoting the point estimate alone. Ask for the interval on any headline figure. A vendor who has never computed one has told you how carefully everything else was done.
Distribution shape
- What it measures
- Averages hide shape. Skewness measures which tail the returns lean toward: negative skew is many small wins against rare large losses, positive skew is the reverse. Kurtosis measures how fat the tails are, in other words how often extreme results turn up compared with a normal curve. MFE and MAE, maximum favourable and maximum adverse excursion, record how far each trade moved for and against the position while it was open.
- What good looks like
- Neither skew is disqualifying, but you should know which one you own. Negative skew feels comfortable right up to the tail event. Positive skew feels like slow bleeding until the payoff arrives, and many people abandon it first. MFE set against the actual exits shows whether the exits capture what the trades offered, and MAE set against the stop shows whether the stop sits where trades genuinely fail or somewhere arbitrary.
- How it gets gamed
- Negative skew is the classic subscription product, because months of small steady gains build a track record and a customer base before the loss that was always coming. Kurtosis goes unquoted for the same reason. MFE and MAE can be gamed in a backtest as well: an exit tuned until every historical trade left near its peak excursion is curve fitting in its purest form, and it will not repeat. Ask what the worst single day in the test was, and how often the shape of the distribution says to expect another one.
Then check ours
Then apply the same test to us. The report covers the full walk forward study and the live forward test as two separate records, labelled as such and never blended, and it goes to approved applicants before any payment is taken rather than after.
