Silex Strategies

How to evaluate a strategy

Win rate and profit factor are the two numbers this industry quotes most, and the two that tell you least. Either can be pushed wherever a seller wants it by choosing the stops, the exits, the test window and the start date, which is why they lead almost every sales page. The sections below explain what each common measure actually says and, more usefully, how each one gets inflated or misread. Nothing here is a pitch. The point is to make you harder to sell to, by us or by anyone else.


Contents

  1. 01Win rate
  2. 02Profit factor
  3. 03Expectancy and payoff ratio
  4. 04Maximum drawdown and drawdown duration
  5. 05Sharpe and Sortino
  6. 06Alpha, beta and R squared
  7. 07Correlation between components
  8. 08Out-of-sample and walk-forward testing
  9. 09Confidence intervals
  10. 10Distribution shape

01

Win rate

What it measures
The fraction of closed trades that made money. Nothing else. It carries no information about how much a win pays or how much a loss costs, so a system can win nine trades out of ten and still lose money overall.
What good looks like
Anything from 30% to 70% can belong to a sound system. Trend-following approaches often win under 40% because they take many small losses while they wait for a large move. Mean-reversion approaches can win above 70% while carrying occasional large losses. The figure only means something when it sits next to the average win and the average loss.
How it gets gamed
This is the easiest number on the page to inflate. Widen the stop and cut winners early: the win rate climbs at once, because trades that would have been small losses get held until they recover, and gains are banked before they can reverse. The bill arrives later, as one rare and very large loss that gives back months of small wins. When a vendor leads with a high win rate, ask for the average loss and the largest single loss. Refusal to show those two numbers is the answer.

02

Profit factor

What it measures
Take every dollar the system won and divide it by every dollar it lost. A profit factor of 2 means two dollars gained for each dollar given back over the whole test. The entire history gets compressed into one ratio, and the compression is the weakness, because the ratio says nothing about when the losses arrived or how deep the equity curve sank on the way.
What good looks like
Above 1 the system made money over the period tested. Below 1 it did not. Live systems that pay real commissions and slippage tend to sit somewhere between about 1.2 and 2. Treat anything above 3 on a meaningful trade count with suspicion, because ratios like that in liquid futures markets usually mean a short favourable window or a fitted result.
How it gets gamed
The quickest inflation is the window. Start the test after a losing stretch and end it after a winning one, and the ratio improves without the strategy changing at all. Leaving out commissions and slippage helps too, since costs come straight out of the gross profit. A high profit factor over a small sample is noise, because thirty trades can produce almost any ratio by chance. Ask for the trade count and the exact date range, and confirm costs are in the figures.

03

Expectancy and payoff ratio

What it measures
What one trade is worth on average, wins and losses netted together. Multiply the win rate by the average win, subtract the loss rate multiplied by the average loss, and the result is expectancy. Payoff ratio is one of its two moving parts, the average win divided by the average loss. Between them they expose the trade-off that win rate hides.
What good looks like
Positive, after costs, is the floor. A useful second check is expectancy divided by the average loss, which tells you what fraction of a typical loss each trade earns back. On payoff ratio, high win rates normally pair with low payoffs and low win rates with high ones, and neither pairing is better in itself. What matters is that the combination nets clearly positive.
How it gets gamed
Expectancy is harder to fake than win rate, which is exactly why fewer vendors quote it. The remaining tricks are familiar. Quote it before commissions and slippage, so a thin per-trade edge that costs would erase looks real. Or quote it from a stretch that leaves out the losing regime. It is also an average, so one enormous win can carry a sample that otherwise loses. Ask whether the number still holds with the single best trade removed. A real edge does. A lucky one does not.

04

Maximum drawdown and drawdown duration

What it measures
Maximum drawdown is the deepest peak-to-trough fall in the equity curve, in currency or as a percentage. Drawdown duration is the time the curve spent below its previous high before making a new one. Depth is the worst moment. Duration is how long the worst moment lasted.
What good looks like
Judge depth against return, since a 10% worst drawdown means something different on a system earning 30% a year than on one earning 8%. Then look harder at duration, because it is the number that decides whether a human being keeps running the system. Most traders can absorb one sharp loss. Very few can watch an account go nowhere for months, and the ones who cannot usually switch the system off at the exact bottom.
How it gets gamed
A backtested drawdown is a floor, not a ceiling. History only contains the bad periods that happened to occur, and the worst drawdown of a live system is routinely deeper than anything in its test. Quoting drawdown from closed-trade equity is another softener, since it hides how far open positions sank before they were closed. Duration is simply left out, almost universally, because a months-long flat stretch sells worse than a tidy depth figure. If duration is missing, assume it is missing on purpose.

05

Sharpe and Sortino

What it measures
Both divide return by a measure of variability. Sharpe divides return above the cash rate by the standard deviation of all returns, upside and downside alike. Sortino divides by downside deviation only, so a system is not penalised for making its money in bursts. Quoting them together is the point, because the gap between them describes the shape of the returns.
What good looks like
For a single futures strategy, a live Sharpe above 1 is respectable and above 2 is uncommon. Sortino runs higher by construction, so never compare one system's Sharpe against another's Sortino. Read the gap. When Sortino sits far above Sharpe, most of the volatility being punished is upside volatility, which is the kind you want. When the two are close, gains and losses are shaped alike.
How it gets gamed
Sharpe rewards smoothness, and smoothness can be manufactured. Strategies that collect small steady gains against a rare large loss print lovely Sharpe ratios right up until the loss arrives. Frequency matters too: a ratio computed from monthly returns hides swings that a daily calculation would expose, and annualising a short sample flatters it further. Ask what return frequency the ratio was computed from and how many years it covers. A high Sharpe over eighteen calm months is a weather report, not a verdict on the system.

06

Alpha, beta and R squared

What it measures
Three numbers out of one regression, the system against the market it trades. Beta is the slope: how far the system moves when the market moves. Alpha is the intercept: the return left over once the market has been paid for its share, usually quoted annualised. R squared is how much of the variance in the returns the market explains at all. Between them they answer the thing almost nobody asks a vendor, which is whether this is doing something of its own or is long exposure with extra steps.
What good looks like
A strategy that claims an edge independent of market direction should show beta near zero, a low R squared, and an alpha that is positive after costs. Low R squared says the returns come from the strategy's own decisions rather than from the index it happens to trade on. Beta near 1 with a high R squared says you are paying a subscription for exposure an index future would give you for the cost of commissions, and whatever alpha is left has to cover the subscription before it has done anything for you. Read alpha last, because it only means something once the other two say it is measuring the strategy rather than the market.
How it gets gamed
The common abuse is silence. A system that is mostly repackaged market exposure looks like skill for as long as the market rises, so the regression is simply never run. You can run it yourself from a trade log in a spreadsheet in an afternoon. Where the numbers are published, the benchmark is the first thing to check, because regressing a stock index strategy against an unrelated market makes beta look flatteringly low and alpha correspondingly high. The second is frequency: alpha computed on daily returns, weekly and monthly should land in the same neighbourhood, and a figure that only survives at one of the three is an artefact of the window rather than a property of the system. Ask what the returns were regressed against, over what period, and at what frequency. A blank look is also an answer.

07

Correlation between components

What it measures
When a system is built from several models or signals, pairwise correlation measures how similarly the components behave. Near +1, they win and lose together. Near zero, each one contributes something the others do not, and the combined equity curve comes out smoother than any single part on its own.
What good looks like
The lower the better, and it has to hold over the full history, not a chosen window. Genuinely low correlation between components is one of the hardest numbers in a report to fake, because it has to survive thousands of trades through every kind of market. High correlation means the diversification is cosmetic. One idea wearing several hats, and every hat comes off in the same storm.
How it gets gamed
A convenient window is the first trick, so ask for the figure over the entire test. Padding is the second: components that barely trade drag the average down without adding anything real. The label itself can lie, since seven independent models that all read the same moving average with different settings will show their kinship the moment anyone measures it. If a vendor claims diversification, the pairwise correlation matrix is the receipt. Ask for it.

08

Out-of-sample and walk-forward testing

What it measures
Not a metric. A method, and the most important thing on this page. In-sample data is the history a strategy's rules were fitted to, and performance there is close to meaningless, because with enough parameters any rule set can be made to fit any past. Out-of-sample data is history the rules never saw during development. Walk-forward testing repeats the split in rolling windows: fit on one stretch, test on the untouched stretch that follows, step forward, do it again. The stitched-together out-of-sample results are the only backtest numbers worth reading.
What good looks like
Out-of-sample results should sit within shouting distance of the in-sample ones. Some decay is normal. A collapse is a verdict. Coverage matters as much as length, so a walk-forward that spans a crash, a grinding bear, a melt-up and a quiet drift says far more than a longer test of one kind of weather.
How it gets gamed
This is the concept every other number on the page hides behind, because any metric computed on fitted data inherits the fit. The subtler abuse is repetition: run twenty variants, keep the one with the prettiest out-of-sample stretch, and the out-of-sample has quietly become in-sample. Some vendors also call a test walk-forward when the parameters were tuned once over the entire dataset, which is an ordinary curve fit with a better name. Ask how many variants were tried before this one, and ask which data the final rules never touched. The honest answer to the first question is rarely one.

09

Confidence intervals

What it measures
A backtest produces one number from one sample of history. A confidence interval turns that point estimate into the range where the true value can reasonably sit. Bootstrap methods build the range by resampling the actual trades many thousands of times and recomputing the statistic each time.
What good looks like
Narrow beats wide, and wide beats absent. The bottom of the range is what to read first, because a Sharpe whose interval reaches below zero belongs to a system whose entire edge might be luck. More trades and more years tighten the range, which is one more reason a short glossy backtest deserves suspicion.
How it gets gamed
Omission is the main trick, since a single confident number sells better than an honest range. Where an interval does appear, check the sample behind it, because one built on fifty trades is wide enough to drive a truck through, and quoting only its optimistic end is the same sleight of hand as quoting the point estimate alone. Ask for the interval on any headline figure. A vendor who has never computed one has told you how carefully everything else was done.

10

Distribution shape

What it measures
Averages hide shape. Skewness measures which tail the returns lean toward: negative skew is many small wins against rare large losses, positive skew is the reverse. Kurtosis measures how fat the tails are, in other words how often extreme results turn up compared with a normal curve. MFE and MAE, maximum favourable and maximum adverse excursion, record how far each trade moved for and against the position while it was open.
What good looks like
Neither skew is disqualifying, but you should know which one you own. Negative skew feels comfortable right up to the tail event. Positive skew feels like slow bleeding until the payoff arrives, and many people abandon it first. MFE set against the actual exits shows whether the exits capture what the trades offered, and MAE set against the stop shows whether the stop sits where trades genuinely fail or somewhere arbitrary.
How it gets gamed
Negative skew is the classic subscription product, because months of small steady gains build a track record and a customer base before the loss that was always coming. Kurtosis goes unquoted for the same reason. MFE and MAE can be gamed in a backtest as well: an exit tuned until every historical trade left near its peak excursion is curve fitting in its purest form, and it will not repeat. Ask what the worst single day in the test was, and how often the shape of the distribution says to expect another one.

Then check ours

Then apply the same test to us. The report covers the full walk forward study and the live forward test as two separate records, labelled as such and never blended, and it goes to approved applicants before any payment is taken rather than after.

Futures and forex trading contains substantial risk and is not for every investor. An investor could potentially lose all or more than the initial investment. Risk capital is money that can be lost without jeopardizing ones financial security or life style. Only risk capital should be used for trading and only those with sufficient risk capital should consider trading. Past performance is not necessarily indicative of future results.

HYPOTHETICAL PERFORMANCE RESULTS HAVE MANY INHERENT LIMITATIONS, SOME OF WHICH ARE DESCRIBED BELOW. NO REPRESENTATION IS BEING MADE THAT ANY ACCOUNT WILL OR IS LIKELY TO ACHIEVE PROFITS OR LOSSES SIMILAR TO THOSE SHOWN; IN FACT, THERE ARE FREQUENTLY SHARP DIFFERENCES BETWEEN HYPOTHETICAL PERFORMANCE RESULTS AND THE ACTUAL RESULTS SUBSEQUENTLY ACHIEVED BY ANY PARTICULAR TRADING PROGRAM. ONE OF THE LIMITATIONS OF HYPOTHETICAL PERFORMANCE RESULTS IS THAT THEY ARE GENERALLY PREPARED WITH THE BENEFIT OF HINDSIGHT. IN ADDITION, HYPOTHETICAL TRADING DOES NOT INVOLVE FINANCIAL RISK, AND NO HYPOTHETICAL TRADING RECORD CAN COMPLETELY ACCOUNT FOR THE IMPACT OF FINANCIAL RISK OF ACTUAL TRADING. FOR EXAMPLE, THE ABILITY TO WITHSTAND LOSSES OR TO ADHERE TO A PARTICULAR TRADING PROGRAM IN SPITE OF TRADING LOSSES ARE MATERIAL POINTS WHICH CAN ALSO ADVERSELY AFFECT ACTUAL TRADING RESULTS. THERE ARE NUMEROUS OTHER FACTORS RELATED TO THE MARKETS IN GENERAL OR TO THE IMPLEMENTATION OF ANY SPECIFIC TRADING PROGRAM WHICH CANNOT BE FULLY ACCOUNTED FOR IN THE PREPARATION OF HYPOTHETICAL PERFORMANCE RESULTS AND ALL WHICH CAN ADVERSELY AFFECT TRADING RESULTS.

NinjaTrader® is a registered trademark of NinjaTrader Group, LLC. No NinjaTrader company has any affiliation with the owner, developer, or provider of the products or services described herein, or any interest, ownership or otherwise, in any such product or service, or endorses, recommends or approves any such product or service.