Why adding criteria makes a screen look better and work worse
Each criterion added to a screen improves how it looks on the data you built it on and usually worsens how it performs on data you did not. Why that happens, and how to tell fit from edge.
2 min read
A screen with three criteria returns too many names. Add a fourth and the list improves. Add a fifth and it improves again. Keep going and you arrive at a screen that would have selected an excellent set of companies over the period you tested it on, and has no particular reason to select anything useful next year.
Every criterion is a parameter
A threshold is a choice, and a choice made by looking at the outcome is a parameter fitted to that outcome. "Debt-to-equity below 0.5" is not a law; it is a number that happened to work. The screen has not learned something about balance sheets. It has learned something about this sample.
Illustrative, not measured — the crossover point depends entirely on the data and the criteria. The shape is the durable part: the two lines separate, and the gap between them is the fit.
Why the gap opens
01Each filter removes names. Removing names that performed badly in the sample raises the sample's average by construction, whether or not the criterion means anything.
02The universe shrinks, so the remaining set is small and its average is noisy — a handful of names now determine the result.
03Criteria correlate. Six filters on profitability are not six independent tests; they are one idea, expressed six times, and they narrow far more than they diversify.
04The criteria that survive your review are the ones that improved the number, which is a selection process operating on noise.
Telling fit from edge
Hold out a period before you start and do not look at it until the screen is final. Once you have tuned against it, it is training data.
Round every threshold. If the screen collapses when 0.47 becomes 0.5, it was resting on the decimal, not the idea.
Count what you tried, including the abandoned versions. A screen found on the fortieth attempt is weaker evidence than the same screen found on the first.
Check whether the result survives in an adjacent market or an adjacent decade. A real effect usually leaves traces elsewhere.
Prefer fewer criteria that you can each justify aloud, without reference to what they did to the backtest.
A screen built this way returns a longer, less impressive list. That is the correct outcome: it has stopped pretending to know things it only fitted, and the names on it are there for reasons that will still exist next year.
Common questions
How many criteria should a stock screen have?
Few enough that you can justify each one without referring to what it did to the backtest. There is no fixed number, but criteria correlate heavily — six filters on profitability are one idea expressed six times — so additional filters usually narrow the universe far more than they add independent information.
How do I know if my screen is overfitted?
Round the thresholds and re-run. If performance collapses when 0.47 becomes 0.5, the result was resting on the decimal rather than the idea. Then test on a period you have never tuned against; a large gap between that and your development period is the fit.
Why does adding a filter always improve backtest results?
Because filters are chosen by looking at outcomes. Removing names that did badly in the sample raises the sample's average by construction, whether or not the criterion has any predictive content.
Related product
VriddhiX Screener
Find and analyse opportunities using your own criteria.
Most screens return either four companies or four hundred. How to choose criteria that narrow a universe without accidentally encoding a single sector, and how to sanity-check a screen before trusting it.
Survivorship bias, look-ahead bias and overfitting each inflate backtested returns in ways that vanish in live trading. What each one is, how it creeps in, and how to test for it.