The method — Compustat, deciles and sixty years of data
Chapter 1 of 13 · 10 min
Before O'Shaughnessy shows what works, he shows how to find out. The book builds on the Compustat database with thousands of American stocks from 1951 onward, divided into ten equally large portfolios — deciles — that are rebalanced once a year. This is the template for all serious factor analysis, including AK1A's own backtests.
The setup is pure mechanics. Every year on 31 December the entire universe is divided into ten equally large groups — deciles — by a factor, for example P/S. Decile 1 contains the ten percent of stocks with the lowest value on the factor, decile 10 the ten percent with the highest. You buy the decile portfolio, hold for exactly one year and repeat the procedure.
No opinions, no adjustments along the way, no knowledge of the future. Then you compare the deciles' aggregate results with the universe as a whole over decades. This is how you test a factor — and it is exactly the way AK1A works when a variable is to earn its place in the system.
Two universes underlie almost all tests. All Stocks is all companies of sufficient size — thousands of stocks in modern editions. Large Stocks is the larger and more traded companies. During the main period 1951–1996, All Stocks returned about 13,6 percent per year while Large Stocks returned about 11,9 percent — the small-cap premium is thus already in the base data, something chapter 10 returns to.
The fourth edition extends the tests to 2009 and in several chapters back to 1927 — and the main conclusions hold essentially unchanged through the IT bubble and the financial crisis.
The book's most important methodological choice concerns survivorship bias. An ordinary stock database contains only companies that still exist — those that went into bankruptcy or were delisted have disappeared. If you test factors on such a sample, the result is systematically overstated, because all the dead companies are missing from the statistics. O'Shaughnessy uses the database's research file, which includes historical, now deceased companies, precisely so that the result will mirror reality: whoever bought decile 1 in 1955 could not know which companies would survive to 1996.
In addition, he looks for MONOTONIC relationships — that the result improves or worsens evenly across the deciles — because a gradual pattern over ten steps is hard to explain with luck, while a single top reading is easily data mining.