
Inside AlphaNova's Research: Signal Diversity, Correlation, and the Geometry of Forecast Libraries
The Space of Signals Is Bigger Than It Looks: Correlation, Diversity, and the Research Behind AlphaNova
What happens when you combine thousands of forecasts, and how do you know whether they are really different?
For a long time, I have been pushing a simple idea at AlphaNova: there is room for far more useful market signals than we tend to imagine. Even a simple approach, such as giving each signal an equal weight, deserves serious attention.
At AlphaNova, we run prediction competitions in financial markets. People bring different datasets, methods, intuitions and ways of looking at the same assets. Each submission is a signal: a set of forecasts about how those assets will perform relative to one another. The interesting question is what we can build from the collection.
Suppose you already have a good forecast. Someone arrives with another that is weaker on its own but makes different mistakes. It could improve the combination. Someone else arrives with a brilliant-looking forecast that largely repeats what you already have. Its incremental contribution could be smaller.
That makes the search for signals a collective problem. We need to understand what each new contribution adds to the whole.
Trying to make that intuition precise led to three connected papers. They ask what signal diversity actually measures, what happens as a library grows, and how much room remains when we insist that signals stay sufficiently different. Together, they make me more interested in the opportunity, and more careful about how we describe it.
Start with the simplest possible combination. Put the signals on a comparable scale, give each one an equal weight, and average their forecasts for each asset.
There is a sensible intuition behind this. If several forecasts contain a little useful information, and enough of their errors cancel, the average can preserve useful content while reducing some of the surrounding variation. But the cancellation is an assumption to investigate. Shared mistakes survive averaging just as shared insight does.
Nor does averaging automatically increase raw profit. With fixed combination weights and before costs, the blend's profit and loss is simply the corresponding average of its components' profit and loss. The attraction is the possibility of a more reliable forecast or a better balance between return and risk. Position sizing then becomes a separate decision.
Equal weighting is a useful starting point because it makes the behaviour of the collection relatively easy to inspect. It also gives more elaborate combination methods something concrete to beat.
The first complication appears in a word we use constantly: correlation.
There are two different questions hiding inside it. When we compare signals, we ask whether they express similar views across assets at the same time, then average that similarity over time. When we compare profit-and-loss histories, we ask whether the strategies tend to have good and bad periods together.
Those questions sound related. They are. But one answer does not determine the other.
Imagine two forecasts that look quite different across a basket of assets. Once returns arrive, it turns out that the parts of those forecasts that generated gains and losses behaved almost identically. Much of their visible disagreement was in directions that contributed nothing to that period's payoff. Their signal correlation can be low while their profit-and-loss correlation is high.
The first paper, "Signal Correlation, IC, and PnL Dependence", makes this distinction exact. Looking back at each period, we can separate a forecast into its alignment with the realised return pattern and everything perpendicular to that pattern. The alignment is measured by the information coefficient, or IC: the cross-sectional correlation between the forecast and what subsequently happened.
For the normalised portfolios in the paper, the period's gross payoff depends on that alignment and on how widely asset returns were spread. Correctly distinguishing winners from losers matters more when the winners and losers finish far apart.
Signal similarity also sees the rest of the forecast. Profit-and-loss correlation involves further changes, including the size of each period's return opportunity and the removal of each strategy's average profit. There is therefore no universal conversion from one correlation to the other.
This is especially relevant to our contests. When ranking submissions, we enforce a ceiling of 50% on pairwise time-averaged cross-sectional signal correlations. That encourages contributions with different forecast histories. It gives us no corresponding promise that their profit-and-loss correlations will stay below 50%.
The rule has a useful job. We still need to examine how the resulting strategies win and lose together.
And the parts of a signal that contributed nothing to one realised payoff are not permanently irrelevant. They can affect risk, behave differently at another horizon, and matter when we combine forecasts and rescale the resulting portfolio. The decomposition helps explain past outcomes; it uses realised returns, so it cannot tell us tomorrow's winning direction in advance.
The second complication is what we mean by a large library.
One observation motivating the second paper was that an equal-weight blend of roughly 3,000 signals across 20 assets appeared strongly aligned with a familiar reversal pattern: buying recent relative losers and selling recent relative winners. The paper treats that observation as a question to explain.
Was this simply inevitable because so many signals were competing for so little space?
At a single moment, the dimensional constraint is real. Twenty asset scores, after subtracting their average, have only nineteen independent directions available. You cannot keep adding new linearly independent directions to that snapshot forever.
But a signal is an entire sequence of forecasts. Two signals can express similar views today and behave very differently tomorrow, or respond differently when volume, volatility or market conditions change. A small asset universe can support a very rich collection of histories.
Twenty assets do not imply twenty trading ideas.
What the equal-weight blend approaches depends on how that collection is populated. Under suitable assumptions about a stable research process and dependence between submissions, the average approaches the mean forecast produced by that process.
This has a practical consequence: equal weights across submissions can give very unequal influence to research ideas.
If one family of methods contributes 900 signals and another contributes 100, equal weighting allocates 90% of the submission weight to the first family. Those entries may differ, yet still carry a common tendency. The blend reflects that imbalance.
If the research process keeps producing variations on reversal, a large average may reveal reversal very clearly. Understanding that outcome requires understanding the research and selection process. The number of submissions alone cannot explain it.
There is another subtlety here. The average forecast need not point along the strongest axis in the collection, which the papers study using principal components.
Picture a collection of arrows pointing mostly upwards or downwards, all leaning slightly to the right. The upward and downward contributions cancel in the average, leaving a small arrow pointing right. Yet the strongest axis of the collection is vertical.
That is how an equal-weight average can point in one direction while the dominant pattern of signal activity lies elsewhere. The analogous calculations on profits introduce another distinction, because converting forecasts into payoffs discards and reweights information, as the first paper explains.
So whenever someone says a large ensemble has converged to “the main factor”, it helps to ask which object they measured: the average forecast, the strongest axis among the signals, a component of asset returns, or a common pattern in strategy profits. Agreement between them needs an explanation.
Once forecasts have been normalised, similarity can be represented geometrically: similar signals point in similar directions. Comparing complete histories lets us express an upper correlation cap as a minimum separation between those histories. We can then ask how many sufficiently separated histories fit.
The answer supports the intuition that the space is large. The paper constructs collections whose size can grow exponentially with history length while respecting an upper correlation cap. Even with a fixed number of assets, the possible variety over time is enormous.
Those are mathematical constructions. They count possible histories, without establishing how many can be generated in advance and remain predictive after trading costs. The distinction matters: the geometry demonstrates capacity, while research and live evaluation have to establish usefulness.
Separation also leaves plenty of room for a shared component. Signals can carry a common useful tendency while differing substantially elsewhere. If those other contributions cancel sufficiently well in the average, that tendency can survive clearly in the blend. This is one reason a diversity constraint and a useful collective forecast can coexist.
But the cap does not force that outcome. The third paper constructs large, separated collections with positive realised predictive alignment in which the average agrees with the leading signal component, and others in which the two point in perpendicular directions. How the remaining variation is distributed makes the difference.
Even reaching a point where no further candidate passes the filter would tell us something about the candidate pool and the signals already admitted. It would not establish that we had exhausted the space of possible ideas. Which candidates arrive, and which are admitted first, can affect the collection we end up with.
There are several practical lessons I take from the trilogy.
First, we should judge additions by their contribution to a combination as well as their standalone performance. A promising submission deserves testing alongside what we already have, at a comparable risk budget and with realistic costs.
Second, we should track both forecast similarity and the co-movement of profits. The 50% rule helps preserve variety in the forecasts. Testing profits, drawdowns and exposures tells us more about the diversity that reaches a trading account. An average correlation over the whole sample can also conceal periods of much closer agreement.
Third, we should pay attention to the composition of the library. A growing submission count may represent new approaches, more variations within existing families, or a mixture of both. Those are different kinds of growth.
Finally, adding signals does not add market history. Thousands of candidates tested on the same period still share the same limited evidence. Even random forecasts can look encouraging if we retain the ones that happened to work. A slightly positive historical IC needs scrutiny, especially after a large search. Selection decisions need evaluation on fresh periods, with allowance for persistence and overlapping forecast horizons.
The same discipline applies to implementation. Forecasts become positions through sizing, constraints and trading decisions. A pleasing geometric picture has to survive that journey. Turnover and costs belong in the evaluation of the implemented blend.
These papers provide theory and synthetic illustrations; they do not establish live profitability or prove that equal weighting is optimal. What they do provide is a clearer way to investigate the ambition behind AlphaNova.
I still believe there is substantial room for people to contribute useful signals, including people working on the same assets with relatively simple methods. The opportunity comes from the variety of views they can form, how those views change through time, and what survives when we combine them carefully.
That is the question I want our competitions to keep exploring: what can your signal help the whole collection do better?
The research behind this article: Signal Correlation, IC, and PnL Dependence, Large Signal Libraries: Equal-Weight Limits and the Divergent Spectra of Signals and PnL, and Separated Signal Libraries: Packing, Saturation, and Joint Spectral Limits.