The first notable feature of the full-sample results is the near-zero market beta across all portfolio variants. The CAPM beta for the EW prepared remarks portfolio is 0.033, and 0.043 for the EW Q&A portfolio. The VW counterparts register betas of 0.08 and 0.015, respectively. These confirm that the long-short strategy is effectively market neutral, and the returns are therefore not compensation for bearing broad market risk.
Examining the cumulative returns charts (Figures 2 - 5), the strategy shows clear regime dependence. Performance is strong during the 2006 - 2012 period window surrounding the 2008 financial crisis, then flattens during the 2012-2016 period, before generally deteriorating from 2017 onwards (Figures 6-7). This pattern is consistent with the hypothesis that NLP-based earnings call analysis has become increasingly commoditized. As hedge funds and quantitative desks deployed similar FinBERT-style models in the late 2010s, the sentiment mispricing alpha was progressively arbitraged away. This motivated a formal pre/post 2017 split to separate the period where the signal was relatively proprietary from the period where it had become widespread.
The pre/post 2017 breakdown in Table 4 makes the regime dependence precise. Before 2017, the prepared remarks EW portfolio generated an FF5 alpha of 0.631% per month and the Q&A EW portfolio generated 0.477% per month. While they still fall short of the 5% significance threshold given the limited sub-period sample, they are more significant than the alpha over the full sample period. The same can be seen for the VW portfolios. After 2017, both strategies flip to negative alpha (-0.414% and -0.304% per month for prepared remarks, respectively), confirming that the full-sample average is being pushed upwards by the strong pre-2017 regime and diluted by the post-2017 deterioration.
One natural hypothesis for this post-2017 deterioration is that FAANG stocks drove it: as Apple, Amazon, Google, Facebook (now Meta), and Netflix grew to dominate the top 100 by market capitalization, their sustained strong performance while sitting in the short leg could have mechanically suppressed value-weighted returns. However, Figure 9 (see the Appendix) rules this out as it shows that the FAANG contribution to the short leg returns (pink line) is mostly close to zero throughout the entire sample, while the non-FAANG names account for virtually all return variations in both regimes. The post-2017 deterioration is visible in the non-FAANG contribution alone, pointing instead to the broader commoditization of the NLP sentiment signal.
The monthly window decomposition in Table 5 provides a further layer of insight by asking where within the 3-month holding period the return is actually generated. For the prepared remarks signal, the M1 → M2 window produces the largest alpha (0.810% per month, EW), but it is not statistically significant (p = 0.127). For the Q&A signal, however, the picture is more striking: the M2 → M3 window generates an EW FF5 alpha of 0.771% per month, significant at the 5% level (p = 0.044). The M0 → M1 and M1 → M2 windows for the Q&A sentiment produce much smaller alphas of 0.201% and 0.206%, respectively. This suggests that the Q&A sentiment signal is a slow-moving predictor: the market takes two full months to begin meaningfully incorporating the information embedded in the unscripted analyst exchange, with most of the price correction occurring in the third month.
Table 6 further decomposes the Q&A M2 → M3 window by regime. Pre-2017, the alpha is 0.611% (p = 0.186%), which is positive but not independently significant. Post-2017, the alpha rises to 1.175% per month (p = 0.022), achieving strong statistical significance. This is the most striking finding of the study. Rather than being eliminated by the spread of NLP tools, the M2 → M3 Q&A window appears to have strengthened after 2017. One potential explanation is that as fast-moving algorithmic strategies more aggressively front-run the immediate post-earnings sentiment signal (depressing early-window returns, as evidenced by the M0 → M1 alpha of -2.186% post-2017), the residual longer-horizon information in the Q&A section becomes relatively more valuable and less-contested. The price discovery process for Q&A-embedded information may be more gradual and less susceptible to crowding than the immediate transcript-scanning that characterizes modern earnings-event strategies.
Taken together, the results show that the overall strategy is suggestive of alpha within the sentiment-driven mispricing during earnings calls but not fully conclusive given the statistical insignificance. Further analysis of the results show, however, that the aggregate 3-month strategy obscures a powerful, statistically significant signal concentrated in the final month of the holding window on Q&A sentiment. With further testing and refinements, this can be a viable trading strategy for hedge funds to adopt and make significant alpha.