News headlines transforming into economic graphs

Can News Headlines Predict the Economy? How Textual Data Is Changing Tail Risk Forecasting

"Uncover the hidden signals in news articles and how they're reshaping macroeconomic forecasts in real-time, offering a new edge in predicting economic tail risks."


In times of economic uncertainty, like the Global Financial Crisis and the COVID-19 pandemic, accurately predicting tail risks becomes essential. Macroeconomic forecasts need to be reliable so that policymakers and central banks can get a better grasp on when the economy is heading for a period of high economic risk. Quantile predictions, which offer a detailed view of potential outcomes, are increasingly becoming important.

Recent work in macroeconomic forecasting is using textual data, analyzing news articles and reports to find important economic signals. Textual data can provide more timely information. Researchers use this data to understand the narratives shaping economic events, since narratives can influence economic outcomes. Quantifying these narratives is a valuable task.

A recent study analyzes whether textual data adds value to macroeconomic quantile predictions. The study uses a data-driven method to analyze news articles along with economic indicators, providing monthly tail risk forecasts for employment, industrial production, inflation, and consumer sentiment. It uses a range of quantiles to assess the benefits of text-based predictors compared to traditional methods.

AI Search Multiple angles on this topic

Defining the Economic Problem at the Core of Forecasting

Economics, as Merriam-Webster defines it, concerns the production, distribution, and consumption of goods and services, while the study of economics centers on decisions and choices made to attain the best possible outcome. Within this framework, economic agents can be individuals, businesses, organizations, or governments, and economic transactions occur when parties agree on the value or price of a good or service, commonly expressed in a currency. Headlines and current events about these transactions and indicators form the raw material that both news organizations and analysts track continuously, as reflected in live economic news coverage that pushes out the latest data, blogs, and video. This statistical and narrative stream is what textual approaches attempt to convert into forward-looking signals about the economy.

The Standard Approach and Where It Falls Short

The standard approach to economic analysis, as summarized by Investopedia, treats economics as the study of how societies manage scarce resources to produce, distribute, and consume goods and services, and how individuals, businesses, and governments allocate those scarce resources. Traditional tail-risk forecasting under this framework relies largely on structured quantitative indicators such as output, inflation, and financial market data. The limitation of this accepted method is that it is anchored to published statistics that lag real-world sentiment and emerge only after events unfold, leaving early warning signals from language and sentiment largely untapped.

An Evolving Record of Macroeconomic Forecasting

The historical record suggests that efforts to anticipate economic turning points have been a recurring ambition of economists and market participants, though no single milestone established a definitive breakthrough. Early forecasting relied on narrative interpretation and heuristics long before systematic quantitative models, and those narrative roots make the recent revival of textual analysis less a novelty than a return to older traditions. That said, the broader history is not well documented in accessible sources, so this summary should be read as a general characterization rather than a settled timeline of achievements.

Decoding the News: How Textual Data Enhances Economic Forecasting

News headlines transforming into economic graphs

The research uses news-based data along with FRED-MD economic indicators to make quantile predictions for several factors, such as employment and consumer sentiment. The results show that news data contains valuable information not found in standard economic indicators. By using this information, forecasters can improve tail risk predictions.

The study uses the Correlated Topic Model (CTM) to analyze about 800,000 newspaper articles from The New York Times and The Washington Post, extracting numerical text-based predictors. The CTM identifies word clusters, or topics, and their proportions, indicating media coverage. The analysis also uses tone-adjusted topic proportions to determine whether the tone of a topic is positive or negative.

  • Macroeconomic Predictors: Uses FRED-MD database.
  • Unadjusted Text-Based Predictors: Incorporates raw topic proportions from news articles.
  • Tone-Adjusted Text-Based Predictors: Combines topic proportions with sentiment analysis to gauge positive or negative tones.
AI Search Multiple angles on this topic

An Emerging but Still Nascent Research Front

Recent work on using headline and news text for economic prediction appears to be an active and growing area, but the specific findings, methods, and model families involved are not consistently documented in the sources currently at hand. The general trend seems to be toward combining language-derived signals, such as sentiment and narrative framing, with conventional economic indicators to improve near-term forecasts. Until peer-reviewed reviews establish the contours of this research, it is safer to describe it as promising and incomplete rather than to treat any individual result as settled.

Evidence of Failure and Reasonable Skepticism

There are legitimate reasons to doubt whether text-derived signals can reliably predict tail events, including the possibility that news simply mirrors data already reflected in prices rather than anticipating them. Forecasts of rare, extreme outcomes are especially prone to failure because such events are by definition unusual and poorly represented in training or historical data. A balanced treatment should acknowledge that the case against textual tail-risk forecasting is as plausible as the case for it, and that documented failures remain an honest open question rather than a settled verdict.

Comparing Textual Signals with Conventional Indicators

In principle, a comparison of textual forecasting with conventional quantitative approaches would highlight trade-offs: numeric indicators are precise and well-understood but lag the present, while text signals can capture sentiment and shifts in tone more quickly but are messier and harder to validate. Such a comparative assessment depends on empirical results that are not available in the current source material, so any relative ranking of the two approaches would be premature. The reasonable interim conclusion is that they are complements rather than strict substitutes, with textual methods most valuable as an overlay on established indicators.

To prevent overfitting, the study uses Bayesian quantile regressions (QRs) with shrinkage priors. These methods are effective for forecasting in high-dimensional settings. The study also uses non-linear models, including Gaussian Process Regressions and QR Forests, to capture complex predictive relationships. These methods help to evaluate the empirical differences between linear and non-linear models.

The Future of Forecasting: Integrating News and Economic Data

The study's findings suggest that combining textual data with economic indicators improves tail risk forecasts, particularly in extreme economic situations. Adding tone-adjusted text-based predictors enhances forecast accuracy compared to using unadjusted predictors alone. Non-linear models capture predictive relationships better than linear models. By using textual data and advanced analytical methods, forecasters can gain valuable insights for predicting economic tail risks.

AI Search Multiple angles on this topic

Bringing the Threads Together

Across the material considered, the case for using news headlines to forecast tail risk rests on a plausible intuition: economic decisions and confidence are shaped by the stories people read as much as by the numbers they see. The strongest framing is integrative, treating textual signals as supplements to standard quantitative economics rather than replacements for it. Expert commentary, where it exists, appears to reflect this measured view, but direct quotations and settled expert consensus are not captured in the current sources and should not be fabricated.

Where the Field May Head Next

Looking ahead, it seems likely that the frontier of tail-risk forecasting will involve increasingly sophisticated language models applied to news text, with attention to speed of signal extraction and robustness against misleading headlines. Realistically, progress will depend on data availability, validation against extreme events, and careful handling of the noise inherent in text. These are projections about an emerging field rather than findings, so they should be treated as informed speculation about plausible directions.

Systemic Constraints on Textual Forecasting

Any serious use of textual data in forecasting must contend with broader systemic challenges, including the reliability of news sources, the amplification of outliers and misinformation, and the fact that headline behavior can change across media cycles and markets. These issues mean that a model trained on one period or outlet may not generalize cleanly to the next. Acknowledging these systemic risks is important because they directly affect whether headline-derived tail-risk signals can be trusted at the scale and speed where they would be most useful.

What This Means for People, Not Just Models

Behind every tail-risk forecast is a human consequence: headlines shape the confidence, spending, and investment decisions of real economic agents such as individuals, businesses, and governments. If textual signals can genuinely anticipate extreme outcomes, they could give people and institutions earlier warning to adjust expectations and behavior. Given the uncertainty in the methods, the human value of this research ultimately depends on honest communication of both its potential and its limits.

About this Article -

Written with AI assistance from published research, and reviewed by the Mystum team. See our About page for more information.

This article is based on research published under:

DOI-LINK: https://doi.org/10.48550/arXiv.2302.13999,

Title: Forecasting Macroeconomic Tail Risk In Real Time: Do Textual Data Add Value?

Subject: econ.em

Authors: Philipp Adämmer, Jan Prüser, Rainer Schüssler

Published: 27-02-2023

Everything You Need To Know

1

How can news headlines be utilized to forecast economic downturns?

News headlines can be analyzed to forecast economic downturns by extracting valuable signals from textual data. Researchers leverage news articles and reports to identify important economic signals. The process involves using a data-driven method to analyze news articles along with established economic indicators. This approach provides monthly tail risk forecasts for key areas like employment, industrial production, inflation, and consumer sentiment. By incorporating this textual data, forecasters aim to gain timely insights into potential economic risks, enhancing the accuracy of predictions compared to relying solely on traditional economic indicators like those from the FRED-MD database.

2

What specific types of textual data are used to enhance macroeconomic forecasts?

The study uses several types of textual data to enhance macroeconomic forecasts. The primary source is news articles from The New York Times and The Washington Post, comprising about 800,000 articles. This data is processed using the Correlated Topic Model (CTM) to extract numerical text-based predictors. These predictors include raw topic proportions, representing the frequency of certain word clusters or topics within the news articles. Additionally, tone-adjusted text-based predictors are used. These combine topic proportions with sentiment analysis to gauge whether the tone of a topic is positive or negative, adding another layer of insight to the analysis.

3

What is the role of the Correlated Topic Model (CTM) in analyzing news articles for economic forecasting?

The Correlated Topic Model (CTM) plays a crucial role in analyzing news articles for economic forecasting. The CTM identifies word clusters, known as topics, within the vast amount of textual data from news articles. It determines the proportions of these topics, indicating how frequently different themes are covered in the media. This process allows researchers to quantify the narratives shaping economic events. By understanding these narratives, forecasters can identify hidden signals relevant to economic trends. CTM is essential for extracting numerical data from text, which can then be used in conjunction with economic indicators to improve the accuracy of economic forecasts.

4

How do the different types of text-based predictors (unadjusted vs. tone-adjusted) impact the accuracy of economic forecasts?

The study highlights that tone-adjusted text-based predictors enhance forecast accuracy compared to using unadjusted predictors alone. Unadjusted text-based predictors incorporate raw topic proportions from news articles, offering a basic view of media coverage. However, by combining topic proportions with sentiment analysis, tone-adjusted predictors offer a more nuanced understanding. This adjustment allows forecasters to assess the positive or negative tone associated with a topic, providing additional context that can be crucial in predicting economic tail risks. The added information from sentiment analysis refines the forecasts, making them more sensitive to market sentiment and economic events.

5

What analytical methods are employed to process textual data and economic indicators for forecasting tail risks, and why are they chosen?

The study utilizes several advanced analytical methods to process textual data and economic indicators for forecasting tail risks. Bayesian quantile regressions (QRs) with shrinkage priors are used to prevent overfitting, which is a risk when dealing with high-dimensional data. Non-linear models, including Gaussian Process Regressions and QR Forests, are employed to capture complex predictive relationships that linear models might miss. These methods are chosen for their ability to handle the complexity of the data and provide a detailed view of potential economic outcomes, which is particularly valuable for predicting tail risks. This detailed view is provided by a range of quantiles, allowing for a comprehensive assessment of the potential economic scenarios.

Newsletter Subscribe

Subscribe to get the latest articles and insights directly in your inbox.