AI did and econometrics
Econometrics | Data Analysis | Forecasting

When Words Become Data: How AI Turns News, Speeches, and Reports into Economic Signals

Economic information is increasingly hidden inside words rather than spreadsheets. AI and natural language processing can convert news articles, central-bank speeches, corporate reports and earnings calls into numerical measures of sentiment, uncertainty and policy direction that economists can test alongside traditional economic data.

By Outsider Advisory · September 29, 2026

Economists traditionally work with numbers: GDP growth, inflation, unemployment, interest rates, investment, wages and consumer spending. Yet enormous amounts of economically relevant information arrive first as words rather than numbers, including newspaper stories, central-bank speeches, corporate earnings calls and regulatory reports. A CEO saying that customers are becoming cautious or a central banker describing inflation as “persistent” may contain information that will appear in official statistics only weeks or months later. AI allows economists to transform this qualitative language into quantitative variables that can be analyzed alongside conventional economic data.

This idea is often described as “text as data.” Instead of treating thousands of documents as material that must be read individually, natural language processing can classify their topics, measure sentiment, identify uncertainty and detect changes in language over time. The resulting output might be a daily news-sentiment index, a monetary-policy tone score or a company-level measure of concern about labor shortages. Words that once existed outside the econometric dataset can therefore become variables inside it.

This does not mean AI literally understands the economy simply because it can analyze language. The system is converting patterns in text into structured measurements that researchers can subsequently compare with observable economic outcomes. Econometric analysis is then needed to determine whether those measurements actually contain useful information about inflation, GDP, financial markets or corporate behavior.

The approach has moved well beyond experimental applications. A 2025 IMF study used a fine-tuned large language model to classify 74,882 documents from 169 central banks spanning 1884 to 2025, measuring dimensions including topic, communication stance, sentiment and intended audience. AI is making it possible to quantify economic communication at a scale that would have been extremely difficult to achieve through manual reading alone

Central-Bank Language Can Become a Monetary-Policy Indicator

Central banks provide an ideal example because monetary policy depends heavily on communication. Interest-rate decisions matter, but so do statements about inflation, economic growth, employment, financial conditions and the probable direction of future policy. Investors therefore analyze individual words and phrases for indications that policymakers have become more concerned about inflation or more worried about economic weakness.

Natural language processing can systematize this process. Researchers can break a central-bank statement or press conference into sentences, classify each sentence by topic and assign numerical scores describing its tone or policy direction. A paragraph about inflation can therefore be converted from qualitative language into a measurable time series.

The European Central Bank has demonstrated one version of this approach. Its researchers use natural language processing to categorize individual messages in monetary-policy statements and press conferences into topics such as monetary policy, the economic outlook and inflation, and then quantify the direction and strength of their tone. In the ECB’s example, language saying that underlying inflation “declined further” receives a dovish score, while stronger language about inflation falling rapidly receives an even more dovish numerical value.

Once researchers have constructed these scores, they can compare them across meetings. If the inflation component steadily becomes more hawkish while the growth component becomes more pessimistic, economists have a numerical description of a shift that previously required subjective interpretation. The words have effectively become observations in a dataset.

Federal Reserve researchers have applied similar methods to FOMC communication. A 2025 Federal Reserve paper analyzed financial-press coverage surrounding FOMC meetings and created a sentiment index covering conventional monetary policy, asset purchases and forward guidance. Surprises in that sentiment measure helped explain variation in major asset prices beyond conventional measures of monetary-policy surprises.

Research from the Federal Reserve Bank of Kansas City similarly found that qualitative descriptions in FOMC statements can matter for financial markets. Its natural-language measure indicated that changes in descriptions of economic conditions and risks could affect bond prices substantially, and the authors found that the tone of statements could influence financial conditions even when the policy rate itself did not change. Markets respond not only to what central banks do, but also to what central banks say about what may happen next.

Speeches provide another source because policymakers communicate between formal meetings. A 2025 Journal of Econometrics study used U.S. Federal Reserve speeches and machine-learning methods to extract implied revisions to forecasts for GDP, inflation and unemployment. Those speech-derived revisions helped explain volatility and tail risk in both equity and bond markets, showing how apparently qualitative remarks can contain measurable economic information.

News and Corporate Language Can Reveal Economic Changes Earlier

Central-bank communication is only one source of text. Newspapers continuously describe layoffs, consumer confidence, housing markets, bankruptcies, supply disruptions, wage pressures and corporate investment. Because news is published every day, it can potentially reveal economic changes before monthly or quarterly statistics become available.

Researchers can collect millions of articles and use natural language processing to determine whether economic coverage is becoming more positive, negative or uncertain. Daily observations can then be aggregated into a news-sentiment index and incorporated into forecasting models. The economic question becomes testable: does adding the text-based indicator improve predictions relative to models using only conventional variables?

Evidence suggests that it sometimes does. A 2024 Journal of Applied Econometrics study analyzed more than 27 million articles from 26 newspapers in five languages and found that news-based sentiment indicators contained useful information for forecasting macroeconomic variables in five major European economies. Importantly, their predictive information remained relevant after controlling for other indicators available to forecasters in real time.

ECB research has reached related conclusions about GDP nowcasting. Newspaper-based sentiment measures have shown useful relationships with economic activity and can improve real-time assessments, particularly early in a quarter when many conventional indicators for that quarter have not yet been published. Text data can therefore help fill the information gap created by slow official statistics.

Corporate communication provides another enormous source of information. Earnings calls contain executives’ comments about customer demand, hiring, costs, investment, supply chains, pricing and risks, frequently before these developments become obvious in aggregate statistics. Instead of manually reading thousands of transcripts, researchers can search and classify language across companies, industries and time.

Federal Reserve researchers have used earnings-call transcripts to study how businesses discuss artificial intelligence. Research highlighted by the San Francisco Fed found that sentences mentioning AI increased more than fivefold during the year following ChatGPT’s public release, with positive language more common than negative language while uncertainty surrounding AI also increased. Corporate text can reveal not only what companies are doing, but how management perceptions and expectations are changing.

The same method can be applied to many economic questions. Researchers could measure references to labor shortages, recession, supply-chain disruptions, tariffs, inflation or investment and then examine whether those indicators predict subsequent employment, capital spending, prices or earnings. A CEO’s sentence becomes economically useful when it can be classified consistently across thousands of companies and connected statistically to later outcomes.

AI Turns Language Into Variables—but Econometrics Tests Whether They Matter

The transformation from document to economic signal involves several stages. First, researchers need to collect and clean the relevant text, preserving important information such as publication dates, speakers, companies and topics. The quality of the final indicator depends heavily on whether this underlying text actually represents the economic phenomenon being studied.

Next comes classification. Older approaches frequently relied on dictionaries that label particular words as positive, negative, uncertain, hawkish or dovish, while machine-learning systems can learn more complicated linguistic relationships from training data. Modern transformer models and large language models can go further by analyzing sentences in context rather than treating every word independently.

Context matters because identical words can carry very different meanings. “Inflation remains high” and “the risk of high inflation has diminished” contain overlapping vocabulary but communicate different economic messages. The major advantage of more advanced language models is their ability to use surrounding words and sentence structure when determining what a statement actually signals.

The model’s output can then be converted into a numerical variable. Researchers might calculate the proportion of negative articles each day, average the hawkishness of central-bank statements or measure how frequently corporate executives discuss recession risk. Once those observations are ordered through time, they become a conventional dataset suitable for statistical analysis.

Econometrics then becomes essential. Researchers can test whether today’s news sentiment predicts tomorrow’s economic activity, whether a shift in central-bank language changes bond yields or whether increasing corporate uncertainty precedes lower investment. AI creates the measurement; econometrics tests whether that measurement has explanatory or predictive value.

A useful model might compare GDP forecasts with and without the text-derived variable. If forecast errors become consistently smaller when news sentiment is included, researchers have evidence that the textual information adds value beyond conventional indicators. Similar tests can examine inflation, unemployment, investment, market volatility or interest rates.

Causality requires even greater caution. Negative news may predict a recession because journalists detect deteriorating conditions early, but negative coverage could also influence consumer confidence and behavior. Both news sentiment and economic activity might additionally respond to a third event, such as an energy shock. A correlation between language and the economy does not by itself establish that the language caused the economic outcome.

The Biggest Opportunity—and Risk—Is Measuring What Was Previously Invisible

Text analysis expands the economist’s information set dramatically. Traditional datasets tell us how many people are unemployed or how quickly prices changed, while text can potentially reveal expectations, uncertainty, confidence, concern and changes in narrative before those concepts become visible in official statistics. This makes language particularly valuable for nowcasting and detecting turning points.

But textual indicators introduce new measurement problems. News organizations decide which events deserve coverage, executives choose what to emphasize on earnings calls and central bankers deliberately craft language to influence expectations. The text being analyzed is therefore not a random or perfectly neutral sample of economic reality.

Language itself also changes. A phrase interpreted as hawkish in one monetary-policy environment may carry a different implication in another, while terminology surrounding technology, inflation or financial risk evolves over time. A text indicator can suffer from structural breaks just as a conventional economic model can.

Model choice creates additional uncertainty. A dictionary-based sentiment system, a specialized financial-language model and a general-purpose large language model may assign different scores to the same paragraph. Researchers therefore need validation samples, robustness tests and transparent definitions showing exactly what the index is intended to measure.

There is also a danger of confusing precision with truth. An AI model might assign a central-bank speech a hawkishness score of 0.73, but the additional decimal places do not guarantee that the underlying concept has been measured accurately. Turning words into numbers makes language easier to analyze; it does not eliminate uncertainty about what those words mean.

Large language models nevertheless expand what researchers can attempt. The IMF’s 2025 central-bank project classified individual sentences across topic, stance, sentiment and audience using a fine-tuned LLM and produced a directional communication index capturing signals about future policy-rate changes and unconventional monetary measures. This illustrates how text analysis is moving from simple word counting toward multidimensional economic measurement.

The strongest applications will therefore combine AI with careful economic design. Researchers need to define the concept being measured, validate the model against human judgments or observable outcomes and then test whether the resulting variable adds explanatory or predictive information. The goal is not to replace economic statistics with language, but to extract additional signals from information that traditional datasets leave behind.

Economic analysis has traditionally been constrained by information that can be recorded conveniently as numbers. GDP, inflation, employment and interest rates fit naturally into spreadsheets and econometric models, while speeches, newspaper articles and corporate discussions have historically been much harder to quantify systematically. AI and natural language processing are changing that boundary. Language itself is becoming a measurable economic dataset.

Central-bank statements can be transformed into indicators of policy tone, inflation concerns and economic expectations. News articles can become daily measures of sentiment and uncertainty, while corporate reports and earnings calls can reveal changes in demand, investment, labor conditions and management expectations. These signals can then be compared with subsequent economic and financial outcomes.

The value is particularly significant when conventional statistics arrive slowly. GDP may be released quarterly and subsequently revised, while newspaper stories, speeches and earnings calls appear continuously. Textual data can provide an economic signal while official statistics are still being collected.

However, AI-generated indicators are not automatically reliable simply because they are numerical. Researchers must examine sampling, model choice, context, language changes and whether the resulting variable actually improves explanation or forecasting outside the data used to construct it. Econometric testing remains essential because a sophisticated language model can still produce an economically useless indicator.

The most important development is therefore not that AI can read millions of documents faster than economists. It is that information previously regarded as too qualitative, unstructured or extensive to analyze systematically can now be transformed into testable variables. AI converts words into measurements; economics determines whether those measurements tell us something meaningful about the real world.

That combination creates a powerful new form of economic analysis. Instead of waiting exclusively for the next official release, economists can observe how policymakers are speaking, what companies are discussing and how the tone of economic news is changing in near real time. When words become data, the economy can begin sending measurable signals before the traditional numbers arrive.