HOME>RESEARCH>ECONOMICS

Opportunities, challenges for economic statistics in digital–intelligent era

Source:Chinese Social Sciences Today 2026-09-21

Big data and AI are reshaping economic statistics in the digital-intelligent era, enabling more intelligent, real-time, and data-driven measurement. Image generated by AI

As a core tool for observing, describing, and analyzing economic activity, economic statistics faces unprecedented systemic disruption and an urgent need for transformation in the digital–intelligent era. Traditional economic statistics rely heavily on structured, standardized survey data, but long collection cycles, high costs, and limited coverage make such data poorly suited to the rapid evolution of emerging activities such as the platform economy, online consumption, and digital-asset trading. Economic big data is vast in scale and highly diverse in both source and form. Meanwhile, artificial intelligence (AI) technologies have further enhanced the ability to extract information from unstructured data such as text, images, audio, and video, construct indicators, and even generate simulated data. It is therefore necessary to systematically examine the profound implications of big data and AI for economic statistics and, accordingly, develop forward-looking theoretical perspectives and practical approaches.

Innovation in economic measurement paradigm

Economic measurement is an indispensable foundation for empirical economic research and plays a crucial role in advancing economic theory. The integration of big data and AI is driving a transformation of the economic measurement paradigm, giving rise to four major shifts in economic measurement and statistical oversight.

First, economic statistical indicators are moving toward higher-frequency, real-time measurement. The high velocity of big data, together with AI’s powerful capabilities, is driving traditional economic indicators toward higher frequency and greater real-time availability. Ubiquitous sensors, including IoT devices, GPS systems, and smart meters; continuously operating online platforms, including e-commerce, social media, and search engines; and automated collection systems can generate massive amounts of data almost continuously. AI enables this information to be processed immediately and its value extracted. Economic indicators can thus evolve from low-frequency and lagged measures into high-frequency, continuous, and even real-time “dynamic images” of economic activity. This transformation has three important benefits. High-frequency and real-time indicators can reflect economic conditions and changing trends in near real time, providing timely evidence for policymaking. They can also capture dynamic lead–lag relationships more precisely and expand the scope of economic research.

Second, unstructured data enable the measurement of subjective factors. The widespread application of big data is making multidimensional statistical measurement of economic and social activity increasingly feasible, extending measurement beyond conventional, observable indicators such as GDP and CPI to less tangible dimensions, including social psychology and social welfare. In the digital–intelligent era, vast amounts of unstructured information, particularly text, can be used to construct subjective indices. With the development of technologies such as Latent Dirichlet Allocation (LDA) and generative large language models (LLMs), natural language processing techniques applied to textual information have become an important means of extracting insights into the psychological states of societies and groups, enabling the construction of more timely and broadly representative subjective indices. Compared with psychological indicators based on traditional statistical surveys, such as market expectations and confidence indices, big-data-based sentiment analysis can be conducted at high frequency and in near real time while offering broader coverage. It can therefore effectively address limitations such as sample-selection bias and low-frequency measurement lags. Such approaches have already been widely studied in areas including investor sentiment and economic policy uncertainty.

Third, big data and AI enable economic measurement based on estimation and prediction. Traditional empirical economics relies primarily on econometric models built on restrictive assumptions such as linearity, variable independence, and normality. Yet the complexity of real-world economic systems, including nonlinear relationships and high-dimensional interactions among variables, makes it difficult for conventional econometric models to capture the complex dynamics of economic change. The integration of big data and machine learning provides a more flexible, data-driven measurement paradigm that can substantially reduce estimation and prediction errors. AI, particularly deep learning, is not tied to a prespecified model form, allowing for much greater flexibility in measurement models, reducing model-specification error, and enabling them to better capture complex economic structures. The rapid development of generative LLMs, particularly their powerful capabilities in pattern recognition, predictive generation, and modeling complex relationships, is creating opportunities to transform economic measurement in multiple areas of empirical research. By constructing LLM-based simulated economic systems, researchers can also test the predictive performance of different economic theories in virtual environments, thereby identifying their boundary conditions and contexts in which they apply.

Fourth, big data and AI are enabling new forms of statistical oversight. Statistical oversight is a fundamental institutional safeguard for ensuring data authenticity and maintaining public trust in government. Traditional statistical oversight generally relies on ex post audits and sample-based inspections, which have limited coverage and timeliness. Big data and AI are moving statistical oversight toward a new stage characterized by routine, intelligent monitoring. By integrating information from multiple sources—including government departments such as taxation, customs, electricity, and finance, as well as online platforms and sensors—an intelligent anomaly-detection and early-warning system can be established. When reported statistical data deviate substantially from signals revealed by other sources within this system, an alert can be issued automatically. Generative large models can analyze vast amounts of historical statistical information and policy texts to establish a “statistical baseline” for what constitutes a reasonable range, helping determine whether newly reported data conform to expected patterns.

Limitations of observational data

The widespread application of big data and AI creates new opportunities for economic statistics. However, economic big data is fundamentally observational in nature and therefore remains vulnerable to sample-selection bias. Moreover, because it is observational, the statistical relationships it reveals cannot be equated with causal relationships. Causal inference thus remains highly challenging.

Economic big data constitutes a passive record of socioeconomic activity in digital space. Its generation depends on predetermined conditions of observation, including levels of technological development, platform rules, and patterns of user behavior, as well as the prevalence and usage characteristics of digital devices. This non-random recording process creates sample-selection bias. Digital divides commonly exist across regions, social groups, and generations, creating structural gaps in the socioeconomic picture captured by big data. As a result, big data may fail to represent the population as a whole, particularly the actual conditions facing vulnerable groups. In addition, patterns of human behavior make the generation of big data inherently non-random. Even when sample sizes are extremely large, user-behavior data generated through voluntary participation or passive recording may exhibit substantial non-randomness.

Because big data is observational and its generation is not controlled by researchers, confounding factors cannot be effectively controlled for or isolated as they can in randomized controlled experiments. Statistical relationships derived from big data may consequently be influenced by confounding variables rather than represent genuine causal effects. Even rigorous quasi-experimental methods depend on a series of critical assumptions. When these assumptions are not satisfied, the validity of causal-inference conclusions is substantially weakened. Many mainstream machine-learning methods remain grounded in statistical theory, including linear regression, logistic regression, decision trees, naive Bayes classifiers, and support vector machines. Deep neural networks, often regarded as “black boxes,” can perform well in many prediction tasks, but their internal decision-making processes are often opaque, making it difficult for researchers to clearly characterize relationships among variables.

Directions for economic statistics in digital–intelligent era

In response to the opportunities and challenges created by big data and AI, economic statistics require systematic innovation in their theoretical foundations, methods and technologies, indicator systems, measurement of the digital economy, and evidence-based policy evaluation.

First, a system of high-frequency, real-time economic statistical indicators should be developed. As the pace of economic and social activity accelerates, such a system has become essential to meeting increasingly rapid decision-making needs and strengthening economic governance. The key is to break free from the periodic constraints of traditional statistical surveys and move toward a new statistical paradigm based on a closed loop of “data–algorithm–indicator–validation.” This requires integrating multiple high-frequency sources, including platform transaction logs, search indices, and sensor information; developing new statistical methods for extracting and processing traditional sources of information, such as tax and logistics records, at higher frequencies; and creating models and algorithms capable of transforming real-time big data—which is often unstructured and noisy—into economically interpretable indicators. More importantly, dynamic optimization mechanisms are needed so that indicator models can adapt to changes in economic structure.

Second, traditional economic statistical indicators need to be transformed and upgraded to keep pace with changing economic and social realities and provide an accurate and comprehensive picture of their complexity. This requires reforming indicators whose representativeness has been weakened by historical inertia and correcting measurement biases embedded in existing indicators.

Third, disaggregated economic statistical indicators should be developed. A fundamental challenge for economic statistics is to capture both the overall development and future trajectory of a country or economy and the heterogeneity within it, thereby enabling governments to formulate more targeted policies and improve the welfare of disadvantaged groups. Given China’s large population, vast territory, substantial economic scale, and diverse development conditions, promoting balanced economic and social development requires not only aggregate statistics but also a system of disaggregated statistics. Such statistics can reveal the relationships among different groups and the heterogeneous characteristics concealed by aggregate indicators. Disaggregation primarily involves economic statistics broken down by region, industry, firm size, income level, urban–rural status, gender, age, and other dimensions.

Fourth, the measurement of data as a factor of production and as an asset needs to become more precise. Traditional economic accounting systems often broadly classify digital technologies and data inputs as “ICT capital” or intangible investment, obscuring the important contribution of data as a new factor of production. Digital-economy statistics therefore need to measure data as both a factor of production and an asset more precisely to fully capture its contribution to economic growth. This cannot be achieved through simple adjustments to accounting or statistical techniques. Rather, it requires fundamental innovation in economic theory and a transformation of the statistical measurement paradigm.

Fifth, evidence-based policy evaluation should be strengthened through big data and AI. Accurate identification of the causal relationship between policies and outcomes is essential for formulating targeted and effective policies. Big data and AI can strengthen causal inference and evidence-based policy evaluation. Methods such as Causal Random Forests and Bayesian Additive Regression Trees can substantially improve the effectiveness and precision of causal inference and policy evaluation. LLMs can also analyze large volumes of textual data to identify potential confounding variables, instrumental variables, and mediating variables, and their use in causal inference is attracting increasing attention.

Sixth, data ethics and governance systems should be strengthened. The development of economic statistics in the digital–intelligent era requires robust systems of data ethics and governance, including standards for data privacy protection, mechanisms for assessing algorithmic fairness, and clearly defined boundaries for data ownership and usage rights.

 

Hong Yongmiao is a professor from the School of Economics and Management at the University of Chinese Academy of Sciences. Luo Liangqing is a professor from the School of Statistics and Data Science at Jiangxi University of Finance and Economics. Zhang Ming is a lecturer from the School of Economics and Management at Inner Mongolia University. This article has been edited and excerpted from Statistical Research, Issue 3, 2026.

 

 

 

Editor:Yu Hui

Copyright©2023 CSSN All Rights Reserved

Copyright©2023 CSSN All Rights Reserved