Five Questions About the Token Economy: Defining Its Boundaries

Deep News
Oct 06

On August 3, the Beijing Economic-Technological Development Area released the "Ten Measures for Tokens." Under the plan, by 2030, Beijing's Yizhuang district will build four ten-thousand-card token factories and create a "super gateway" that spans multiple models and computing power. Companies that call large models in real-world scenarios can receive subsidies of up to 50% based on token consumption. Two thousand kilometers away, Shantou has already completed a cross-border token transaction. An AI toy sold to Singapore received a user question. The question traveled back to Shantou through international communication links, a model deployed domestically completed the inference, and the answer was sent back to the overseas user. The entire service consumed about 100,000 tokens, priced at 2 yuan per million tokens, making the transaction worth roughly two jiao. Beijing is building factories and competing for gateways, while Shantou is selling an intelligent service worth two jiao. From grand industrial policy to a tiny cross-border order, a technical unit once hidden in the back end of models has suddenly stepped into the center of the economic stage. Guangdong has established a token trading and service center, and Shanghai has issued model and token vouchers. Local layouts for the token economy have moved from conceptual discussion into the policy and project phase. The token economy has truly arrived. Questions have followed: Is it a new economic form, or another round of conceptual packaging? To answer this question, five boundaries must first be drawn.

First question: What exactly is a token? A token is not new oil. Tokens run through corpus processing, model training, and inference services. Processed and rights-confirmed corpora may have asset attributes, but the asset object is the corpus and its usage rights. The current token economy mainly revolves around the inference side. Inference tokens are generated and consumed immediately with each call and do not possess independent asset attributes. Calling them "new oil" exaggerates the value of a unit of measurement. When a person inputs a sentence to a large model, the model first splits text, punctuation, numbers, and even image descriptions into fragments, then processes them one by one. These fragments are tokens. User input consumes tokens, and model-generated answers also consume tokens. Model companies usually charge based on input and output volume, making tokens the most common billing unit for large model services. But one Chinese character does not necessarily correspond to one token, and one token does not necessarily correspond to a complete word. The same sentence processed by different models may produce different numbers of tokens; consuming the same ten thousand tokens may yield vastly different answer quality across models. Tokens relate to resource consumption, yet they cannot directly show how many chips a model used, how much electricity it consumed, or how much value it created. Calling tokens "new oil" easily overestimates their asset attributes. Oil has relatively stable physical properties and can be stored, transferred, and held long term. Tokens are usually generated and consumed immediately during model calls and have no independent scarcity. More accurately, the token economy is installing an "electricity meter" on intelligent services: the API billing system measures, and the token count is the reading on the meter. This reading allows model services to be priced by volume, but the measurement standard is not yet unified. The same task may produce different token counts across models, and output quality may vary greatly. Tokens record processing scale; final value still depends on task results. This is the first boundary for understanding the token economy.

Second question: How do tokens become a business? Charging by token is not the same as buying and selling tokens. When enterprises pay by token, they are purchasing the service of a model completing a task. Tokens only record how much of that service was used. The so-called token economy is essentially the measurement, pricing, and settlement of intelligent services. There is no batch of tokens waiting to be bought, sold, and appreciated. Tokens were originally just a counting unit inside models. Each time a piece of input is processed and each time a response is generated, the system records the corresponding token count. Model companies then charge based on this number, turning a technical processing step into a settleable transaction. For a counting unit to support a business, three conditions are needed: the service can be split, the price can be calculated, and the result can be delivered. Large model APIs happen to fill these three conditions. In the past, when enterprises wanted to use artificial intelligence, they often had to purchase servers, deploy models, and build technical teams. Project cycles were long, and upfront investment was heavy. After the emergence of large model APIs, enterprises can buy intelligent services per use: pay once to identify a contract, pay once to generate a piece of code, pay once to complete a round of customer service dialogue. What enterprises buy is the model's ability to complete tasks. Tokens record how much of this service was used. Computing power, data, and models are thus packaged into the same bill. Enterprises do not need to know how many chips were used in the background or which data was called; they only need to compare calling prices, response speeds, and output quality. This is the foundation on which the token economy stands. An electricity meter does not create electricity, but it supports the electricity market. Tokens do not create intelligence, but they lower the transaction cost of intelligent services. They allow AI capabilities to be divided, quoted, and continuously purchased, and they shift model services from one-time software project delivery to on-demand basic services. Whoever controls the measurement standard, the calling gateway, and the settlement system may influence the rules of the next generation of intelligent infrastructure. Tokens themselves do not appreciate. Where they truly play a role is in helping intelligent services complete transactions. This is the second boundary.

Third question: Do more tokens mean greater value? Call volume is not value volume. Call volume only shows how much a model was used; it cannot directly show how many problems were solved. One hundred million tokens may complete a complex task, or they may be consumed in lengthy output and ineffective loops of an agent. This is currently the most easily confused account. The National Data Resource Survey Report (2025) shows that in 2025 China's token call volume was approximately 21,100 trillion, with average daily calls growing from more than 1 trillion at the beginning of the year to 100 trillion by year-end. In March this year, domestic average daily token calls further exceeded 140 trillion. Growth indeed exists, and at an astonishing speed. It at least shows that large models are moving from occasional trials to high-frequency calls. But call volume can only prove market activity; it cannot directly prove value growth. Imagine two companies processing the same ten thousand contracts. The first company consumes one hundred million tokens. The second company, through model optimization, consumes only twenty million tokens. If only call volume is considered, the first company is larger; if efficiency is considered, the second is clearly stronger. Models repeatedly producing nonsense, or agents falling into ineffective loops, will also generate large numbers of tokens. Consumption rises, the task is not completed, but the bill still grows. Local policies are amplifying this contradiction. Beijing Yizhuang is laying out along the entire industrial chain: building ten-thousand-card token factories upstream, creating a unified calling and settlement "super gateway" midstream, and stimulating application demand downstream through model vouchers and token vouchers. What Beijing is competing for is token supply capacity and service distribution gateways. Guangzhou Nansha has chosen to enter from market infrastructure. The Guangdong Token Trading and Service Center, built on the Guangzhou Data Exchange, plans to carry out measurement rules, compliance attestation, and trading services in three zones: inclusive, industry, and cross-border. What Guangzhou is competing for is measurement standards and trading order. Shanghai issues computing power vouchers, model and token vouchers, and corpus vouchers, focusing on reducing the cost for enterprises to call models. Yangzhou has launched token vouchers, cross-regional token operation centers, and computing-power-electricity coordination centers, hoping to connect local computing power, new energy, and the application needs of small and medium-sized enterprises. These local policies have different emphases: Beijing focuses on production and gateways, Guangzhou on standards and trading, Shanghai and Yangzhou on application consumption, and Shantou and Lingang on cross-border delivery. The real concern is that these routes may ultimately converge on the same indicator: how many tokens were called. Factories look at output, platforms look at traffic, localities look at growth rates, and enterprises look at subsidies. As long as policy rewards are directly tied to token consumption, lower model efficiency may actually bring more subsidies. One company consumes one hundred million tokens to complete a task, while another uses only twenty million tokens to complete the same task. Assessed by call volume, the former may receive higher evaluation; assessed by production efficiency, the answer is completely opposite. The token economy has established a quantity account, but the quality account remains blank. To evaluate token policy, the key is real revenue, task effectiveness, token efficiency, and energy consumption. Without these indicators, trillion-level call volume may be nothing more than statistical prosperity. This is the third boundary: call volume represents usage scale and cannot alone represent economic value.

Fourth question: What exactly is sold when tokens go overseas? Tokens going overseas is not the same as computing power going overseas. Tokens going overseas and computing power going overseas are at different levels of the same industrial chain. Computing power going overseas mainly delivers computational resources, while tokens going overseas deliver pay-per-use model inference services. The former provides production capacity; the latter provides task results processed by models. Tokens going overseas are often described as an upgraded version of computing power going overseas: computing power going overseas sells "shovels," while tokens going overseas sell "gold." This phrase is vivid enough, but the boundary must be drawn more precisely. Computing power and tokens are on the same industrial chain. Chips, servers, and data centers provide underlying computing capacity; models process computing power into inference services; tokens handle measurement and settlement. When chips are deployed domestically, overseas customers initiate requests through APIs, and domestic models complete inference and return results, what is delivered across borders is model service. Tokens serve as the billing unit for this service. Such transactions are closer to exports of pay-per-use AI inference services. Compared with directly exporting hardware, this can use domestic chips, electricity, and models to deliver higher value-added services to overseas customers, and it can also reduce dependence on cross-border hardware transportation. Shantou's advantage comes from industrial scenarios. The city has a vast toy industry belt and a commercial network connecting overseas markets. After AI toys are sold overseas, voice dialogue, story generation, and companionship services still require continuous calls to domestic models. Traditional toys can only generate one-time product revenue. After connecting to large models, each conversation may bring new service revenue. In other words, one toy completes one export, but the intelligent service can be sold many times. As of July this year, token services disclosed by Shantou have reached a daily average of tens of billions, forming multiple application scenarios such as AI toys and cross-border e-commerce. Shanghai Lingang has taken another path. It relies on international data processing pilots, cross-border communication links, and model service platforms to explore the complete process of "overseas demand access, domestic model inference, and cross-border result delivery." Together, the two places have verified one thing: China's model capabilities can enter overseas markets in the form of digital services. But this path does not automatically eliminate geopolitical and compliance risks. The risks have merely changed position. Hardware exports face chip and equipment controls; model service exports also face cross-border data, content security, intellectual property, algorithm transparency, service responsibility, tax attribution, and cross-border settlement. Where data comes from, where inference is completed, who is responsible for generated content, and in which region revenue is recognized all require rules to answer. Charging by token will not automatically create a wholly new category of international trade. What specific type of digital service a transaction belongs to still depends on the contractual relationship, data flow, and actual delivery method. Therefore, whether tokens going overseas can form an industry cannot be judged only by call counts. It also depends on whether there are real customers, sustained revenue, stable repeat purchases, and affordable compliance costs. That two-jiao order in Shantou first proved that the technology, communication, and payment links can run through. To form a mature industry, two jiao must be turned into long-term orders, and demonstration projects must become replicable business models. This is the fourth boundary: tokens going overseas export pay-per-use model services, and tokens are merely the settlement unit of that service.

Fifth question: What rules does the token economy need? A rule gap is not a legal vacuum. The token economy is constrained by existing rules on copyright, personal information, data security, generated content, and market competition. The current gaps are mainly in measurement standards, quality evaluation, call volume auditing, capital constraints, and cross-border responsibility linkage. The token economy is not in a legal vacuum. Training data is constrained by copyright, personal information, and data security rules; model-generated content is constrained by AI service and content governance rules; cross-border calls must also comply with data outbound, network communication, and international trade systems. What is truly lacking at this stage are market rules suitable for token services. First, measurement rules must be transparent. How platforms calculate input and output tokens, whether cached calls are charged, and how tool calls and long texts are priced should all be verifiable by enterprises. Token counts across different models cannot be directly compared, and so-called "unified measurement" cannot simply stipulate how much one token is worth. What regulation needs to unify is billing disclosure, verification methods, audit processes, and dispute resolution rules. When government subsidies and public procurement are involved, independent audits should also be introduced to prevent fabricated call volume,无效 traffic, and fund extraction through related-party transactions. Second, quality evaluation must keep up. One token from different models may carry vastly different amounts of intelligence. The market needs to evaluate accuracy, completion rate, response speed, stability, and unit task cost in specific tasks. In June this year, the country's first Token Service Performance Monitoring Platform was released, beginning to monitor indicators such as call success rate, first-token latency, and output speed. This shows that the industry's focus has begun shifting from call quantity to service quality. But technical performance is only the first step. A model answering quickly does not mean it answers correctly; a low token price does not mean the total cost of completing a task is low. Ultimately, it still depends on whether the model solved the problem. Third, green costs must be accounted for. The longer a model's output, the more computing power and energy it usually consumes. When localities build token factories, they cannot assess only output; they must also disclose energy consumption per token, energy consumption per task, and the proportion of clean energy used. If a model creates ten times the token consumption with one time the task value, such growth can hardly be called high-quality development. Fourth, responsibility must also be traceable. A service may simultaneously involve overseas customers, domestic models, third-party data, cloud platforms, and payment institutions. Once data leakage, infringing output, or service interruption occurs, the responsibility chain, tax attribution, and dispute resolution mechanism must be clarified in advance. There is another easily overlooked issue: measurement power may become new platform power. If a platform simultaneously controls the model, interface, measurement rules, and settlement gateway, it is both player and referee. Future regulatory priorities should cover not only price but also interface openness, billing transparency, data portability, and fair market competition. Fifth, capital also needs rules. Data centers, model optimization, industry applications, and cross-border services all require long-term investment, and capital should of course enter. What it should most support is high-quality data, underlying technology, and applications that can solve real problems. It is also most prone to one mistake: directly converting token output into future revenue, and then packaging call volume into valuation. How many cards a token factory has and how many tokens it can generate each year only show supply capacity. Without paying customers and stable tasks, even huge output is just idle capacity. What is even more necessary to guard against is a kind of capital idling: the government issues token vouchers, enterprises use the vouchers to purchase model calls, the platform inflates call volume, and then uses that call volume to seek financing and the next round of subsidies. Every link has growth numbers, but at the end there are no truly paying market customers. Therefore, when capital values token companies, it should at least look at real revenue, paying customers, repurchase rates, task completion rates, and unit task costs. Enterprises should also separately disclose fiscal subsidies, free calls, related-party transactions, and market-based revenue, avoiding putting traffic of different natures into the same growth curve. Government guiding funds and market funds must also have a clear division of labor. Fiscal funds are better suited to investing in basic capabilities such as measurement standards, public evaluation, green computing power, and cross-border compliance, while market funds bear the risks of model commercialization and application competition. If localities compete to invest in similar token factories, they may ultimately repeat the problems of duplicated data center construction and low utilization rates. A line must also be drawn for financial impulses. Tokens are the billing unit of model services and cannot be packaged into quasi-financial products that can be hoarded, speculated on, and promised appreciation. Pre-sales, pledges, or high-yield financing using future token output as a gimmick also need strict constraints. This is the final boundary: the token economy needs new rules, focused on measurement, evaluation, auditing, capital constraints, and responsibility linkage, that is, managing the five accounts of quantity, quality, cost, responsibility, and funds, without creating an entirely new legal system from scratch. The token economy deserves attention because intelligent services are entering a stage of large-scale transactions. It also deserves caution. Over-mythologizing will package a unit of measurement as an asset, while blind denial will miss the real changes taking place in how intelligent services are traded. Tokens can count intelligence, but they cannot price intelligence on their own. The first half of the token economy focused on capacity, price, and call scale. Competition in the second half will revolve around a more difficult question: who can use fewer tokens to complete more complex tasks and create higher, verifiable value. The author of this article is Li Enhan, director of the Token Digital Economy Research Center at the China Development Institute (Shenzhen) and a postdoctoral fellow in economics.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10