China's Internet Corpus 4.0 Released for AI Development
China's Internet Corpus 4.0, containing 120GB of data, was officially released today. The dataset aims to support the training of large AI models and advance artificial intelligence development.
China's Internet Corpus 4.0, a dataset of 120 gigabytes, was officially released today in Jinan during a forum on AI security governance, a sub-event of the 2026 National Cybersecurity Week. The data is now available on the "Chinese Internet Corpus Resource Platform" for registered and authenticated users.
The initiative was led by the China Cyberspace Security Association in collaboration with the National Internet Emergency Center. Several prominent technology companies, including Baidu, iFlytek, and Zhihu, participated in the data collection and processing efforts. This latest version builds upon previous releases (1.0, 2.0, and 3.0) and is intended to provide a trusted data foundation for training large-scale AI models.
The release of the corpus is presented as a significant step in enhancing the supply of high-quality Chinese language data resources. According to a representative from the Cyberspace Security Association, the dataset enriches the domestic supply and supports the high-quality development of the AI industry.
The association intends to continue working with other organizations and companies to further develop the corpus. The goal is to continuously improve the quality and quantity of data to bolster the growth of the AI sector.