
The most extensive dataset for training recommendation systems has been created in Russia. European researchers utilized this dataset to devise a novel methodology that dramatically accelerates AI training—by factors of ten—without sacrificing outcome quality. This information was shared with TASS by the press service of “Yandex.”
In their statement, the authors of the research explain that for a considerable period, the research community faced restricted access to large-scale, industry-level datasets. “Yandex” inaugurated Yambda, becoming a pioneer in bridging this gap by furnishing a distinctive instrument capable of fueling a global advance in this domain.
In the early summer of 2025, “Yandex” specialists engineered and publicly released one of the world’s largest collections of data intended for advancing recommendation engines. The complete iteration contains five billion items, having been constructed from anonymized data sourced from “Yandex.Music.” This set encompasses aggregated listening statistics, endorsements, rejections, and various attributes pertaining to the musical recordings.
This training data repository was recently leveraged by academics from the University of Amsterdam to develop an innovative technique for training recommendation systems built upon the SEATER algorithm, originally conceived by Chinese researchers. This piece of software facilitates the organization of all items or tracks into a distinct, hierarchical catalog, similar to a folder structure on a computer.
In principle, such a catalog enables the system to generate suggestions more swiftly and accurately; however, its training demands a highly significant investment of time. In deployed applications, this constraint impedes the frequent refreshment of recommendations and hinders rapid adaptation to shifts in user taste. Experts from the Netherlands formulated two substitute methods that enable the activation of catalog preparation, which they subsequently validated using “Yandex’s” data.
The validation process demonstrated that one of the new algorithms managed to reduce the data preparation duration from 82 minutes down to just 83 seconds—a speed increase of nearly 60 times. This modification had a negligible effect on the quality of the recommendations, meaning the product developed by the Chinese specialists remains regarded as the industry leader for such systems. “Yandex” experts highlight that the entire code base for the enhanced SEATER model has been made publicly available, clearly illustrating the benefits derived from publishing and employing substantial datasets in the development and instruction of artificial intelligence.