Yahoo Releases Record Machine Learning Dataset

Thirteen Terabytes of anonymized user-news item interaction data has been made available for developers to use in machine learning applications. This is the largest ever set of data to be made available for general use. It began life as user-news interaction data, collected by recording the user-news item interactions of about 20 million Yahoo users from February 2015 to May 2015. The dataset contains around 100 billion events. The Yahoo news Feed dataset was drawn from the news feeds of several Yahoo properties, including the Yahoo homepage, Yahoo News, Yahoo Sports, Yahoo Finance, Yahoo Movies, and Yahoo Real Estate. Writing about the dataset, Suju Rajan of Yahoo Labs said: “Our goals are to promote independent research in the fields of large-scale machine learning and recommender systems, and to help level…


Link to Full Article: Yahoo Releases Record Machine Learning Dataset