数字经济时代,规模化的在线平台发展日益成熟。对于以快手、YouTube等为代表的用户生成内容(User Generated Content, UGC)平台,数以亿计的参与用户、海量的内容推荐候选集和异步的信息传输方式使得识别可能流行的内容并在运营中加以应用对于平台而言意义重大。本文对UGC流行度预测问题进行深入讨论并从两方面对现有机器学习模型提出改进:第一,在模型中融入用户间在线社交网络的视角,即充分挖掘平台社交功能,在特征工程中对内容创作者的粉丝网络结构特点进行提炼;第二,在工程中融合聚类思想,即对创作者上传内容在一定时间内的流行度变化趋势进行时间序列聚类,而后据由算法识别出的不同模式分别建立预测模型。基于某短视频平台脱敏抽样后的88万位用户的活动数据进行实证检验,结果显示相较于基线模型,考虑用户在线社交网络规模、紧密程度等结构特征的实验模型预测精度更高,在测试集上的AUC达到0.76。其次,相比基于全体样本构建的整体模型,对创作者进行聚类模式识别后分别建立的子模型具有更好的整体性能和正例识别能力:对于聚类算法识别出的缓慢波动上升型和剧烈波动上升型两种典型作者,子模型的AUC、准确率和召回率相比整体模型分别平均增长7.07%、60.17%和15.62%。 此外,与一些黑箱式的机器学习模型相比,本文所提出的预测模型结论具备可解释性和指导用户生成内容平台运营的实际意义。借由SHAP(Shapley Additive Explanations)值的方法,本文发现在线社交网络的结构对于网络中用户所生成内容的流行度增长具有重要的直接影响,且对于平台通过推荐主动增加内容曝光和流行度这一过程具有正向的调节效应。据此本文提出,在舆情控制、商业营销等场景中,除传统的流量分配策略外,平台可以通过算法对用户网络中的关键创作者和紧密团簇进行识别,从而有针对地进行流量倾斜,使曝光手段的有效性进一步提升。 本文的研究具备一定的理论创新和实践意义。一方面,为社交网络视角对UGC流行度预测问题的贡献提供了实证证据,同时揭示了针对复杂场景进行聚类模式识别的重要性;另一方面,为平台运营损益估计、创作环境优化乃至数字经济发展和智慧公共决策均提供了有价值的参考。
In the era of digital economy, online platforms have grown into considerably giant sizes. For user generated content (UGC) platforms like Kuaishou and YouTube which have billions of users and recommend contents to users from an unlimited selection in an asynchronous way, identifying contents that will gain great popularities is essential to improving user engagement and community health. In this work, we conduct an extensive analysis of UGC popularity prediction and complement previous machine learning models from two aspects. First, we enhance the model from a new perspective of user network, that is to say, build fans’ network for content uploaders, extract features about network structure, and embed them into the machine-learned classifier. Second, apply time-series clustering to recognize different patterns for popularity trend, so as to identify various types of uploaders and build classifiers for them respectively. Based on a sample from short-video platform which contains 880 thousand users’ activity data, we find that compared to baseline models, adding social-network-structure feature group to classifier can dramatically increase the prediction accuracy, which achieves an AUC of 0.76 on test set. Furthermore, compared to the overall model trained from the whole sample, the sub-models built on clustering results present better performance and can also identify popular contents more accurately, which achieves average increases of 7.07%, 60.17%, and 15.62% on AUC, precision, and recall respectively. In contrast to black box classifiers, the predicting model we propose is explainable and can shed light on maintaining such UGC platforms. By calculating SHAP (Shapley Additive Explanations) value, we find that the structure of online social network shows direct influence on the popularity of contents generated by users within this network, and plays as a positive moderator when platform tries to promote the contents. From this point of view, in scenarios such as public opinion control and online marketing, instead of simply giving more exposures, platforms can target at key uploaders and user groups through algorithms to improve the efficiency of promoting strategies. This paper shows certain theoretical innovation and practical value. On the one hand, it provides empirical evidence that the perspective of online social network contributes to UGC prediction problem, and shows how clustering and pattern recognition can lift model performance in a complex context. On the other hand, not only does it give a valuable reference for platform to estimate operating cost/benefit and cultivate an optimal environment for UGC creation, but it could be applied to digital economy development and smart public decision-making as well.