基于密度Canopy的评论文本主题识别方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Topic recognition method of comment text based on density Canopy
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    融合Sentence-BERT和LDA的评论文本主题识别(SBERT-LDA)方法,将LDA的主题数作为K-means算法中的k值,导致算法可解释性较差、主题一致性较低。为了解决上述问题,提出基于密度Canopy的SBERT-LDA优化方法(SBERT-LDA-DC),利用密度Canopy改进K-means算法。实验结果表明,提出的方法在一致性指标上要优于使用K-means以及K-means++对特征向量聚类的同类方法;与SBERT-LDA方法相比,在1 852条戏剧评论数据集上,一致性指标值提高了22.9%。因此,所提出的SBERT-LDA-DC方法是有效的,对产品或服务提供者更好地了解用户意见、完善自身产品或提升服务水平提供了新方法,具有较强的实际应用价值。

    Abstract:

    The method, which combines Sentence-BERT and LDA, takes the topic number of LDA as the k value in K-means algorithm, resulting in poor interpretability and low topic consistency. To solve this problem, a Sentence-BERT and LDA optimization method based on density Canopy(SBERT-LDA-DC) was proposed, which used density Canopy to improve the K-means algorithm. The experimental results indicate that this method is superior to similar methods using K-means and K-means++ to cluster feature vectors on the consistency index. Compared with the SBERT-LDA method, the consistency index is improved by 229% on the 1 852 drama comment dataset. The proposed SBERT-LDA-DC method is effective, which provides a new method for product or service providers to better understand user opinions and improve their own products or services, and has strong practical application value.

    参考文献
    相似文献
    引证文献
引用本文

刘 滨,詹世源,刘 宇,雷晓雨,杨雨宽,陈伯轩.基于密度Canopy的评论文本主题识别方法[J].河北科技大学学报,2023,44(5):493-501

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2023-09-04
  • 最后修改日期:2023-10-08
  • 录用日期:
  • 在线发布日期: 2023-11-02
  • 出版日期:
文章二维码