加载中...
Guided topic model based product defect information detection method and applications
如何从日益增多的线上投诉中挖掘出潜在的产品缺陷信息已成为制造商和监管部门质量管理工作的重点和难点。为此,本研究提出了一个全新的基于主题模型的产品缺陷信息识别模型。该模型首先利用引导式主题模型的优势,通过缺陷词典事先确定缺陷主题数量,提取出与行业标准内容一致的缺陷主题,并根据这些主题对投诉进行多标签分类;另外,针对投诉附带信息的非结构化特点以及投诉内容缺陷词的稀疏性分布规律,模型融入了产品特征建模以及自适应性词加权两种改进机制,进一步提升了缺陷识别的性能。本研究收集了国内汽车权威平台115,668条线上投诉进行实验,实验结果表明所提模型相比于潜在狄利克雷分配等现有主题模型能生成与行业的实际缺陷分类一致的缺陷主题,还能分析产品的品牌、能源类型等特征对投诉中缺陷主题概率分配的影响。模型在F1指数、AUC指数、汉明损失、排序损失上明显优于对比的其他主题模型。模型的产品特征建模、自适应性词加权和引导式主题模型结构确保了本模型在各评价指标上的精度表现。该研究可帮助制造商和监管部门及时监管产品质量水平,保护消费者合法权益。
Abstract
Extracting potential product defect information from the increasingly abundant online complaints has become a focal point and challenge in quality management for both manufacturers and regulatory agencies. To this end, this study proposes a novel product defect information detection model based on topic modeling. Firstly, the model takes advantage of the guided topic model to determine the number of defect topics through a defect dictionary, extract defect topics consistent with industry standards, and perform multi-label classification of complaints based on these topics. In addition, the model incorporates two improvement mechanisms, namely product feature modeling and adaptive term weighting, to further enhance defect recognition performance by addressing the unstructured nature of complaint information and the sparse distribution of defect words. This study collected 115,668 online complaints from a leading domestic automotive platform for experimentation. Experimental results indicate that compared to existing topic models such as latent Dirichlet allocation, the proposed model can generate defect topics consistent with actual industry defect classifications. Additionally, it can analyze the influence of product brand, energy type, and other features on the probability distribution of defect topics in complaints. The model outperforms other comparative topic models significantly in F1 score, AUC score, Hamming loss, and ranking loss. The product feature modeling, adaptive term weighting, and guided topic model structure of the model ensure its accuracy in all evaluation metrics. This research can help manufacturers and regulatory agencies monitor product quality levels in a timely manner and protect the legitimate rights and interests of consumers.
DDGTM在缺陷识别各项指标上优于对比主题模型
在通过交叉验证确定分类阈值的机制下,DDGTM的F1 micro为73.53%,F1 macro为66.78%,AUC micro为78.33%,AUC macro为85.88%,Hamming loss为22.76%,Ranking loss为19.06%,各项指标均优于GLDA、LLDA、DLDA等对比模型。
自适应性词加权和产品特征均提升缺陷识别性能
将DDGTM与不考虑自适应性词加权(NW)、不考虑产品特征(ND)以及同时不考虑两者(NDW)的变体模型进行对比,DDGTM各项指标均最佳。NW和ND性能依次下降,NDW性能最差,说明自适应性词加权和产品特征建模均对缺陷识别有贡献,且词加权的影响更大。
DDGTM与SVM结合时缺陷识别效果依然最佳
将DDGTM与SVM结合进行缺陷识别,F1 micro为76.37%,F1 macro为74.73%,AUC micro为81.41%,AUC macro为83.69%,Hamming loss为15.29%,Ranking loss为32.68%,各项指标均优于LDA、STM等无监督主题模型与SVM结合的效果。
不同产品特征的投诉缺陷主题分配存在显著差异
以投诉品牌和能源类型为例,上汽通用别克的投诉最关注发动机系统(影响值0.14),东风标致最关注安全系统(影响值0.15);燃油动力产品投诉主要关注制动系统、安全系统和附件,纯电动产品投诉主要关注电气系统、发动机系统、传动系统和悬挂系统,混合动力产品投诉主要关注发动机系统和附件。
DDGTM可提取与行业标准一致的缺陷主题并识别隐含问题
DDGTM根据缺陷词典事先确定11个主题(10个缺陷主题+1个非质量缺陷主题),前10个主题分别对应制动系统、电气系统、发动机系统、安全系统、结构与车体、转向系统、传动系统、悬挂系统、灯光系统和附件,与行业缺陷分类标准一致,且可通过主题-词分布中低概率词识别隐含问题。
核心解释变量
投诉文本内容(问题简述、投诉内容、典型问题)、产品特征向量(品牌、品牌属性、国别、车型属性、发动机、变速箱、能源类型、年份)
被解释变量
投诉的缺陷主题标签(11类:10个缺陷部件类别+缺陷无关)
样本与数据
车质网2010至2020年消费者投诉共223,319条,涉及323个汽车品牌;选取投诉量前10品牌共115,668条投诉作为实验数据,构建722组MMY文档,随机分为433组训练集(60%)、144组验证集(20%)、145组测试集(20%)
识别方法 / 模型设定
提出DDGTM(缺陷识别引导式主题模型),包含四层贝叶斯结构(文档-段落-主题-词),融合GLDA引导式结构、自适应性词加权(基于信息增益)和产品特征建模(DMR),运用EM算法进行参数推断,E步通过吉布斯采样(引入CRP)更新主题分配,M步通过L-BFGS优化超参数
内生性及稳健性检验
设计对比实验:(a)DDGTM、(b)不考虑自适应性词加权的DDGTM(NW)、(c)不考虑产品特征的DDGTM(ND)、(d)同时不考虑产品特征和词加权的DDGTM(NDW)与(e)GLDA、(f)LLDA、(g)DLDA、(h)LDA、(i)STM进行对比;采用F1 micro、F1 macro、AUC micro、AUC macro、Hamming loss、Ranking loss六项指标;通过十折交叉验证确定分类阈值μ;设置两种缺陷识别机制(阈值机制和SVM结合机制)进行验证
更多相关数据正在补充