加载中...
When Behavioral Data Betray Users: A Diagnostic and Protective Framework Against Social Interaction Leakages
数字平台通常公开分享行为数据(如评分和评论)以提升服务质量,然而这些看似无害的数据也可能泄露用户之间隐藏的社会关系。本文提出一个两阶段框架:首先通过从行为数据中推断潜在的社会互动来诊断这种泄露风险,然后在保持数据效用的同时降低风险。我们刻画了隐藏关系何时可从观测行为中识别,并利用路易斯安那州和宾夕法尼亚州的公开Yelp数据表明,攻击者在10%的假阳性率下可恢复约一半的真实社会关系,在20%的假阳性率下可恢复超过60%。这些推断出的关系会显著增加网络风险:当用于鱼叉式钓鱼时,估计的攻击回报率在500次假冒尝试的活动中为109%,在10,000次尝试的活动中升至1,098%。为减少此类泄露,我们提出两种具有正式差分隐私保证的扰动机制:一种直接向发布的动作矩阵添加高斯噪声,另一种在重新合成发布数据前向学习到的表示添加拉普拉斯噪声。两种机制均降低了链接推断准确率并大幅降低估计的攻击者回报;在我们的主要保护设置下,较小规模活动的回报变为负值,较大规模时也远低于未保护情形。总体而言,本文为诊断社会互动泄露以及评估平台分享行为数据时的隐私-效用权衡提供了实用框架。
Abstract
Digital platforms often share publicly visible behavioral data, such as ratings and reviews, to improve service quality. Yet, these seemingly innocuous data can also reveal hidden social ties among users. We develop a two-stage framework that first diagnoses this leakage risk by inferring latent social interactions from behavioral data and then mitigates the risk while preserving data utility. We characterize when hidden ties are identifiable from observed actions and show, using public Yelp data from Louisiana and Pennsylvania, that an attacker can recover about half of true social ties at a 10% false-positive rate and more than 60% at a 20% false-positive rate. These inferred ties can materially increase cyber risk: when used for spear phishing, the estimated return on attack rises from 109% for a campaign with 500 impersonation attempts to 1,098% for one with 10,000 attempts. To reduce this leakage, we propose two perturbation mechanisms with formal differential privacy guarantees on the released representations: one adds Gaussian noise directly to the released action matrix, and the other adds Laplace noise to the learned representation before resynthesizing the released data. Both mechanisms reduce the link-inference accuracy and substantially lower estimated attacker returns; under our main protection settings, returns become negative for smaller campaigns and remain much lower at larger scales. Overall, the paper provides a practical framework for diagnosing social interaction leakage and evaluating the privacy-utility trade-off when platforms share behavioral data.
行为数据可泄露隐藏社交关系
利用Yelp数据,在10%假阳性率下,HNL算法分别恢复LA和PA地区49.4%和48.9%的真实社交连接;在20%假阳性率下恢复率超过60%。这表明看似无害的评论长度等行为数据足以揭示用户的私人社交关系,构成严重的社会隐私风险。
社交关系泄露大幅提升钓鱼攻击回报
基于推断出的社交关系进行鱼叉式钓鱼,攻击回报率(ROA)随攻击规模增加而显著上升。在宾州数据中,500次假冒尝试时ROA为109%,10,000次时升至1,098%;在路易斯安那州,ROA从500次时的52%升至10,000次时的984%,凸显了泄露的经济激励。
高斯扰动机制有效降低链接推断精度
直接机制向发布的词频矩阵添加高斯噪声。当扰动强度τ_G=0.05时,AUC在LA和PA分别下降0.18,全矩阵75分位绝对词数变化约17-18词。该机制在保持数据一定效用的同时,显著削弱了攻击者恢复社交连接的能力。
拉普拉斯扰动机制更有效抑制攻击收益
间接机制向学习到的网络和倾向添加拉普拉斯噪声再重合成数据。在PA数据中,AUC降低最多达0.26。在较小规模攻击(如500次)下,可使ROA降至负值(如PA中间接机制为-80%),有效抑制攻击者的经济动机。
核心解释变量
观测到的用户行为数据(如Yelp评论长度构成的用户-商家动作矩阵A)
被解释变量
潜在的社会互动网络G(用户间的朋友关系)及内在行为倾向B
样本与数据
来自Yelp公开数据集的2020年数据,包括路易斯安那州(新奥尔良)的2,065名用户和6,094家商家,以及宾夕法尼亚州(匹兹堡)的2,234名用户和15,965家商家。用户需至少撰写一条评论且属于最大连通分量。
识别方法 / 模型设定
将用户行为建模为线性二次博弈的纳什均衡,其中行为受社会互动影响。通过求解一个双凸优化问题(同质性网络学习,HNL),交替更新网络结构G和内在倾向B,以最小化观测动作与均衡预测的差异。
内生性及稳健性检验
在Yelp数据上进行了广泛的实验,包括与相关性、图Lasso、变分图自编码器等多种基线方法的对比,并评估了不同隐私预算和扰动强度下的隐私-效用权衡。结果在不同州的数据集上表现一致,验证了方法的有效性和稳健性。
更多相关数据正在补充