关 键 词 :跨文化传播;冒犯性评价;受众差异;平台治理;内容审核;计算传播学科分类:新闻学与传播学--传播学
全球平台常以统一模型处理不同文化背景受众的内容评价,但同一表达在不同地区和个体之间可能获得截然不同的冒犯性评分。现有研究多关注文本特征,而对受众地区与个体价值的影响,以及模型在未见地区中的表现均讨论不足。本研究基于 D3CODE[1] 数据集的 4554 条英文评论、 4309 名标注者和 144730 条有效评分,通过消融比较、留一地区外验证、误差切片与 SHAP 分析,考察文本、地区和道德价值对冒犯性评价预测的影响。研究发现,受众地区带来的预测改善高于个体价值;同时纳入地区、价值与预设交互关系的模型表现最佳,但个体价值并未降低八个地区的宏平均误差。评论分歧越大,模型预测误差越高,宗教、社会群体、性别及关系规范等内容也更易出现偏差。研究表明,冒犯性评价并非仅由文本决定,统一内容审核模型难以充分覆盖跨文化受众差异。平台除报告总体性能外,还应分别检验不同地区、议题与分歧水平下的模型表现。
Global platforms often rely on unified models to moderate content for culturally diverse audiences, yet the same expression may receive sharply different offensiveness ratings across regions and individuals. Existing research emphasizes textual features, paying less attention to regional background, individual values, and model performance in unseen regions. Drawing on 4,554 English comments, 4,309 annotators, and 144,730 valid ratings in D3CODE, this study uses ablation analysis, leave-one-region-out validation, error slicing, and SHAP to examine how text, region, and moral values contribute to predicting offensiveness ratings. Regional information yields greater predictive gains than individual moral values. The model combining region, values, and prespecified interactions performs best, but individual values do not reduce macro-averaged error across the eight regions. Prediction error rises with rating disagreement and is higher for content concerning religion, social groups, and gender and relationship norms. Offensiveness judgments therefore cannot be inferred from text alone, and unified moderation models do not fully capture cross-cultural audience variation. Platforms should report performance separately by region, topic, and level of disagreement, alongside aggregate metrics.