<p>Authorship attribution is often treated as a problem of maximizing classification accuracy, yet high accuracy alone does not show whether a model has captured authorial style rather than topic, plot, or semantic similarity. This study returns to the evidentiary basis of attribution through a multilingual, idiolect-based method centered on the frequencies of commonly used words and characters. Using Chinese, English, and Japanese literary corpora, we test whether these high-frequency features can represent habitual writing style across languages. The first experiment evaluates pairwise classification among single works by different authors; the second tests attribution across multiple works by the same authors with varied content and genre. A lightweight multilayer artificial neural network is used as the main classifier, alongside simpler machine-learning models. The proposed features achieve high performance, with average final-test accuracy above 98% across languages. SHAP analysis shows that discriminative evidence is concentrated in function words, particles, pronouns, and punctuation rather than content-specific vocabulary. These findings demonstrate the value of stylistic scholarship for authorship research: carefully selected, theoretically meaningful stylistic features can support high attribution accuracy while reducing dependence on technically demanding, resource-intensive models. Reliable attribution therefore depends less on model complexity than on interpretable representations of authorial habit.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Idiolect-based textual features as evidence for authorship attribution

  • Jianjun Shi,
  • Jingyin Tian

摘要

Authorship attribution is often treated as a problem of maximizing classification accuracy, yet high accuracy alone does not show whether a model has captured authorial style rather than topic, plot, or semantic similarity. This study returns to the evidentiary basis of attribution through a multilingual, idiolect-based method centered on the frequencies of commonly used words and characters. Using Chinese, English, and Japanese literary corpora, we test whether these high-frequency features can represent habitual writing style across languages. The first experiment evaluates pairwise classification among single works by different authors; the second tests attribution across multiple works by the same authors with varied content and genre. A lightweight multilayer artificial neural network is used as the main classifier, alongside simpler machine-learning models. The proposed features achieve high performance, with average final-test accuracy above 98% across languages. SHAP analysis shows that discriminative evidence is concentrated in function words, particles, pronouns, and punctuation rather than content-specific vocabulary. These findings demonstrate the value of stylistic scholarship for authorship research: carefully selected, theoretically meaningful stylistic features can support high attribution accuracy while reducing dependence on technically demanding, resource-intensive models. Reliable attribution therefore depends less on model complexity than on interpretable representations of authorial habit.