Publisher:ISCCAC
Weixiao Hu, Hou Shu, Sihan Zhou, Baokun Yang, Shangge Li, Yujia Wang, Wei Song
Hou Shu
August 23, 2026
Stele inscriptions, Character prediction, Pre-trained language models, GuwenBERT.
The identification of missing characters in stele inscriptions is a critical factor for digital reconstruction of ancient texts. However, pure automated prediction has a limited accuracy for complicated contexts, and the single output of candidate 1 cannot provide satisfactory practical judgment. In order to address this problem, this paper constructs an evaluation corpus of inscription recovery based on 20 texts to compare the candidate generation ability between general Chinese models and ancient Chinese domain models using random masking. Experiment results indicate that GuwenBERT performs best among in epigraphic scenarios (Top1 accuracy: 50.95%, Top5 coverage: 71.73%). On this basis, this paper proposes a Top5 candidate probability assisted judgment scheme for missing characters. This method will display candidate characters, along with its probability of the bar chart superimposed on the original text. The results of a controlled user experiment prove that its effectiveness, i.e., when the information is probability, the accuracy of user judgment is increased from 41.50% to 57.50%, and the subjective confidence of the judgment score has increased from 3.04 to 3.60, That is, it can effectively help the identification of defective stele characters text through the display of candidate probabilities.
© 2026, the Authors. Published by ISCCAC
This is an open access article distributed under the CC BY-NC license