Full text loading...
, Brett Hashimoto2
and Veronika Laippala1
Abstract
The utility of historical language databases is hindered by their complexity, text variety, and lack of register information. We investigate the feasibility of deriving linguistically motivated and reliable register predictions for unannotated, long historical texts. We fine-tune BERT-based deep learning models using register-annotated data from the Corpus of Founding Era American English and predict registers for different text parts of unannotated Eighteenth Century English Online documents. We determine the model’s effectiveness in capturing pervasive linguistic features across different text sections (e.g. beginnings vs. endings), analyzing how document internal variation affects model performance and identifying which sections are most effectively predicted. Additionally, we employ the Stable Attribution Class Explanation method to extract and compare keywords from various text parts to determine the quality of the predictions. Our findings indicate that text beginnings consistently provide more reliable classifications.
Article metrics loading...
Full text loading...
References
Data & Media loading...