Annice Chang
Research Mentor(s): Christi Merrill
Mentor Department: Comparative Literature/Asian L &C
Authors: Ali Bolcakan, Annice Chang, Amy Shi, Yuna Miyoshi, Ziqiao Wang, Kaeshav Krishna, Christi Merrill
Session: Session 4 (1:00pm – 1:50pm)
Presentation Type: Poster 110
Abstract
Rada Mihalcea In the past two decades, there have been great strides in the realms of text processing and analysis, machine learning and translation, and large language models. Still, working with historical and multilingual textual data remains a major challenge because of linguistic variations and related shortcomings in text recognition and parsing. Natural Language Processing (NLP) relies on high-accuracy Optical Character Recognition (OCR), which is particularly difficult with historical and multilingual datasets. Thus, we seek to do our experiments without relying on an understanding of the textual material but instead intend to identify reusable semi-stable anchor points for our comparative computational analyses. Taking Daniel Defoe’s English language novel Robinson Crusoe (1719) as our test case because of its significant impact and availability in translation, we look at a variety of Japanese, Mandarin, and Tamil translations from the late 19th and early 20th centuries, sourced through humanistic archival research. Our primary objective is to come up with reproducible computational frameworks, methods, and tools to analyze the distinctions between the source text and its translations, even with non-Latin script material in under-resourced languages. This foundational work should facilitate the discovery of additional resources and enable analyses of variations in tone, style, and narrative shifts across languages in historical textual data. This project seeks to contribute to both traditional and computational humanities, translation and reception studies, computational linguistics, and computer and information sciences.



