Multilingual Metadata Sorting – UROP Spring Symposium 2024

Multilingual Metadata Sorting

Mrunmayee Jere

Pronouns: she/her

Research Mentor(s): Christi Merrill
Research Mentor School/College/Department: Comparative Literature/Asian L &C / LSA
Program:
Authors: Mrunmayee Jere, Yatin Bichala, Ivan Li, Mason Young, Christi Merrill, Ali Bolcakan
Session: Session 5: 2:40 pm – 3:30 pm
Poster: 95

Abstract

This research project addresses the prevalent issue of de-emphasis on non-Western English language works within online databases and library corpora, particularly focusing on the lack of comprehensive metadata surrounding these works. The project targets current and future researchers utilizing databases like HathiTrust, seeking to streamline search results and enhance the acknowledgment of non-Western authors. The proposed methodology entails leveraging Python to parse HathiFiles, followed by a meticulous sorting and categorization of similar works employing the Levenshtein distance metric. Emphasis is placed on evaluating the similarity between author and title metadata fields to pinpoint duplicate works effectively. The challenge of dealing with inaccurate or inconsistent metadata is tackled with the aid of the UM-GPT API. Ultimately, this project endeavors to create a “central list” containing comprehensive metadata for identical works, enhancing accessibility and accuracy for researchers, and addressing a longstanding gap in the representation of non-Western authors within online databases.

Arts and Humanities, Interdisciplinary

lsa logoum logo