Yutong Ai
Pronouns:
Research Mentor(s): Stefan Larson
Co-Presenter:
Research Mentor School/College/Department: SkySync / NonUM
Presentation Date: April 20
Presentation Type: Poster
Session: Session 5 – 3:40pm – 4:30 pm
Room: League Ballroom
Authors: Yutong Ai, Gordon Lim, Stefan Larson
Presenter: 77
Abstract
Automated document classification is a task that has wide-ranging use cases, especially in the industry where it can be used to apply labels to massive amounts of documents. Many contemporary models achieve high classification accuracy scores on the commonly-used RVL-CDIP benchmark dataset. The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes: letter, memo, email, file folder, form, handwritten, invoice, advertisement, budget, news article, presentation, scientific, publication, questionnaire, resume, scientific, report, specification. This dataset is publicly available. The RVL-CDIP corpus is the de facto standard benchmark for document classification, yet almost all studies that use this corpus do not include an evaluation of out-of-distribution documents. This paper reports on a work-in-progress and evaluates document classifiers trained on RVL-CDIP and tested on a new set of over 3000 out-of-distribution documents. We train several image-based document classifiers on the full RVL-CDIP dataset and then evaluate these models on our new out-of-distribution dataset. We find that standard image-based models (i.e., CNNs) struggle to differentiate between in-distribution RVL-CDIP test documents and documents from this new evaluation set. Our findings suggest that while image-based document classifiers may perform well on in-distribution inputs, they struggle to differentiate these from out-of-distribution documents.
Engineering



