Data Collection and Machine Learning for Document Processing – UROP Spring Symposium 2023

Data Collection and Machine Learning for Document Processing

Nicole Cornehl Lima

Nicole Cornehl Lima photo

Pronouns: she/her

Research Mentor(s): Stefan Larson
Research Mentor School/College/Department: DryvIQ / NonUM
Program: UROPF
Session: Session 6 (3:40pm – 4:30pm)
Authors: Nicole Cornehl Lima, Stefan Larson

Abstract

With no large scale, free, and public data set to train and evaluate machine learning models for the task of document industry classification, a research group collected data to construct such a dataset. The need for this type of data set comes from anyone aiming to perform some type of document industry classification using machine learning. Data is the fuel for machine learning, so, firstly, data needs to be collected and correctly labeled in order to be efficiently used. Thus a lot of the work put into this project was gathering documents and accurately sorting them into different types of industries. For example, a doctor’s resume would be labeled in the healthcare industry, or a job posting for a butcher in the meat industry. Then that collected data can be used to train models or test algorithms. Once the models are trained they can be evaluated and the quality of data can be analyzed. This can be difficult as there is no metric for this, but a good proxy for the quality is model accuracy on the data set. These efforts resulted in an accessible data set that can be used to train models and algorithms with the purpose of industry classification.

Engineering

lsa logoum logo