Projects in Machine Learning for Document Processing – UROP Spring Symposium 2023

Projects in Machine Learning for Document Processing

Temi Okotore

Temi Okotore photo

Pronouns: she/her

Research Mentor(s): Stefan Larson
Research Mentor School/College/Department: DryvIQ / NonUM
Program: UROPF
Session: Session 6 (3:40pm – 4:30pm)
Authors: Temi Okotore , Stefan Larson

Abstract

In today’s world, even with all the various resources available, there still seems to be no large data set that is available to use in training machine models for document classification. The goal of this project is to develop a taxonomy of industry categories to cover a large range of industry types. We are using data collection through search engines to create a large-scale set of data. After reviewing pre-existing industry taxonomy we started data collecting documents in other industries and its company representatives. In conducting our research, we have found a substantial number of various taxonomies for large scale industries and are using large-scale web-search for constructing the data set. We hope to write a paper with the results of our research that can act as an open dataset to use and improve on for the task of document industry classification. There are not yet results for the research because we are currently in the data collection stage, but it is anticipated that we will be able to create a classifying system. In conclusion, the project is a resource for companies and organizations to have to organize their documents, providing an accessible way to find documents.

Life Science

lsa logoum logo