Projects in Machine Learning for Document Processing – UROP Spring Symposium 2023

Projects in Machine Learning for Document Processing

Junjie Shen

Junjie Shen photo

Pronouns: He/him

Research Mentor(s): Stefan Larson
Research Mentor School/College/Department: DryvIQ / NonUM
Program: UROP
Session: Session 6 (3:40pm – 4:30pm)
Authors: Jiayou Shen, John Buckley

Abstract

Machine learning has revolutionized the way we process and analyze large amounts of data, including text. Document processing is one of the fields that has benefited greatly from the advancements in machine learning, providing efficient and effective solutions for tasks such as text classification, entity recognition, and sentiment analysis. With the growing volume of written material available in digital form, machine learning algorithms can be used to automatically extract information and insights from documents, enabling organizations to make data-driven decisions. In this project, a taxonomy of industry categories that has broad enough coverage to capture a reasonably wide/large amount of industry types was developed by reviewing several already existing taxonomy and examining their coverage and accuracy.

Engineering

lsa logoum logo