Sara Haycox

Research Mentor(s): Stefan Larson
Research Mentor School/College/Department: DryvIQ
Presentation Date: 08/03/2022
Presentation Type: Poster
Poster Number: 51
Session: Session I: 12:30 – 1:20pm
Room: League Ballroom
Authors: Sara Haycox, Stefan Larson
Abstract
This research project is aiming to improve the accuracy of DryvIQ’s machine learning model. A machine learning (ML) model can only perform as well as the data it’s provided, which early on in the process can cause many false positives when detecting objects. To combat this, more data needs to be provided to the model and analyzed before it can perform better. The model can currently detect certain images and entities which is beneficial to companies handling private information such as social security cards, ID cards, and other documents with personal identifiable information or PII. Companies need to keep up with regulations surrounding PII, so this model is a way to detect such information. The goal is to develop an object detection tool which will be used to detect objects within an image. We’re currently working on improving the accuracy of detecting objects like social security cards, ID cards, fingerprints, and others that may be important for a company to identify. My part of this research has been working on model validation to ensure the accuracy of the models predictions. We hope to see that our tool is able to generate a large amount of images that could easily be mistaken as being a different object, so we can better understand how the model is performing and measure its false positive rate. It is challenging to test ML models, so this project introduces a new approach to generating challenging data to test and help evaluate the false-positive performance. Other ML researchers/practitioners could benefit from our approach as they could incorporate the same tactic to better understand how well their models are performing. ML models that aren’t calibrated effectively can potentially be dangerous, so this approach can help in detecting where the model may fail before it’s used to test real data.



