Enhancing Text-to-Image Generative AI Models for Clean Energy Technology Image Generation – UROP Symposium

Enhancing Text-to-Image Generative AI Models for Clean Energy Technology Image Generation

Ye Jin Son

Research Mentor: Majdi Radaideh
Mentor Department: Nuclear Engineering and Radiological Sciences, Engineering
Author(s): Majdi Radaideh
Session: Session 6 (3:00 PM – 3:50 PM)
Presentation Type: Poster 39

Abstract

Scientific pipelines for analyzing nuclear document imagery increasingly depend on large collections of figures automatically extracted from PDFs, which are used to support document understanding and downstream machine learning tasks. However, these collections are highly unorganized with low-quality scans, and non-target, irrelevant figures being able to enter the dataset which weakens the model performance and forces manual review. To close this gap between data quality and domain adaptation, we fine-tuned Stable Diffusion to generate domain-specific text-to-image outputs that better match the visual conventions of the target domain and can be used for augmentation when real examples are scarce. Then, we applied transfer learning with EfficientNet-B0 to train a binary classifier for filtering PDF-scraped nuclear images. This model separates target figures, such as 2D plots and diagrams, from non-target figures, improving the relevance and consistency of training data used by downstream models. To evaluate the model, the binary classifier achieved 97.9% precision for diagrams, 93.3% precision for plots, and 95.6% overall accuracy, indicating reliability for large-scale pipelines. By combining cleaner datasets with domain-specific synthetic image generation, this research enables faster experimentation and more robust downstream modeling, and it generalizes to other specialized domains where visual data derived from documents is abundant but difficult to curate at scale.

lsa logoum logo