Enhance the computational efficiency of physical simulators with CUDA – UROP Symposium

Enhance the computational efficiency of physical simulators with CUDA

Julian Whittaker

Research Mentor: Cheng Li
Mentor Department: Climate and Space Sciences and Engineering, Engineering
Author(s): Not Available
Session: Session 1 (9:00 AM – 9:50 AM)
Presentation Type: Poster 61

Abstract

As Graphics Processing Unit (GPU) adoption grows, GPU acceleration is becoming increasingly prevalent as a means of finding performance enhancements. Many legacy programs written in C or C++ could be significantly enhanced through the use of highly parallel processing such as that achievable through GPUs. These legacy programs often contain dynamic allocation or other programming patterns that do not translate well to the specialized GPU interfaces and architecture. This paper introduces a lightweight, fully concurrent solution to this dynamic allocation problem that allows for easy enhancement from CPU driven programs to GPU accelerated programs. We utilize per-warp shared memory as heap storage and allocate a linked data structure with minimal overhead to manage the heap. Across several benchmarks, our allocator achieves performance improvements of 40 to 100 times when compared to the default CUDA malloc. These results demonstrate that our scheme significantly outperforms existing GPU dynamic allocators while maintaining low overhead and ease of use, enabling GPU acceleration for legacy programs without requiring substantial engineering effort.

lsa logoum logo