TY - GEN
T1 - BrIM
T2 - 35th Edition of the Great Lakes Symposium on VLSI 2025, GLSVLSI 2025
AU - Maczynski, Stefan
AU - Ganguly, Amlan
AU - Indovina, Mark
AU - Sutradhar, Purab
AU - Manoj Pudukotai Dinakarrao, Sai
AU - Bavikadi, Sathwika
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/6/29
Y1 - 2025/6/29
N2 - While memory-centric (MC) AI accelerators have shown significant promise in AI computing by overcoming memory bottlenecks, these architectures still rely on traditional processors for some of the critical AI-computing tasks. Most notably, Tokenization, which is an integral part of processing Large Language Model (LLM) AI algorithms, involves decomposing the input data into smaller meaningful units (tokens), a task of a highly branching nature. Existing memory-centric systems, which typically do not support branching tasks, offload such workloads to separate general-purpose 'host' processors. However, such task offloading incurs inevitable back-and-forth movement of data between the external processor and the in-memory AI accelerator, leading to expected latency and energy overheads and a loss of performance. To alleviate such issues, we propose a programmable Look-up Table (LUT)-based in-DRAM processing architecture that efficiently processes branching kernels as well as data-parallel AI-oriented workloads so as to support Tokenization, along with the rest of the LLM workloads within the same memory chip. Our proposed solution offers remarkably superior performance optimization of tokenization compared to the traditional CPU+MC systems, with up to 8 × higher utilization of the compute bandwidth.
AB - While memory-centric (MC) AI accelerators have shown significant promise in AI computing by overcoming memory bottlenecks, these architectures still rely on traditional processors for some of the critical AI-computing tasks. Most notably, Tokenization, which is an integral part of processing Large Language Model (LLM) AI algorithms, involves decomposing the input data into smaller meaningful units (tokens), a task of a highly branching nature. Existing memory-centric systems, which typically do not support branching tasks, offload such workloads to separate general-purpose 'host' processors. However, such task offloading incurs inevitable back-and-forth movement of data between the external processor and the in-memory AI accelerator, leading to expected latency and energy overheads and a loss of performance. To alleviate such issues, we propose a programmable Look-up Table (LUT)-based in-DRAM processing architecture that efficiently processes branching kernels as well as data-parallel AI-oriented workloads so as to support Tokenization, along with the rest of the LLM workloads within the same memory chip. Our proposed solution offers remarkably superior performance optimization of tokenization compared to the traditional CPU+MC systems, with up to 8 × higher utilization of the compute bandwidth.
KW - Large Language Model
KW - Processing in Memory
KW - Tokenization
UR - https://www.scopus.com/pages/publications/105017674856
U2 - 10.1145/3716368.3735168
DO - 10.1145/3716368.3735168
M3 - Conference contribution
AN - SCOPUS:105017674856
T3 - Proceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI
SP - 982
EP - 989
BT - GLSVLSI 2025 - Proceedings of the Great Lakes Symposium on VLSI 2025
Y2 - 30 June 2025 through 2 July 2025
ER -