Skip to main navigation Skip to search Skip to main content

BrIM: A Branching In-Memory Accelerator

  • Stefan Maczynski
  • , Amlan Ganguly
  • , Mark Indovina
  • , Purab Sutradhar
  • , Sai Manoj Pudukotai Dinakarrao
  • , Sathwika Bavikadi
  • Rochester Institute of Technology
  • George Mason University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

While memory-centric (MC) AI accelerators have shown significant promise in AI computing by overcoming memory bottlenecks, these architectures still rely on traditional processors for some of the critical AI-computing tasks. Most notably, Tokenization, which is an integral part of processing Large Language Model (LLM) AI algorithms, involves decomposing the input data into smaller meaningful units (tokens), a task of a highly branching nature. Existing memory-centric systems, which typically do not support branching tasks, offload such workloads to separate general-purpose 'host' processors. However, such task offloading incurs inevitable back-and-forth movement of data between the external processor and the in-memory AI accelerator, leading to expected latency and energy overheads and a loss of performance. To alleviate such issues, we propose a programmable Look-up Table (LUT)-based in-DRAM processing architecture that efficiently processes branching kernels as well as data-parallel AI-oriented workloads so as to support Tokenization, along with the rest of the LLM workloads within the same memory chip. Our proposed solution offers remarkably superior performance optimization of tokenization compared to the traditional CPU+MC systems, with up to 8 × higher utilization of the compute bandwidth.

Original languageEnglish
Title of host publicationGLSVLSI 2025 - Proceedings of the Great Lakes Symposium on VLSI 2025
Pages982-989
Number of pages8
ISBN (Electronic)9798400714962
DOIs
StatePublished - 29 Jun 2025
Event35th Edition of the Great Lakes Symposium on VLSI 2025, GLSVLSI 2025 - New Orleans, United States
Duration: 30 Jun 20252 Jul 2025

Publication series

NameProceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI

Conference

Conference35th Edition of the Great Lakes Symposium on VLSI 2025, GLSVLSI 2025
Country/TerritoryUnited States
CityNew Orleans
Period30/06/252/07/25

Keywords

  • Large Language Model
  • Processing in Memory
  • Tokenization

Fingerprint

Dive into the research topics of 'BrIM: A Branching In-Memory Accelerator'. Together they form a unique fingerprint.

Cite this