Research: AI Hardware

AI Hardware

(Return to Current Research Outline)

Our research focuses on developing specialized systems that efficiently support the growing demands of artificial intelligence workloads. As AI algorithms become more complex and data-intensive, we recognize the challenges that traditional hardware architectures face in terms of performance, power consumption, and scalability. To address these limitations, we explore innovative hardware approaches and technologies that make AI processing faster, more efficient, and secure. Our work spans across Compute-In-Memory and Compute-Near-Memory architectures, neuromimetic devices, co-design, AI-driven secure hardware, and AI for hardware design, each contributing to the next generation of AI systems.

A. Compute-In-Memory and Compute-Near-Memory (CIM and CNM)

The increasing disparity between computation and data movement energy in modern technology nodes, often referred to as the memory wall problem, has made data movement a significant performance bottleneck. To address this challenge, we focus on integrating computation directly within or near memory arrays, through approaches like Compute-In-Memory (CIM) and Compute-Near-Memory (CNM). These methods reduce the latency and energy costs associated with transferring data between separate memory and processing units, enabling more efficient and faster AI processing.

Publications:
  1. Gaurav Kumar K, Yatharth Agarwal and Kaushik Roy. "HyFPCiM: A 65-nm 417-?W Error-Sensitivity-Aware FP8 Compute-in-Memory Macro." IEEE Solid-State Circuits Letters (SSCL), 2026.

    Abstract: This letter presents HyFPCiM, a 65-nm FP8 compute-in-memory (CiM) macro that enables sub-mW floating-point (FP) inference using error-sensitivity-aware FP partitioning (EAP). EAP maps exponent processing to a digital CiM (DCiM) path and mantissa accumulation to an analog CiM (ACiM), avoiding the power- and area-intensive adder-tree-based accumulation used in prior FP-CiM designs. The proposed DCiM employs dual-rail inverter-based sensing for exponent processing, while the ACiM path performs preanalog-to-digital conversion (ADC) charge-domain accumulation using switched-capacitor (SC) circuits. By amortizing one ADC conversion across eight multiply-and-accumulate (MAC) units, the proposed ACiM reduces ADC energy by 8× , ADC + transimpedance amplifier (TIA) power by 75%, and total macro power by 31% versus ADC-first designs. Fabricated in TSMC 65-nm, HyFPCiM consumes 417 ? W at 1 V and 50 MHz, achieves 40 GFLOPS/W, and incurs < 0.55% accuracy loss across multiple classification workloads. To the best of our knowledge, HyFPCiM achieves the lowest reported power among silicon-measured FP-CiM macros.

  2. Mukherjee, Mainakh, Ayan B. Pranta, Utkarsh Saxena, Anushka Mukherjee, Deepika Sharma, Gaurav Kumar K, Kaushik Roy. "MIRAGE:MRAM-Based Near ADC-Less Compute-In-Memory Macro for Deep Learning Acceleration." Design, Automation & Test in Europe Conference (DATE), 2026.

    Abstract: Non-volatile memory (NVM) based Compute-in-Memory (CiM) architectures have emerged as a promising compute primitive for accelerating deep neural networks (DNNs) by performing in-situ matrix?vector multiplications (MVMs). Among various NVMs, STT-MRAM (Spin Transfer Torque based Magnetoresistive Random Access Memory) shows potential due to its high endurance, low energy consumption and high density. However, existing STT-MRAM CiM designs typically rely on multi-bit analog-to-digital converters (ADCs) at the peripherals to digitize accumulated bit-line currents. While enabling high-precision computation, ADCs add substantial energy, latency, and area overheads. To alleviate such problems, we propose a system-technology co-design approach to a Near ADC-Less CiM design with ternary partial-sums called MIRAGE. The accuracy is maintained by considering hardware level partial sum quantization in the training loop. Specifically, we develop an STT-MRAM based CiM macro which features differential bitcells and an adaptive threshold sensing that is amenable to the requirements posed by ternary partial-sum quantization. We do a thorough energy, area, latency, and sense margin analysis along with robust benchmarking against conventional 1T-1MTJ (1 transistor-1 Magnetoresistive Tunnel Junction) based MRAM CiM. The proposed CiM macro occupies ? 20% less area, consumes 1.8× less MVM energy and shows 5× better latency with improved distinguishability compared to 1T-1MTJ CiM macro while achieving better accuracy.

  3. Holla, Amod, Sumedh Chatterjee, Sutanu Sen, Anushka Mukherjee, Fernando Garcia-Redondo, Dwaipayan Biswas, Francesca Iacopi, Kaushik Roy. "LIMO: Low-power in-memory-annealer and matrix-multiplication primitive for edge computing." npj Unconventional Computing, 2026.

    Abstract: Combinatorial optimization (CO) underpins critical applications in science and engineering, ranging from logistics to electronic design automation. A classic example of CO is the NP-complete Traveling Salesman Problem (TSP). Finding exact solutions for large-scale TSP instances remains computationally intractable; on von Neumann architectures, such solvers are constrained by the memory wall, incurring compute-memory traffic that grows with instance size. Metaheuristics, such as simulated annealing implemented on compute-in-memory (CiM) architectures, offer a way to mitigate the von Neumann bottleneck. This is accomplished by performing in-memory optimization cycles to rapidly find approximate solutions for TSP instances. Yet this approach suffers from degrading solution quality as instance size increases, owing to inefficient state-space exploration. To address this, we present LIMO, a programmable mixed-signal computational macro that implements an in-memory annealing algorithm with reduced search-space complexity. The annealing process is aided by the stochastic switching of spin-transfer-torque magnetic-tunnel-junctions (STT-MTJs) to escape local minima. For large instances, our macro co-design is complemented by a refinement-based divide-and-conquer algorithm amenable to parallel optimization in a spatial architecture. Consequently, our system comprising several LIMO macros achieves superior solution quality and faster time-to-solution on instances up to 85,900 cities compared to prior hardware annealers. The modularity of our annealing peripherals allows the LIMO macro to be reused for other applications, such as vector-matrix multiplications (VMMs). This enables our architecture to support neural network inference. As an illustration, we show image classification and face detection with software-comparable accuracy, while achieving lower latency and energy consumption than baseline CiM architectures.

  4. Holla, Amod, Mainakh Mukherjee, Anushka Mukherjee, Kaushik Roy. "ROSETTA: ROM-Overlaid STT-MRAM for Efficient MVM and Softmax Operations Toward Accelerating Transformer Inference." IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS), 2026.

    Abstract: Deep learning architectures are increasingly limited by the memory wall ? the ?von Neumann bottleneck? that forces frequent data transfers between memory and processing units, throttling throughput and energy efficiency. To alleviate this bottleneck, researchers have explored Compute-in-Memory (CiM) architectures to accelerate the core computing primitives such as the matrix?vector multiplications (MVMs), within memory arrays. Yet, large language models based on transformers also require many non-linear softmax operations, creating a new bottleneck. From a technology perspective, non-volatile STT-MRAMs are promising for CiM thanks to their endurance, non-volatility, and area footprint. However, conventional 1T?1R STT-MRAM CiM arrays place bit-cells in parallel, yielding lower MVM energy efficiency than 8T-SRAM CiM due to low device resistance; line parasitics and low OFF/ON ratios of STT-MRAM devices further limit row parallelism and scalability. In this context, we propose ROSETTA, a CiM array based on 3T?2R STT-MRAM bit-cells that uses series-resistance sensing and time-to-digital conversion (TDC) for low-power MVMs. An added word line also stores a ROM bit, enabling in-array tables for fast softmax, without increasing bit-cell area. We introduce a novel palindromic encoding of inputs and weights to mitigate the data-dependent nonlinearity in STT-MRAM CiM arrays based on TDC sensing. This enables 8× higher row-level parallelism than standard 1T?1R MRAM CiM and facilitates scaling to larger CiM arrays. Our macro occupies ~47% less area and consumes 3.1× less MVM energy than an equivalent ROM-overlaid 8T-SRAM macro, at similar latency and accuracy. At the system level, a spatial architecture comprising our macros exhibits 11.7?14.5× and 1.44?1.5× higher energy efficiency than equivalent architectures comprising 1T-1R MRAM and 8T-SRAM ROM-overlaid CiM macros, respectively.

  5. Negi, Shubham, and Kaushik Roy. "HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference." arxiv, 2025.

    Abstract: The rapid adoption of Large Language Models (LLMs) has driven a growing demand for efficient inference, particularly in latency-sensitive applications such as chatbots and personalized assistants. Unlike traditional deep neural networks, LLM inference proceeds in two distinct phases: the prefill phase, which processes the full input sequence in parallel, and the decode phase, which generates tokens sequentially. These phases exhibit highly diverse compute and memory requirements, which makes accelerator design particularly challenging. Prior works have primarily been optimized for high-batch inference or evaluated only short input context lengths, leaving the low-batch and long context regime, which is critical for interactive applications, largely underexplored.

    We propose HALO, a heterogeneous memory centric accelerator designed for these unique challenges of prefill and decode phases in low-batch LLM inference. HALO integrates HBM based Compute-in-DRAM (CiD) with an on-chip analog Compute-in-Memory (CiM), co-packaged using 2.5D integration. To further improve the hardware utilization, we introduce a phase-aware mapping strategy that adapts to the distinct demands of the prefill and decode phases. Compute bound operations in the prefill phase are mapped to CiM to exploit its high throughput matrix multiplication capability, while memory-bound operations in the decode phase are executed on CiD to benefit from reduced data movement within DRAM. Additionally, we present an analysis of the performance tradeoffs of LLMs under two architectural extremes: a fully CiD and a fully on-chip analog CiM design to highlight the need for a heterogeneous design. We evaluate HALO on LLaMA-2 7B and Qwen3 8B models. Our experimental results show that LLMs mapped to HALO achieve up to 18x geometric mean speedup over AttAcc, an attention-optimized mapping and 2.5x over CENT, a fully CiD based mapping.

  6. Sharma, Deepika, Shubham Negi, Trishit Dutta, Amogh Agarwal, and Kaushik Roy. "A 65nm 5 TOPS/W Digital CIM Accelerator with Reconfigurable Precision and Temporal Pipelining for Spiking Neural Networks." IEEE European Solid-State Electronics Research Conference (ESSERC), 2025.

    Abstract: Spiking Neural Networks (SNNs) have the inherent ability to process highly sparse event data from dynamic vision sensors (DVS). However, many existing SNN accelerators support only fixed network architectures, limited data precision, and fail to efficiently handle membrane potential dynamics and varying input sparsity. This work presents SpiDR, a reconfigurable digital compute-in-memory (CIM) SNN accelerator that addresses these limitations through: (1) fused weight-Vmem macros for reduced data movement, (2) staggered layout and reconfigurable peripherals for variable precision support, (3) zero-skipping to exploit unstructured input sparsity, and (4) asynchronous pipelining to maintain throughput despite variable computation times. Fabricated in TSMC's 65 nm CMOS technology, SpiDR achieves up to 5 TOPS/W at 95 % input sparsity and supports diverse event-based workloads including gesture recognition and optical flow estimation. SpiDR achieves 2× to 1000× higher area efficiency and up to 7.6× higher energy efficiency compared to prior designs.

  7. Kim, Dong Eun, Tanvi Sharma, Anushka Mukherjee, Mainakh Mukherjee, and Kaushik Roy. "MemRaptor: Magnetoresistive Array as Matrix Vector Multiplication and Transcendental Function Operator for NLP Applications." IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), 2025.

    Abstract: Compute-In-Memory (CiM) is emerging as a promising paradigm to design energy-efficient hardware accelerators for AI, addressing the processor-memory data transfer bottleneck. The popularity of CiM can be attributed to their ability to perform massively parallel in-situ matrix vector multiplications (MVMs), the dominant computation in neural networks (NNs). However, NNs used in NLP applications such as Long-Short Term Memory (LSTM) and transformers also frequently perform other operations such as transcendental functions (tanh, sigmoid, and softmax). To that effect, we present MemRaptor, utilizing CiM with magnetoresistive random access memory (MRAM) technology, that can perform both MVM and transcendental functions in the same memory array. MemRaptor overlays a read only memory (ROM) on an MRAM array through hard-wiring the connection of bit-cell with an additional bitline (a bit-cell connected to either of the bitlines but not to both), incurring no array area overhead and a minimal peripheral area overhead. Note, the bitline connection of bit-cell stores the ROM value while the magnetic tunnel junction (MTJ) in the bit-cell stores the RAM data. Particularly, the magnetization state of the 1T-1MTJ bit-cell in the array stores the weight value (RAM data) of the neural network, and the bitline connection of the bit-cell stores the look-up table (ROM data) used for computing transcendental functions. We demonstrate the working of our proposed design through circuit-level simulations for a 64×64 array, using a compact model of CoFeB/MgO PMA MTJ with 120% tunnelling magnetoresistance, 5k? RON , in 65nm technology. Further, we showcase the advantage of MemRaptor over standard MVM-based CiM accelerator architecture, PUMA, through comprehensive system-level evaluations for LSTM, BERT, and GPT models. Our results show up to 30% and 5.3% improvements in terms of throughput and energy-efficiency, respectively, on an average across different workloads.

  8. Yoo, Sangmin, Amod Holla, Sourav Sanyal, Dong Eun Kim, Francesca Iacopi, Dwaipayan Biswas, James Myers, and Kaushik Roy. "TAXI: Traveling Salesman Problem Accelerator with X-bar-based Ising Macros Powered by SOT-MRAMs and Hierarchical Clustering." ACM/IEEE Design Automation Conference (DAC), 2025.

    Abstract: Ising solvers with hierarchical clustering have shown promise for large-scale Traveling Salesman Problems (TSPs), in terms of latency and energy. However, most of these methods still face unacceptable quality degradation as the problem size increases beyond a certain extent. Additionally, their hardwareagnostic adoptions limit their ability to fully exploit available hardware resources. In this work, we introduce TAXI ? an inmemory computing-based TSP accelerator with crossbar(Xbar)-based Ising macros. Each macro independently solves a TSP subproblem, obtained by hierarchical clustering, without the need for any off-macro data movement, leading to massive parallelism. Within the macro, Spin-Orbit-Torque (SOT) devices serve as compact energy-efficient random number generators enabling rapid ?natural annealing?. By leveraging hardware-algorithm co-design, TAXI offers improvements in solution quality, speed, and energy-efficiency on TSPs up to 85,900 cities (the largest TSPLIB instance). TAXI produces solutions that are only 22% and 20% longer than the Concorde solver?s exact solution on 33,810 and 85,900 city TSPs, respectively. TAXI outperforms a current state-of-the-art clustering-based Ising solver, being 8× faster on average across 20 benchmark problems from TSPLib.

  9. Sharma, Tanvi, Mustafa Ali, Indranil Chakraborty, and Kaushik Roy. "What, When, Where to Compute-in-Memory for Efficient Matrix Multiplication During Machine Learning Inference." IEEE Transactions on Emerging Topics in Computing, 2025.

    Abstract: Matrix multiplication is the dominant computation during Machine Learning (ML) inference. To efficiently perform such multiplication operations, Compute-in-memory (CiM) paradigms have emerged as a highly energy efficient solution. However, integrating compute in memory poses key questions, such as 1) What type of CiM to use: Given a multitude of CiM design characteristics, determining their suitability from architecture perspective is needed. 2) When to use CiM: ML inference includes workloads with a variety of memory and compute requirements, making it difficult to identify when CiM is more beneficial than standard processing cores. 3) Where to integrate CiM: Each memory level has different bandwidth and capacity, creating different data reuse opportunities for CiM integration. To answer such questions regarding on-chip CiM integration for accelerating ML workloads, we use an analytical architecture-evaluation methodology with tailored mapping algorithm. The mapping algorithm aims to achieve highest weight reuse and reduced data movements for a given CiM prototype and workload. Our analysis considers the integration of CiM prototypes into the cache levels of a tensor-core-like architecture, and shows that CiM integrated memory improves energy efficiency by up to 3.4× and throughput by up to 15.6× compared to established baseline with INT-8 precision. We believe the proposed work provides insights into what type of CiM to use, and when and where to optimally integrate it in the cache hierarchy for efficient matrix multiplication.

  10. Sharma, Tanvi, Indranil Chakraborty, Mustafa Ali, and Kaushik Roy. "Evaluating Compute in Memory Architectures for Matrix Multiplication: A Dataflow-Centric Perspective." IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2025.

    Abstract: Compute in memory (CIM) is a promising technique to reduce data movement costs in traditional hardware by efficiently performing in-situ matrix multiplication, the dominant computation during deep learning (DL) inference. However, the broader question of how CIM architectures compare to tensor-core-like architectures remains largely unexplored. In this work, we take a dataflow-centric approach utilizing classic parameters such as compute latency, bandwidth, capacity and compute/memory access costs to determine throughput and energy consumption of a given CIM architecture. To that effect, we perform an iso-area comparison of tensorcore-like (or PE array) architecture with different CIM integrated architectures for matrix multiplication kernels. Our results demonstrate that CIM integrated memory can improve energy efficiency by up to 3.5× and throughput by up to 11× compared to tensorcore baseline, considering INT8 precision.

  11. Negi, Shubham, Utkarsh Saxena, Deepika Sharma, and Kaushik Roy. "HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads." arXiv preprint arXiv:2403.13577.

    Abstract: Analog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNN). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from crossbars leads to substantial power and area overhead. Moreover, the high area overhead of ADCs constrains the throughput due to the limited number of ADCs that can be integrated per crossbar. An approach to mitigate this issue involves the adoption of extreme low-precision quantization (binary or ternary) for partial sums. Training based on such an approach eliminates the need for ADCs. While this strategy effectively reduces ADC costs, it introduces the challenge of managing numerous floating-point scale factors, which are trainable parameters like DNN weights. These scale factors must be multiplied with the binary or ternary outputs at the columns of the crossbar to ensure system accuracy. To that effect, we propose an algorithm-hardware co-design approach, where DNNs are first trained with quantization-aware training. Subsequently, we introduce HCiM, an ADC-Less Hybrid Analog-Digital CiM accelerator. HCiM uses analog CiM crossbars for performing Matrix-Vector Multiplication operations coupled with a digital CiM array dedicated to processing scale factors. This digital CiM array can execute both addition and subtraction operations within the memory array, thus enhancing processing speed. Additionally, it exploits the inherent sparsity in ternary quantization to achieve further energy savings. Compared to an analog CiM baseline architecture using 7 and 4-bit ADC, HCiM achieves energy reductions up to 28% and 12%, respectively.

  12. Modak, Nirmoy, and Kaushik Roy. "Energy Efficiency Through In-Sensor Computing: ADC-less Real-Time Sensing for Image Edge Detection." ACM/IEEE International Symposium on Low Power Electronics and Design (ISLPED), 2024.

    Abstract: In-sensor computing has revolutionized modern vision-based applications, particularly in scenarios like autonomous vehicles and robotics where real-time or near-real-time processing is crucial. By enabling data processing at the sensor level, in-sensor computing eliminates the need to transmit data to cloud servers, significantly reducing latency and enhancing decision-making speed. Central to the in-sensor computing paradigm, CMOS image sensors (CISs) with edge computing, play a pivotal role in machine vision applications. The need for high resolution, low power, and real-time operation aligns seamlessly with the demands of modern vision-based applications. In this paper, we propose a novel approach for real-time image edge detection with an in-sensor, ADC-less sensing solution that achieves high energy efficiency and speed. The design utilizes the column-parallel architecture of existing CIS and the row-wise pixel readout scheme. Column voltages of three consecutive rows with a delay arrangement extract 4-bit edge pixels without deriving the actual digital image pixels. A time-to-digital conversion (TDC) technique using a 4-bit counter eliminates the requirement of power-hungry ADC. A 256(H) x 256(V) 2D CMOS pixel array with 10 ?m pixel pitch is simulated using Spectre in TSMC 65nm low-power technology. CMOS pixels with wide dynamic range (WDR) capture the light intensity variation up to 92dB [10]. Simulation results show energy consumption of 2pW per pixel per frame, operating at a frame rate of 3.9kfps, all well-contained within a modest 0.5 mW power budget. The resultant frame rate emerges as notably superior in terms of speed, accompanied by a more than tenfold reduction in power consumption per edge frame-pixel compared to the existing prior art.

  13. Ali, Mustafa, Indranil Chakraborty, Sayeed Choudhary, Muya Chang, Dong Eun Kim, Arijit Raychowdhury, and Kaushik Roy. "TOPS/W adaptive-SNR sparsity-aware CIM core with load balancing support for DL workloads." In 2023 IEEE Custom Integrated Circuits Conference (CICC) (pp. 1-2), IEEE.

    Abstract: The growing trends of developing domain-specific accelerators for Deep Learning (DL) applications has led to exploration of compute-in-memory (CIM) primitives based on SRAM [1] ? [5]. Multiple research chips have demonstrated macro and core-level designs supporting multi-bit Matrix-Vector Multiplication (MVM) and sparsity to increase energy-efficiency and performance. However, CIM designs suffer from the following challenges, as shown in Fig. 1: (1) Difficulty in leveraging both input and weight unstructured sparsity in existing DL accelerators. Note, unstructured sparsity is more amenable during DL model training than structured sparsity. Fig. 1 (top) shows input and weight bit-level sparsity of ResNet20 running a CIFAR10 task and mapped on a 64×64CIM macro. We observe that activations and weights of each layer experience different bit-level sparsity, also, sparsity levels vary significantly across layers.(2) Mixed-signal CIM macros suffer from noise and variation-based computation errors and signal-to-noise ratio (SNR) degradation. Moreover, the macro errors get accumulated in scaled-up CIM architectures leading to significant model accuracy drop. (3) Sparsity-aware CIM compute units encounter different sparsity; hence they might finish their corresponding MVMs at different times leading to load imbalance. To overcome the aforementioned challenges, this work proposes a sparsity-aware, adaptive-SNR CIM core based on sparsity aware CIM macros with load balancing support. The proposed core achieves 1.4-6.7 TOPS/W 8b energy efficiency and is fabricated in 65 nm technology. The core contributions are: 1) Input and weight unstructured bit-level sparsity exploitation by dynamically reconfiguring the CIM macros ADC precision. 2) Adaptive HW SNR using reconfigurable Word Line (WL) parallelism to adapt to workload SNR requirements and achieve optimal energy efficiency which provides 2x and 1.78x performance and energy benefits, respectively compared to iso-accuracy baseline where only 8 RWLs are enabled to maximize CIM SNR. 3) Flexible MVM kernel mapping and compiler-level load balancing and its corresponding HW support to balance weight sparsity among Sparse Compute Units (SCUs).

  14. Kim, Dong Eun, Aayush Ankit, Cheng Wang, and Kaushik Roy. "SAMBA: sparsity aware in-memory computing based machine learning accelerator." IEEE Transactions on Computers, 72(9), pp.2615-2627.

    Abstract: Machine Learning (ML) inference is typically dominated by highly data-intensive Matrix Vector Multiplication (MVM) computations that may be constrained by memory bottleneck due to massive data movement between processor and memory. Although analog in-memory computing (IMC) ML accelerators have been proposed to execute MVM with high efficiency, the latency and energy of such computing systems can be dominated by the large latency and energy costs from analog-to-digital converters (ADCs). Leveraging sparsity in ML workloads, reconfigurable ADCs can save MVM energy and latency by reducing the required ADC bit precision. However, such improvement in latency can be hindered by non-uniform sparsity of the weight matrices mapped into hardware. Moreover, data movement between MVM processing cores may become another factor that delays the overall system-level performance. To address these issues, we propose SAMBA, Sparsity Aware IMC Based Machine Learning Accelerator. First, we propose load balancing during mapping of weight matrices into physical crossbars to eliminate non-uniformity in the sparsity of mapped matrices. Second, we propose optimizations in arranging and scheduling the tiled MVM hardware to minimize the overhead of data movement across multiple processing cores. Our evaluations show that the proposed load balancing technique can achieve performance improvement. The proposed optimizations can further improve both performance and energy-efficiency regardless of sparsity condition. With the combination of load balancing and data movement optimization in conjunction with reconfigurable ADCs, our proposed approach provides up to 2.38x speed-up and 1.54x energy-efficiency over state-of-the-art analog IMC based ML accelerators for ImageNet datasets on Resnet-50 architecture.

  15. Ali, Mustafa, Indranil Chakraborty, Utkarsh Saxena, Amogh Agrawal, Aayush Ankit, and Kaushik Roy. "A 35.5-127.2 tops/w dynamic sparsity-aware reconfigurable-precision compute-in-memory SRAM macro for machine learning." IEEE Solid-State Circuits Letters, 4, pp.129-132.

    Abstract: This letter presents an energy-efficient sparsity-aware reconfigurable-precision compute-in-memory (CIM) 8T-SRAM macro for machine learning (ML) applications. The proposed macro dynamically leverages workload sparsity by reconfiguring the output precision in the peripheral circuitry without degrading application accuracy. Specifically, we propose a new energy-efficient reconfigurable-precision SAR ADC design with the ability to form ( n+m)-bit precision using n-bit and m-bit ADCs. Additionally, the transimpedance amplifier (TIA) ?required to convert the summed current into voltage before conversion? is reconfigured based on sparsity to improve sense margin at lower output precision. The proposed macro, fabricated in 65-nm technology, provides 35.5-127.2 TOPS/W as the ADC precision varies from 6 to 2 bit, respectively.

  16. Agrawal, Amogh, Mustafa Ali, Minsuk Koo, Nitin Rathi, Akhilesh Jaiswal, and Kaushik Roy. "IMPULSE: A 65-nm digital compute-in-memory macro with fused weights and membrane potential for spike-based sequential learning tasks." IEEE Solid-State Circuits Letters, 4, pp.137-140.

    Abstract: The inherent dynamics of the neuron membrane potential in spiking neural networks (SNNs) allows the processing of sequential learning tasks, avoiding the complexity of recurrent neural networks. The highly sparse spike-based computations in such spatiotemporal data can be leveraged for energy efficiency. However, the membrane potential incurs additional memory access bottlenecks in current SNN hardware. To that effect, we propose a 10T-SRAM compute-in-memory (CIM) macro, specifically designed for state-of-the-art SNN inference. It consists of a fused weight ( W MEM ) and membrane potential ( V MEM ) memory and inherently exploits sparsity in input spikes leading to ~97.4% reduction in energy-delay product (EDP) at 85% sparsity (typical of SNNs considered in this work) compared to the case of no sparsity. We propose staggered data mapping and reconfigurable peripherals for handling different bit precision requirements of W MEM and V MEM , while supporting multiple neuron functionalities. The proposed macro was fabricated in 65-nm CMOS technology, achieving energy efficiency of 0.99 TOPS/W at 0.85-V supply and 200-MHz frequency for signed 11-bit operations. We evaluate the SNN for sentiment classification from the IMDB dataset of movie reviews and achieve within ~1% accuracy difference and ~ 5× higher energy efficiency compared to a corresponding long short-term memory network.

Click on the expansion arrows in each section to read the publication abstracts.


<<< Return to Current Research Outline