Annual Workshop on Chips and AI

Tuesday, September 22, 2026
(Registration Required)
Purdue Material Sciences and Electrical Engineering Building
501 Northwestern Avenue
West Lafayette, IN 47907

Past Events

Workshop on Chips and AI, September 22, 2026

This event serves as the next annual workshop of the Institute of Chips and AI, building on the momentum established at the Institute?s kickoff workshop in November 2024 and the successful 2025 workshop. As a gathering point for returning and new participants, the workshop provides an opportunity to reconnect, share progress, strengthen collaborations, and identify new directions for research and innovation. This year's event will expand participation through broader registration opportunities and will introduce a hackathon designed to foster hands-on engagement, creativity, and interdisciplinary collaboration.

The Workshop on Chips and AI, sponsored by the U.S. National Science Foundation FuSe-TG: FAB: A Heterogeneous Ferroelectronics Platform for Accelerating Big Data Analytics and the Institute of Chips and AI, will bring together distinguished industry leaders, researchers, and Purdue University students to discuss the latest advances and challenges in topics including Hardware Technologies for Neural Computing, Systems and Architectures for Next Generation AI, and AI for Chip Design. Participants will have the opportunity to connect with emerging talent, explore cutting-edge research, and engage in meaningful conversations that foster collaboration, innovation, and the continued growth of the Institute's research community.

The workshop is organized by a dedicated committee of Purdue students who are actively involved in the Institute's efforts to advance AI research and innovation and are passionate about strengthening Purdue University's leadership in the rapidly evolving global AI and semiconductor landscape. Through this event, the committee aims to cultivate connections across academia, industry, and government while creating opportunities for students to contribute to the development of next-generation AI technologies, support the Institute's mission of moving AI forward, and expand their real-world impact.

Join the Hackathon!

NSF logo

Speaker Bios

Click this link to view all guest speaker bios.

Student Planning Committee


Agenda

Tuesday, September 22, 2026
Time Location Topic Speaker(s)/Participant(s)

8:30-9:00 a.m.

Rice Design Studio (MSEE 190)

Registration and Breakfast

All Participants

9:00-9:40 am

Rice Design Studio (MSEE 190)

Engineering Welcome

Nikhilesh Chawla
Associate Dean for Research and Ransburg Professor, Purdue Engineering

ECE Welcome

Milind Kulkarni
Birck Head and Professor, Purdue ECE

Institute Overview

Kaushik Roy
Director of Institute of Chips and AI and Tiedemann Jr. Distinguished Professor, Purdue ECE

9:40-10:05 a.m.

Rice Design Studio (MSEE 190)

Plenary Talk: Engineering AI Hardware at Scale: From Silicon to Data Centers

Sang Phill Park
Architect, Oracle Cloud Infrastructure (OCI)

AbstractScaling AI performance requires coordinated hardware design across silicon, packages, servers, racks, clusters, and data centers. Increasing compute capacity places demands on communication, power delivery, and heat removal, making integration across these scales central to the performance a system can sustain. Power availability and cooling increasingly set practical limits on system scale.

This talk examines how workload requirements interact with hardware architecture and physical constraints at each boundary. We discuss how integration and connectivity affect communication overhead, how rack and data-center power and cooling limits influence chip and server design, and how physical placement affects deployment and serviceability. Failures and repair time also influence available capacity and operating cost. These relationships lead to research questions about when tighter integration improves system performance, where communication limits scaling, and how to evaluate component advances under common workload and system constraints. The emphasis is on connecting innovations in devices, circuits, and architectures to achievable performance and energy efficiency across the complete AI infrastructure.

10:05-10:30 a.m.

Rice Design Studio (MSEE 190)

Plenary Talk: Enterprise AI Acceleration with IBM Spyre

Swagath Venkataramani
Principal Research Scientist, IBM T.J Watson Research Center

10:30-11:30 a.m.

Rice Design Studio (MSEE 190)

Panel - The Future of AI Computing: Workloads, Architectures, and Systems

Moderator: Timothy Rogers
Associate Professor, Purdue ECE
Chankyu Lee
AI Research Scientist, NVIDIA
Siri Narla
Embedded NVM Technology Architect, Global Foundries
Sang Phill Park
Architect, Oracle Cloud Infrastructure (OCI)
Kaushik Roy
Director of Institute of Chips and AI and Tiedemann Jr. Distinguished Professor, Purdue ECE
Swagath Venkataramani
Principal Research Scientist, IBM T.J Watson Research Center

Break: 11:30 a.m. - 11:45 a.m.

11:45 a.m.-12:50 p.m.

Rice Design Studio (MSEE 190)

Student Lightning Talks

Lunch: 12:50 - 1:15 p.m.

1:15-2:15 p.m.

Rice Design Studio (MSEE 190)

Poster Session & Hackathon Demos

Student Poster List
Student Poster Presenters
Hackathon Finalist List
Hackathon Finalists


2:15-2:35 p.m.

Rice Design Studio (MSEE 190)

Plenary Talk: Memory-Centric AI at the Edge

Siri Narla
Embedded NVM Technology Architect, Global Foundries

AbstractArtificial intelligence is rapidly expanding beyond the cloud into sensors, edge devices, industrial systems, autonomous platforms, and personal electronics. As AI models and sensor data volumes continue to grow, data movement is becoming a fundamental limitation, bringing the von Neumann bottleneck to the edge where power, latency, cost, bandwidth, and reliability constraints are most severe. This talk examines how overcoming these challenges may require rethinking the relationship between memory, compute, and algorithms. We explore emerging directions ranging from in-memory computing to next-generation embedded memories optimized for AI systems and hierarchical AI systems that distribute intelligence across the edge ecosystem. The goal of the talk is to highlight a rich set of research opportunities spanning across the stack to enable the next-generation of physical AI.

2:35 - 2:55 p.m.

Rice Design Studio (MSEE 190)

Plenary Talk: Powering AI is the Quiet Scaling Limit: Reclaiming the Voltage We Ship but Never Use

Charles Augustine
Circuit Design Engineer, NVIDIA Inc.

AbstractEvery chip ships with more voltage than its logic needs. Between IR drop, transient di/dt droop, aging, process spread, and our own measurement uncertainty, 10-20% of the supply is guardband that the logic never uses and the whole system pays for. Because dynamic power scales with the square of voltage, this is expensive: 15% less voltage is roughly 28% less dynamic power — with no new process node, no new architecture, and no software rewrite. At data-center scale the arithmetic is stark. U.S. data centers used 176 TWh in 2023, so a single percent reclaimed is roughly 170,000 homes powered for a year.

AI workloads make this harder rather than easier. Thousands of MAC lanes switch coherently, and layer and all-reduce boundaries repeat on a fixed cadence, so the excitation lands in the mid-frequency gap where neither the voltage regulator nor on-die capacitance holds the rail. The resonance is set by the hardware; the excitation frequency is set by the compiler.

This talk argues that the remaining guardband is bought with knowledge, not silicon: put the instrument on the die, and turn a fixed margin into a control problem. It also maps where machine learning genuinely helps chip design, and where it cannot — because ground truth still requires silicon.

2:55 - 3:15 p.m.

Rice Design Studio (MSEE 190)

Plenary Talk: TBD

Sandeep Thirumala
Senior Engineer - Advanced DRAM Process Integration, Micron Technology

3:15-3:30 p.m.

Rice Design Studio (MSEE 190)

Results and Awards

3:30-4:30 p.m.

Rice Design Studio (MSEE 190)

Panel - AI-Native Chip Design: From Data-Driven EDA to Trustworthy Design Automation

Moderator: Anand Raghunathan
Co-Director of Institute of Chips and AI and Silicon Valley Professor, Purdue ECE
Charles Augustine
Circuit Design Engineer, NVIDIA Inc.
Sumeet Gupta
Professor, Purdue ECE
Vijay Raghunathan
VP and University Ambassador to India and Professor , Purdue ECE
Sandeep Thirumala
Senior Engineer - Advanced DRAM Process Integration, Micron Technology
Srividhya Venkataraman
Fellow Silicon Design Engineer, AMD

4:30-5:00 p.m.

Rice Design Studio (MSEE 190)

Industry Panel - Career Advice

5:00 p.m.

Rice Design Studio (MSEE 190)

Closing + Networking Session

Student Posters

AI for CHIPS
Number Authors Title

1

M. Khan, S. Gupta

Repurposing Batch Normalization for Autonomous and In-Situ Detection and Mitigation of Weight Corruptions in CNN Accelerators

AbstractIn safety-critical deployments, AI hardware must remain reliable against a broad spectrum of threats such as radiation-induced upsets, aging, hard faults, and adversarial attacks (e.g., the progressive bit-flip attack (PBFA)). All of these corrupt stored weights while the chip keeps producing confident (but inaccurate) predictions. Detecting and mitigating such weight perturbations is therefore crucial for safety-critical platforms. To that end, we propose an on-chip, batch normalization (BN)-based technique that continually senses shifts in activation statistics to detect and mitigate a wide variety of weight corruptions on the fly, including random and localized faults as well as adversarial bit flips. The technique operates in two phases: (1) offline pre-characterization of the relationship between activation shifts and inference accuracy drop, and (2) on-chip runtime detection and mitigation of weight corruptions. It is designed so that benign perturbations stay silent while harmful faults are flagged. Upon detection, the flagged layer is re-centered toward its stored clean reference within the same forward pass. The technique is fully autonomous, eliminating the need for host communication, operation halts, or access to fine-tuning data. If the residual shift after mitigation indicates that accuracy has fallen below a user-set floor, a held-out watcher aborts the inference. Evaluated on ResNet-20/50 and MobileNetV2 for CIFAR-10/100, our approach detects harmful corruptions with >99% precision across all fault types. Furthermore, it recovers accuracy from 10% to 85.88% under 0.5% random bit flips (ResNet-50/CIFAR-10), up to 84% for localized faults (MobileNetV2/CIFAR-10), and from random-guess accuracy to 80-83% under PBFA (ResNet-20/CIFAR-10), entirely on chip and within the same inference pass. This comes with <2% latency cost and 0.53% computation overhead.
BioMarzia Khan is a third-year Ph.D. student in the Elmore Family School of Electrical and Computer Engineering at Purdue University, advised by Prof. Sumeet Kumar Gupta. She received her B.Sc. in Electrical Engineering from Bangladesh University of Engineering and Technology (BUET) in 2022. Her research focuses on runtime hardware fault detection and mitigation for quantized DNN accelerators.

2

J. Balasubramanian, A. Raghunathan

PowerScope: ML-based Intra-Cycle Power Estimation

AbstractPower estimation at sub-clock-cycle temporal resolutions is critical for tasks such as power delivery network (PDN) design, dynamic voltage droop analysis, and pre-silicon power side-channel security evaluation. Designers commonly rely on commercial post-layout gate-level power analysis tools for these tasks, but these flows are computationally expensive and scale poorly with design size and workload length. Machine learning (ML)-based power estimation frameworks have shown promise in accelerating power estimation, but prior efforts only address average power or per-cycle power estimation. We propose PowerScope, the first ML-based intra-cycle power estimation framework. PowerScope operates purely on RTL simulation traces at inference time, eliminating the need for post-layout gate-level simulation and power analysis per workload. Across a diverse benchmark suite, PowerScope achieves 3.3% median and 6.7% mean absolute percentage error compared to commercial post-layout gate-level power estimates while running ∼571× faster. We further demonstrate that PowerScope's predictions can be reliably used for the downstream task of pre-silicon power side-channel leakage assessment.
BioJayanth Balasubramanian received his B.Tech and M.Tech in Electrical Engineering from Indian Institute of Technology Madras in 2025. He is currently a PhD student at Purdue University advised by Prof. Anand Raghunathan. His research focuses on the intersection of ML and EDA, with a focus on power estimation for SoCs.

3

G. Yarramneni, A. Raghunathan

Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs

AbstractModern edge SoCs integrate heterogeneous processing units?including CPUs, GPUs, and NPUs?with different performance and energy characteristics. Deploying AI inference workloads on these platforms requires jointly selecting workload mappings and processing-unit configurations, creating a combinatorial design space that makes exhaustive search impractical. Existing black-box design-space exploration methods typically guide the search using only sparse objective values, providing little insight into why a candidate performs as it does. We present TraceDSE, an agentic design-space exploration flow that uses system execution traces as rich feedback. An LLM proposer generates candidate mappings and hardware configurations for evaluation, while an LLM critic analyzes the resulting traces using programmatic tools to identify bottlenecks and recommend targeted refinements. Across four AI inference workloads deployed on an Intel Meteor Lake SoC, TraceDSE improves Pareto-front hypervolume by up to 35% over NSGA-II and 68% over Bayesian optimization while requiring approximately 6-9× fewer hardware evaluations.
BioGeetha Prasuna Yarramneni is a Ph.D. student in the Elmore Family School of Electrical and Computer Engineering at Purdue University, advised by Prof. Anand Raghunathan. Her research focuses on AI for electronic design automation, particularly agentic co-design of hardware, software, and systems for efficient computing. She received her B.Tech. in Electrical Engineering from IIT Madras and previously spent two years at NVIDIA designing and verifying low-power features for GPU SoCs.

4

S. Chandra, K. Roy

2D-ThermAl: TCAD Model Informed Thermal Analysis of Circuits using GenAI

AbstractThermal analysis is increasingly critical in modern integrated circuits, where non-uniform power dissipation and high transistor densities can cause rapid temperature spikes and reliability concerns. Traditional methods such as FEM-based simulations offer high accuracy but computationally prohibitive for early-stage design, often requiring multiple iterative redesign cycles to resolve late-stage thermal failures. To address these challenges, we propose '2D-ThermAl', a physics-informed generative AI framework which effectively identifies heat sources and estimates full-chip transient and steady-state thermal distributions directly from input activity profiles. ThermAl employs a hybrid U-Net architecture enhanced with positional encoding and a Boltzmann regularizer to maintain physical fidelity. Our model is trained on an extensive dataset of heat dissipation maps for more than 200 circuit configurations, ranging from simple logic gates (e.g., inverters, NAND, XOR) to complex designs, generated via COMSOL and Cadence EDA flows. The dataset captures diverse activity patterns, and we note that material-dependent thermal properties may require targeted fine-tuning to ensure accuracy across different fabrication contexts. Experimental results demonstrate that ThermAl delivers precise temperature mappings for large circuits, with a root mean squared error (RMSE) of only 0.71°C and outperforms conventional FEM tools by running up to ~200x faster. We analyze performance across diverse layouts and workloads and discuss its applicability to large-scale EDA workflows. While thermal reliability assessments often extend beyond 85°C for post-layout signoff, our focus here is on early-stage hotspot detection and thermal pattern learning. To ensure generalization beyond the nominal operating range (25-55°C), we additionally performed cross-validation on an extended dataset spanning 25-95°C maintaining a high accuracy (<2.2% full-scale RMSE) even under elevated temperature conditions representative of peak power and stress scenarios.} Limitations such as 2D-only modeling and real-world validation are addressed with concrete future directions, including 3D extension, generalization across technology nodes, and transfer learning strategies.
BioSoumyadeep Chandra received his B.Tech degree in Electrical and Telecommunication Engineering from Jadavpur University, Kolkata, India in 2020. Currently, he is a 6th year Ph.D. student in the School of Electrical and Computer Engineering at Purdue University, West Lafayette, IN, under the mentorship of Prof. Kaushik Roy. His ongoing research focuses on areas such as procedural learning, scene understanding, surgical workflow analysis, and generative AI for hardware.

5

A. Joshi, K. Roy

SHIRE: Enhancing Sample Efficiency using Human Intuition in REinforcement Learning

AbstractThe ability of neural networks to perform robotic perception and control tasks such as depth and optical flow estimation, simultaneous localization and mapping (SLAM), and automatic control has led to their widespread adoption in recent years. Deep Reinforcement Learning (DeepRL) has been used extensively in these settings, as it does not have the unsustainable training costs associated with supervised learning. However, DeepRL suffers from poor sample efficiency, i.e., it requires a large number of environmental interactions to converge to an acceptable solution. Modern RL algorithms such as Deep Q Learning and Soft Actor-Critic attempt to remedy this shortcoming but can not provide the explainability required in applications such as autonomous robotics. Humans intuitively understand the long-time-horizon sequential tasks common in robotics. Properly using such intuition can make RL policies more explainable while enhancing their sample efficiency. In this work, we propose SHIRE, a novel framework for encoding human intuition using Probabilistic Graphical Models (PGMs) and using it in the Deep RL training pipeline to enhance sample efficiency. Our framework achieves 25-78% sample efficiency gains across the environments we evaluate at negligible overhead cost. Additionally, by teaching RL agents the encoded elementary behavior, SHIRE enhances policy explainability. A real-world demonstration further highlights the efficacy of policies trained using our framework.
BioAmogh Joshi received his B.Tech. in Electronics Engineering from the University of Mumbai in 2020. From 2020 to 2021, he was a Project Research Assistant at the Indian Institute of Technology, Bombay where he worked on safety devices for Lighter-than-Air systems, and an avionics and control suite for a medium-weight class UAV. Amogh joined the Nanoelectronics Research Lab in the Spring of 2022, where he is currently pursuing his PhD under the supervision of Prof. Kaushik Roy. His research interests include autonomous and resilient robot learning, reinforcement learning, and neuromorphic computing for autonomous drone applications. He is currently exploring robot synthesis techniques using constrained co-learning approaches.

6

D. Lee, K. Roy

ViP-VLA: Vision-Internal Token Pruning for Efficient Vision-Language-Action Models

AbstractVision-Language-Action (VLA) models enable general robotic control by leveraging pretrained vision-language representations. However, their high inference cost makes closed-loop deployment expensive. At every decision step, hundreds of visual tokens are repeatedly processed through the vision encoder and the downstream model. Reducing this computation requires identifying which visual tokens are important for action generation and how early the less important tokens can be removed. We introduce ViP-VLA, an action-aware visual-token pruning framework that identifies visual tokens important for action generation and removes unimportant tokens inside the vision encoder. ViP-VLA learns token importance at an intermediate vision layer using the pretrained VLA's own action objective. During inference, the selector retains only the highest-scoring visual tokens after the intermediate vision layer. These tokens continue through the remaining vision-encoder and cross-modal layers, while discarded tokens receive no further processing, reducing computation without modifying the pretrained VLA. We evaluate ViP-VLA across multiple VLA architectures and robot manipulation benchmarks. On LIBERO with OpenVLA-OFT, ViP-VLA uses only 29.3% of the total FLOPs required by the original policy without token pruning and achieves a measured 1.85x inference speedup, while its average success rate is 96.8%, compared with 96.2% for the original policy.
BioDonghun Lee (James) received his B.S. in Electrical and Electronic Engineering from Yonsei University, Korea, in 2022. During his undergraduate years, he worked on Computer Vision. Currently, he is pursuing an M.S. degree in Electrical and Computer Engineering at Purdue University under the supervision of Prof. Kaushik Roy. His primary research interests lie in Neuro-Inspired Algorithms and Spiking Neural Networks. In his free time, he enjoys swimming as a hobby.

7

A. Mukherjee, K. Roy

DRC-Aid: Design-Rule Violation Correction via Agentic Framework utilizing Inference-Time Large Language Models

AbstractResolving Design Rule Violations (DRVs) in layouts entails an iterative loop of geometric edits and verification. We present DRC-Aid, a closed-loop agentic framework that automates local DRC repair by formulating it as verification-in-the-loop search. To constrain the combinatorial geometric repair space, a deterministic Rule Engine converts physical verification tool-reported violations into a bounded menu of geometric edits. An off-the-shelf Large Language Model (LLM) evaluates local geometric context to select edits from this menu, with budgeted depth-first search and backtracking. Immediate feedback from verification tools enforces geometric compliance and guards against electrical-topology degradation, while a global Memory Bank prevents cyclic re-exploration. Evaluated on FreePDK45 layouts containing DRVs, DRC-Aid achieves DRC-clean, electrically-equivalent repairs in ∼92.5% of cases, while residual cases yield partially repaired candidates. Under an identical search and verification infrastructure, LLM- based selection outperforms random (54.4%) and deterministic-heuristic (83.3%) policies, with the gap widening on cases more than six violations.
BioAnushka completed her bachelor's degree from TISL, Kolkata in Electronics and Communication engineering in 2021. She worked at Cognizant Technology Solutions from 2021 to 2022. She completed her masters degree from Indian Institute of Engineering Science and Technology, Shibpur in VLSI Design in 2024, where she primarily worked on hardware security and was awarded the silver medal for securing the second position in the M.Tech program. Anushka is currently enrolled in Purdue's PhD program under the direction of Prof Kaushik Roy. She currently works on hardware-software co-design and automation in EDA workflows.

8

T. Chen, Y. Kim

Kernel-Informed Predictive VRM Control for Efficient and Robust GPU Power Delivery

AbstractLarge language model inference on GPUs produces repeated load-current steps at kernel launches. On sub-1 V rails at kilowatt-class power, these steps reach hundreds of amps, and the static voltage guardband that absorbs them is increasingly expensive. On-die responses such as adaptive clocking protect against a droop by stretching the clock, so every event they catch costs performance. Moving the regulator setpoint ahead of the step costs voltage instead and is the response that can shrink the guardband. But given the regulator's limited slew rate, pre-positioning the setpoint takes microseconds of lead, far more than any on-die signal can provide. Kernel launches offer that lead: decode is structurally repetitive, so the step each launch will cause is predictable from the identity of the kernel about to run. This work explores a predictor that exploits this recurrence to pre- position the setpoint before the step arrives. We investigate what information in the launch stream identifies a kernel's load step, how to learn step magnitudes online and decide when a prediction is trustworthy enough to act on. We evaluate the design in AccelWattch coupled to a PDN and regulator model, separating the ideal mechanism bound from interface and learning cost.

9

S. Lee, Y. Kim

Lightweight Early Workload Prediction for Energy-Efficient Audio Event Classification

AbstractAudio event classification (AEC) enables environmental awareness in various real-world applications, but the deployment of always-on AEC on battery-powered low-power microcontrollers (MCUs) remains challenging due to the tight energy budget and strict real-time constraints. Although acoustic machine learning (ML) models are relatively compact, continuous feature extraction and inference consume substantial energy in streaming workloads. Early-exit inference combined with dynamic voltage and frequency scaling (DVFS) has proven effective in energy-constrained real-time systems, yet significant energy-saving opportunities remain under-explored for streaming workloads. In this paper, we exploit the observation that adjacent audio segments from the same event often produce similar feature tensors, while each event class exhibits a distinct and relatively consistent exit-depth requirement. Together, these two signals enable early workload prediction and adaptive inference. Based on these, we develop an energy-efficient inference framework for AEC on MCUs. Through early, lightweight workload prediction, the proposed framework enables proactive DVFS under real-time constraints. Experimental results show that the proposed framework reduces average inference energy by 16.17%-17.55% while preserving accuracy, enabling more practical always-on AEC on commodity ultra-low-power platforms.

10

H. Vu, Y. Kim

DynaBreath: Motion-Robust mmWave Respiration Monitoring

AbstractMillimeter-wave (mmWave) radar enables contactless, privacy-preserving respiration monitoring using commodity chipsets, making it attractive for clinical and home-care settings. However, current systems lose accuracy when the subject moves. Existing approaches reduce radar measurements to a one-dimensional phase trace, which collapses the spatial information needed to tell breathing apart from whole-body movement.

We propose DynaBreath that addresses this challenge by processing radar signals in three dimensions. The key insight is that breathing causes localized chest-wall deformation, while body sway translates the torso approximately uniformly. These two types of motion are geometrically distinguishable in a 3D point cloud. The system tracks the subject's torso pose to establish a body-centric reference frame, generates a spatially dense point cloud from the tracked torso region using interferometric elevation estimation and Doppler compensation, and estimates the respiration waveform through anchor-based spatial encoding with temporal modeling.

We evaluate DynaBreath under continuous periodic oscillation and free random motion. The system achieves mean absolute errors of 1.89 and 3.01 breaths per minute with waveform availabilities of 88.6% and 76.8%, representing a 3-4× error reduction and up to 6.4× availability improvement over four reimplemented baselines. Overall, the results demonstrate that 3D spatial processing on commodity radar hardware enables robust respiration monitoring under the tested continuous-motion conditions.

11

A. Mukherjee, K. Roy

FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding

AbstractFine-grained visual understanding depends on local detail, yet visual encoders face a trade-off between costly full-image high-resolution processing and compact global encoding that can weaken such evidence. Inspired by human active vision, we separate where to look from what to encode. We focus on the latter and introduce FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry. We first isolate this encoding problem using oracle ground-truth crops in a controlled small-object regime. On ImageNet objects with a native maximum side of 96 pixels, FAVE improves Top-1 by 9.4 points over a fixed-resolution ViT on the same crop window with 12.7 times lower FLOPs. Increasing global resolution or backbone capacity does not recover the same operating point. We then integrate FAVE as a complementary local branch in FastVLM. Its local tokens are combined with FastVLM's global visual tokens, while the original global pathway and language model remain frozen. With at most 16 additional local tokens, FAVE improves TextVQA by 1.60 points and achieves a 3.3 times controlled TTFT speedup over SmolVLM2-2.2B. On GQA attribute questions, it improves FastVLM-1.5B by 1.31 points, extending the benefit beyond text while narrowing the gap to FastVLM-7B. Together, these results show that selectively allocating high-acuity local capacity provides an efficient complement to broader global representations and model scaling for fine-grained understanding of small objects, text, and attributes.
BioAmitangshu received his B. Tech in Applied Electronics and Instrumentation Engineering from West Bengal University of Technology in 2016. He graduated from Iowa State University, USA with a Master's in Computer Engineering in Fall 2019 where he worked as a Graduate Research Assistant at the Self-Aware Complex Systems Lab and DICE Lab (now moved to New York University) from Fall 2018 to Spring 2020. His primary research work revolved around topics in Deep Learning such as Domain Adaptation and Adversarial Machine Learning and he published his research works in conferences such as ICCV 2019 (main conference), CVPR and Neurips (Workshops). His M.S thesis project focused on GAN based Data Augmentation techniques for improving performances of Deep Networks in Perception based tasks for Autonomous Driving. In Summer 2020, he joined Purdue University's ECE department to pursue his PhD degree under the guidance of Prof. Kaushik Roy at the Center for Brain-Inspired Computing (C-BRIC). His primary research interests will focus on designing Neuro-inspired Algorithms for Explainable and Robust Learning in Autonomous Systems.

12

M. Chauhan, A. Bera

HEROIC: Heterogeneous Evidential Reasoning for Open-Vocabulary Identification and Cross-Robot Collaboration

AbstractMulti-agent heterogeneous air-ground robot teams are attractive for open world search, with applications for reconnaissance, urban search and rescue missions (USAR), disaster response and recovery, and hazardous environments. These two platforms have different failure modes: aerial robots cover ground quickly but cannot resolve small or occluded targets from altitude, while ground robots can identify objects-of-interest, such as people or hazardous objects, at close range but cover less area. Existing language-tasked teams either have roles fixed prior, or have a language model assign them from hand-written capability tags, so the team is unable to know when within a mission an asset is no longer useful. We present HEROIC, a decentralized heterogeneous multi-agent open-vocabulary search coordination framework that requires agents to communicate in natural language only. HEROIC's initial agent role assignment is derived from sensor properties and a scale law to determine whether targets can be detected with a high confidence. From the mission's natural language prompt alone, this law assigns aerial flight altitudes and sweep spacing. When this calculated height falls below the altitude for safe flight, aerial agents re-task themselves from searcher to aerial triage, escort, and route guide for ground agents. Both robots maintain an evidential belief over the search area (bearing rays for positive evidence, a log-odds posterior for negative evidence) and gate any arrival on close-range verification. In full-stack experiments, HEROIC reaches the target 84\% of the time across all 6 scenes, compares to 35-53\% for vision-language frontier baselines, frontier-based search, lawnmower, and random-walk running the same perception, all while being 2-4x sooner to arrive at the target.
BioMihir Chauhan is pursuing B.S. degrees in computer science and mathematics at Purdue University, West Lafayette, IN, USA. He is an undergraduate researcher in the IDEAS Lab, advised by Professor Aniket Bera, where he works on reinforcement learning for autonomous and collaborative robotics, including adversary-aware navigation, heterogeneous air-ground robot teams, and structured value learning for manipulation. His work has been published at IEEE ICRA and in the IEEE Sensors Journal, where his paper has been nominated for the Best Paper Award. He was previously a Research Intern with the Department of Electrical Engineering and Computer Sciences, University of California at Berkeley. Mihir has twice interned as a Member of Technical Staff at Nutanix, building ML-driven tools for hardware performance analysis and LLM inference on CPUs. In 2020, he co-founded the Virtual Robot Simulator, a robotics learning platform used by more than 60,000 students in 35 countries. His work on a portable water harvesting machine was showcased in the Smithsonian's FUTURES exhibition in Washington D.C. Mihir?s recognitions include FIRST Dean's List Finalist and FIRST Tech Challenge World Championship Finalist.

13

K. Cheng, A. Bera

RTG: Reverse Trajectory Generation for Rigid-Body Manipulation Under Sparse Rewards

AbstractSparse-reward reinforcement learning remains a long-standing challenge for robotics, where extensive exploration is often required before meaningful reward signals can guide the propagation of state-value functions. Prior approaches typically rely on expert offline demonstrations or carefully designed curriculum learning strategies to improve exploration efficiency. Leveraging recent advances in differentiable rigid-body dynamics and trajectory optimization, we compute physically plausible trajectories that end at a goal configuration to provide dense learning signals. We propose Reverse Trajectory Generation (RTG), a replay-generation method that starts from sampled successful configurations and constructs predecessor transitions with a trajectory-optimization-based Reverse Rigid-Body Simulator (RRBS). RTG forward-replays generated transitions to reduce simulator inconsistency, uses beam search to obtain diverse reverse trajectories, and mixes the resulting replay with online data in off-policy RL. Across gathering, sorting, and articulated object manipulation tasks, RTG consistently improves learning efficiency and task performance over off-policy baselines and improved sampling strategies like reverse-curriculum and hindsight-relabeling baselines. We further quantify reverse-generation cost and demonstrate that the generated replay remains useful after forward replay in a 3D PyBullet environment. Results show that RTG is a practical way to exploit rigid-body structure for sparse-reward robot learning.

14

A. Kumar, A. Bera

Fitting Through: Perceptive Posture Reconfiguration for Humanoid Locomotion in Height-Constrained Environments

AbstractHumanoid robots are mechanically capable of entering spaces that are inaccessible to wheeled and quadrupedal platforms, but doing so requires abandoning upright bipedal locomotion in favor of contact modes that recruit the knees and hands. We present a single perceptive policy that enables a Unitree G1 humanoid to traverse vertically constrained environments by continuously reconfiguring between crouched bipedal walking, knee walking, and quadrupedal crawling, spanning base heights from 0.78 m when standing down to 0.3 m when crawling, with smooth transitions between contact modes rather than discrete skill switching. Perceiving overhead geometry across these postures is difficult because the sensor pose relative to the ground changes drastically with contact mode, and because the underside of an overhang is only observable from below it. We therefore introduce a configuration-invariant clearance map, which the policy observes directly: depth returns from multiple body-mounted cameras are projected through the kinematic chain into a gravity-aligned, terrain-anchored grid in which each cell stores the free vertical space between the floor and the ceiling above it. Because clearance is a difference of heights in a common frame, the representation is invariant to the robot's own posture and sensor placement, so a single visual observation is shared across all contact modes. The map carries an explicit observability mask, distinguishing unobserved geometry from free space, so that the policy commits to a posture from the geometry it has resolved rather than treating occluded regions as passable. Conditioned on this map, the policy selects its own posture, transition timing, and forward speed from the geometry ahead, with no external mode arbitration. We train one teacher per mode and per transition using privileged information, distill them into a unified student, and fine-tune with multi-critic PPO, which decouples the conflicting reward structures of the different contact modes.
BioAyush is a Master's student in Robotics at Purdue University, where he has been a member of the IDEAS Lab since 2025. His research focuses on humanoid locomotion control and multi-robot collaboration. Prior to his Master's, he spent 2.5 years in industry working on controls and autonomy for legged robots.

15

Z. Li, A. Bera

PoseShield: Neural Collision Fields for Human Self-Collision Resolution

AbstractSelf-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extreme articulations or stochastic motion synthesis, generated meshes frequently exhibit self-penetrations, leading to physically implausible results. We propose PoseShield, a neural collision constraint defined directly in SMPL pose space. We formulate collision correction as a constrained optimization problem and connect the learned constraint with the Eikonal equation. Enforcing Eikonal regularization ensures non-vanishing gradients near the collision boundary, improving numerical stability and robustness of the optimization process. Unlike prior methods that operate in the mesh space or rely on heuristic penalties, our approach operates directly in the low-dimensional space of human poses and is theoretically grounded. The same learned constraint extends to human motion sequences, providing a generator-agnostic post-hoc collision corrector without retraining the underlying motion model. Experiments on a newly constructed SMPL pose benchmark show that our method achieves a 95.8% success rate and outperforms state-of-the-art baselines.

16

X. Sun, A. Bera

M3P-R1: Reinforcement Learning for Large Language Model Guided Multi-Modal Motion Planning via MIP Code Generation

AbstractMulti-Modal Motion Planning (M3P) requires joint reasoning over continuous motions and discrete mode transitions, making it difficult to solve efficiently. For instance, a bipedal robot may walk to a target location and then use its arms to grasp an object. This scenario captures both mode transitions and continuous dynamics, yielding feasible paths that neither purely discrete nor continuous planners can handle. While Mixed-Integer Programming (MIP) offers a principled framework, constructing tractable formulations for non-convex problems is typically manual and domain-specific, especially in the approximate, discretization-based MIP regime needed for non-convex robotic tasks. We propose MP-R1, a reinforcement learning method that fine-tunes large language models (LLMs) to decompose MP tasks into MIP variables, constraints, and objectives. Instead of directly outputting answers—often prone to hallucination—the model generates executable Python code using MIP optimization libraries and constraint interfaces. This enables solver-backed execution for robust and verifiable solutions.

17

A. Vashisth, A. Bera

CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation

AbstractRobot navigation typically assumes an obstacle-free path exists between start and goal. In real environments, however, clutter may block all routes. We introduce Lifelong Interactive Navigation, where a mobile robot with manipulation capabilities must move objects to forge paths and complete sequential object-placement tasks. Because environment modifications persist, decisions impact future navigability and task difficulty. We propose CoReLIN, an LLM-driven constraint-based reasoning framework with active perception. CoReLIN reasons over a structured scene graph to decide which objects to relocate, where to place them, and where to explore next. A standard motion planner executes reliable navigation and manipulation primitives. To evaluate long-horizon behavior, we introduce 2 new metrics - Long-term Efficiency Score (LES), a unified metric capturing success, execution efficiency, environment optimality, captured by Price of Clutter. In ProcTHOR-10k, CoReLIN outperforms best baseline by 16% under standard metrics and LES, and transfers to real-world hardware.

18

D. Sharma, K. Roy

Input Loss Curvature as a Predictor of Sample Vulnerability to Hardware Noise

AbstractAnalog in-memory computing (AIMC) accelerators can deliver significant energy efficiency over conventional architectures, but their accuracy is limited by device- and circuit-level non-idealities. While prior work has characterized the effects of these non-idealities at model or layer granularity, their impact on individual samples remains largely unexplored. In this work, we show that an input?s vulnerability to such non-idealities can be strongly predicted by its loss curvature, a metric capturing how sharply the loss changes under small input perturbations. Across multiple models, datasets, and non-idealities, our experiments reveal a strong positive correlation between input loss curvature and analog non-ideality-induced failures, with failure rates increasing substantially for high-curvature samples. These findings uncover a previously overlooked, sample-level dimension of hardware robustness and suggest new opportunities for input-aware strategies for addressing non-idealities.
BioDeepika received her B. Tech in Electrical Engineering from Indian Institute of Technology Roorkee, India. During her undergraduate studies, she worked on digital circuit designing for different applications including wireless communication. She joined Purdue in Fall 2019 and currently she is pursuing her PhD under Prof. Kaushik Roy. Her research interests include developing hardware architectures for neuromorphic computing. She loves to read and travel.

CHIPS for AI

19

D. Sharma, K. Roy

SpiDR: A 65nm 5 TOPS/W Digital CIM Accelerator with Reconfigurable Precision and Temporal Pipelining for Spiking Neural Networks

AbstractSpiking Neural Networks (SNNs) have the inherent ability to process highly sparse event data from dynamic vision sensors (DVS). However, many existing SNN accelerators support only fixed network architectures, limited data precision, and fail to efficiently handle membrane potential dynamics and varying input sparsity. This work presents SpiDR, a reconfigurable digital compute-in-memory (CIM) SNN accelerator that addresses these limitations through: (1) fused weight-Vmem macros for reduced data movement, (2) staggered layout and reconfigurable peripherals for variable-precision support, (3) zero-skipping to exploit unstructured input sparsity, and (4) asynchronous pipelining to maintain throughput despite variable computation times. Fabricated in TSMC's 65 nm CMOS technology, SpiDR achieves up to 5 TOPS/W at 95% input sparsity and supports diverse event-based workloads, including gesture recognition and optical flow estimation. SpiDR achieves 2× to 1000× higher area efficiency and up to 7.6× higher energy efficiency compared to prior designs.
BioDeepika received her B. Tech in Electrical Engineering from Indian Institute of Technology Roorkee, India. During her undergraduate studies, she worked on digital circuit designing for different applications including wireless communication. She joined Purdue in Fall 2019 and currently she is pursuing her PhD under Prof. Kaushik Roy. Her research interests include developing hardware architectures for neuromorphic computing. She loves to read and travel.

20

S. Bose, K. Roy

HyFPCiM: A 65-nm 417-µW Error-Sensitivity-Aware FP8 Compute-in-Memory Macro

AbstractEdge AI applications such as wearables, health monitors, and IoT sensors require sub-mW inference, but existing floating-point computing-in-memory (FP-CiM) solutions consume 10-460 mW, far exceeding device budgets. Integer quantization (INT8/INT4) reduces bit-width but often demands high-precision scale-factor multiplications and retraining, which erode efficiency. This work presents HyFPCiM, a hybrid FP8 (1S-4E-3M) CiM architecture that enables sub-mW floating-point inference through error-sensitivity-aware FP partitioning (EAP). Exploiting the 8-15x higher error sensitivity of exponents relative to mantissas, EAP maps exponent processing to a digital CiM (DCiM) macro for exact, low-cost computation, while mantissa accumulation is performed using an analog CiM (ACiM) macro, where moderate circuit non-idealities are tolerable and massive parallelism can be exploited. The proposed DCiM employs dual-rail inverter-based sensing for exponent processing, while the ACiM performs pre-analog-to-digital conversion (ADC) charge-domain accumulation using switched-capacitor (SC) MAC circuits, avoiding the power- and area-intensive adder-tree-based accumulation employed in prior FP-CiM designs. By amortizing a single ADC conversion across eight MAC units and employing a duty-cycled, power-gated transimpedance amplifier (TIA)/ADC, HyFPCiM reduces ADC energy by 8x and ADC+TIA power by 75%, and total macro power by 31% compared to prior FP-CiM designs. Fabricated in 65 nm TSMC LP, HyFPCiM achieves 417 µW at 50 MHz (1 V), achieving 40 GFLOPS/W system efficiency and 18.2 TFLOPS/W/mm² area efficiency. After iso-28 nm normalization, HyFPCiM outperforms the best prior work by >2.75x. FP8 inference on ResNet-18 with CIFAR-10/100 and Tiny-ImageNet shows <0.55% accuracy loss versus software FP8. To the best of our knowledge, HyFPCiM achieves the lowest reported power among silicon-measured FP-CiM macros, demonstrating its suitability for energy-constrained edge AI inference.
BioSouhardya received his Bachelor of Technology (Honours) degree from the Indian Institute of Technology Kharagpur, India, in 2025 with a major in Electrical Engineering and a minor in Electronics and Electrical Communication Engineering. He was awarded the Prof. K Venkataratnam Memorial Award for securing the highest GPA in core electrical courses and their associated laboratories. During his undergraduate years, he has worked on mixed-signal circuits, where he contributed to a tape-out of a RHBD SERDES IC, and compute-in-memory architectures. He is currently pursuing his Ph.D. Degree in Electrical and Computer Engineering at Purdue University under the supervision of Professor Kaushik Roy. His current research interests include hardware architecture design for in-memory computing, algorithm-hardware co-design, AI accelerators, and probabilistic computing. Besides research, he enjoys listening to folk music and Rabindra sangeet, gardening, and painting.

21

A. Holla, K. Roy

LIMO: Low-Power In-Memory-Annealer and Matrix-Multiplication Primitive for Edge Computing

AbstractCombinatorial optimization (CO), exemplified by the NP-hard Traveling Salesman Problem (TSP), is often intractable on von Neumann architectures due to the 'memory wall' and exponential search space growth with problem scale. While compute-in-memory (CiM) architectures using simulated annealing can mitigate these bottlenecks, solution quality degrades as instances scale. To address this, we present LIMO, a programmable mixed-signal macro implementing an in-memory annealing algorithm with reduced search-space complexity. LIMO utilizes stochastic switching of spin-transfer-torque magnetic-tunnel-junctions (STT-MTJs) to escape local minima. For large-scale problems, a refinement-based divide-and-conquer algorithm enables parallel optimization across a spatial architecture. Consequently, LIMO achieves superior quality and faster time-to-solution for instances up to 85,900 cities compared to prior annealers. Furthermore, with annealing functionality implemented as modular peripherals, the compute core can be repurposed for diverse workloads, such as neural network inference. We demonstrate image classification and face detection with software-equivalent accuracy, surpassing baselines in latency and energy efficiency.
BioAmod received his bachelor's degree in Electrical Engineering from the Indian Institute of Technology, New Delhi in 2024. During his undergraduate years, he worked on analog compute-in-memory using spin-based devices and spiking neural network (SNN) architectures for edge-AI applications. His work on crossbar-array-based SNNs for cognitive memory modeling earned him the award for Best Bachelor?s Thesis. His research interests include neuromorphic computing, hardware-algorithm co-design, and circuit design for supporting post-CMOS architectures. Apart from research, he enjoys nature, strength training, endurance running, traveling, gardening, and cooking.

22

R. Koduru, K. Roy

Integrating BEOL IGZO eDRAM as On-die Scratchpad for GPUs via System Technology Co-optimization

AbstractBack-end-of-line (BEOL) embedded DRAM (eDRAM) alleviates off-chip memory bottlenecks, limiting front-end-of-line (FEOL) overhead to peripherals. BEOL-compatible indium-gallium-zinc-oxide (IGZO) transistors offer ultra-low leakage, enabling capacitor-less high-retention gain-cells (4.3s at 85° C). We deploy these IGZO eDRAM arrays as extra on-die scratchpads for GPUs, bridging L2 cache and high-bandwidth memory (HBM). Placing them in intermediate BEOL layers (M6-M10 of in-house 7nm node), directly above their area-matched peripherals (in M1-M5), we preserve M11 and beyond for global routing. Under iso-GPU-die constraint, these peripherals displace baseline compute cores (SMs) and L2 cache, forcing a memory-compute trade-off. We develop a system-technology co-optimization (STCO) framework that co-designs this reallocation with realizable eDRAM capacity, bandwidth and evaluates the resulting scratchpad in an A100-like GPU. Utilizing 2T0C and 3T0C eDRAM macros, we achieve effective densities of 5.30 and 4.05 MB/mm2, yielding up to 800 and 600 MB scratchpads (15-20× baseline L2). Across LLM workloads, this integration yields up to 35% energy and 15% latency savings versus unmodified baseline, with eDRAM refresh energy consuming <0.001% of total system energy. We also identify optimal SM/L2 reductions and IGZO technology targets.
BioRevanth Koduru is a postdoctoral researcher holding a joint appointment between imec USA and Purdue University under the guidance of Prof. Kaushik Roy. He earned his Ph.D. in Electrical and Computer Engineering from Purdue University in December 2025. His doctoral research concentrated on physics-based multi-grain phase-field models of hafnium-zirconium-oxide ferroelectric devices, providing critical insights into their fundamental physical mechanisms and device properties. His current research focuses on Design-Technology Co-Optimization (DTCO) and System-Technology Co- Optimization (STCO) frameworks. Specifically, he investigates the hardware acceleration of artificial intelligence workloads and large language models (LLMs). His approach focuses on optimizing emerging memory technologies and their integration into advanced computing architectures via 3D bonding and monolithic integration.

23

A. Kosta, K. Roy

SG-CPG: Severity-Gated Central Pattern Generators for Adaptive Quadruped Locomotion under Continuous Actuator Degradation

AbstractLegged robots are poised to carry increasingly capable AI into homes, workplaces and open terrain, where they must keep working for long periods. Over such deployments their actuators can weaken, and an overheating or underpowered motor can no longer deliver the torque the gait needs. Existing fault-tolerant controllers majorly model faults such as locked joints or complete power loss, but not a lowered torque ceiling that stays invisible to proprioception until the gait demands the capped torque. Animals, in contrast, sense an injured or weakened limb and readily adapt their walking pattern to accommodate it. Spinal central pattern generators (CPGs), which produce the basic locomotor rhythm, contribute to this adaptation. Inspired by CPGs, we propose SG-CPG, a framework for estimating and adapting to continuous actuator degradation under both torque-ceiling and gain-scaling faults. SG-CPG freezes a healthy CPG policy and adds a residual policy that acts only under a fault and modulates the actions based on fault-severity. We also perform a controlled study of severity estimators showing that the fault mechanism decides which estimator to use. In simulation, SG-CPG keeps a trot at every tested severity under both fault mechanisms. With a joint reduced to 5% of its strength, it survives ≥97% of straight-line episodes at a forward velocity tracking error of ≤9% of the commanded 1.0 m/s. For omnidirectional commands, it still survives 83-94% of episodes with an error of 9-17%. On a real Unitree Go2 with an emulated calf torque ceiling lowered by up to 93%, SG-CPG survives 28 of 29 forward and turning runs (97%). Across these runs, commanded up to 1.0 m/s and 1.0 rad/s, the forward-speed and yaw-rate errors average 26% and 31% of the command, respectively.
BioAdarsh Kumar Kosta is a Ph.D. candidate in Electrical and Computer Engineering at Purdue University, working with Prof. Kaushik Roy in the Nanoelectronics Research Laboratory. His research explores bio-inspired computation for efficient perception, decision-making, and control, spanning event-based vision, spiking neural networks, and reinforcement learning, with recent work extending to fault-tolerant algorithms for legged robot locomotion. He has published in venues including ICRA, IROS, and Frontiers in Science, and previously interned at Samsung AI Center, where he developed a patented acoustic-based tactile sensing system for robots. Adarsh is broadly interested in building algorithms that draw on biology's efficiency ? sparse, adaptive computation that performs well under real-world constraints.

24

M. Mukherjee, K. Roy

MIRAGE:MRAM-Based Near ADC-Less Compute-In-Memory Macro for Deep Learning Acceleration

AbstractNon-volatile memory (NVM) based Compute-in-Memory (CiM) architectures have emerged as a promising compute primitive for accelerating deep neural networks (DNNs) by performing in-situ matrix?vector multiplications (MVMs). Among various NVMs, STT-MRAM (Spin Transfer Torque based Magnetoresistive Random Access Memory) shows potential due to its high endurance, low energy consumption and high density. However, existing STT-MRAM CiM designs typically rely on multi-bit analog-to-digital converters (ADCs) at the peripherals to digitize accumulated bit-line currents. While enabling high-precision computation, ADCs add substantial energy, latency, and area overheads. To alleviate such problems, we propose a system-technology co-design approach to a Near ADC-Less CiM design with ternary partial-sums called MIRAGE. The accuracy is maintained by considering hardware level partial sum quantization in the training loop. Specifically, we develop an STT-MRAM based CiM macro which features differential bitcells and an adaptive threshold sensing that is amenable to the requirements posed by ternary partial-sum quantization. We do a thorough energy, area, latency, and sense margin analysis along with robust benchmarking against conventional 1T-1MTJ (1 transistor-1 Magnetoresistive Tunnel Junction) based MRAM CiM. The proposed CiM macro occupies ∼ 20% less area, consumes 1.8× less MVM energy and shows 5× better latency with improved distinguishability compared to 1T-1MTJ CiM macro while achieving better accuracy.
BioMainakh Mukherjee received his B.Tech. from the West Bengal University of Technology in 2021 and completed his M.Tech. in VLSI Design at the Indian Institute of Engineering Science and Technology (IIEST), Shibpur in 2024. During his master's, he worked on modeling non-linearities and designing mixed-signal circuits, including data converters such as SAR and flash ADCs. He joined Purdue University in Fall 2024 to pursue a Ph.D. under Prof. Kaushik Roy. His research interests lie in device?circuit co-design, mainly circuit designing using emerging memory technologies for compute-in-memory at advanced technology nodes.

25

A. Pranta, K. Roy

To CiM or Not to CiM? A System-Level Study of Digital and Analog Compute-in-Memory

AbstractThe widening gap between compute throughput and memory bandwidth has made data movement a major performance and energy bottleneck in modern AI accelerators. Compute-in-memory (CiM) addresses this by performing computation within or near the memory arrays that store the model weights, reducing weight movement and exploiting the arrays' inherent parallelism. However, these macro-level gains may not translate into system-level benefits once peripheral overheads, compute utilization, buffering, and off-chip memory traffic are included. We examine when CiM is advantageous over conventional compute, and under what conditions that advantage holds by comparing digital CiM (DCiM), ADC-based analog CiM (ACiM), near-ADC-less ACiM, and a conventional digital MAC array. All compute primitives are implemented in 22-nm GlobalFoundries FDX technology, and the CiM primitives are based on SRAM arrays. At the macro level, CiM provides 12.4×-67.2× higher area efficiency and up to 3.2× higher energy efficiency than the conventional digital MAC baseline. However, full-precision ACiM remains energy-inefficient due to ADC overhead. Reducing ADC precision from 6-bits to 4-bits improves ACiM's area and energy efficiency by 1.8× and 1.57×, respectively, while near-ADC-less ACiM achieves the highest area efficiency (27.43 TOPS/W/mm2) and energy efficiency (4.83 TOPS/W). Integrating these primitives into a common accelerator framework under iso-area constraints shows that CiM's system-level benefit is governed by the on-chip compute to off-chip memory traffic ratio. Near-ADC-less ACiM achieves up to 3× speedup on ResNet-18. DCiM and ADC-based ACiM variants achieve 54%-105% speedup over the digital MAC baseline on compute-intensive, prefill-heavy LLM workloads, shrinking to only 1%-2% on the most decode-dominated workloads. As off-chip DRAM traffic dominates, the compute-primitive choice has little effect on system performance. Overall, CiM is most effective when its area efficiency translates into balanced system resources with sufficient array utilization and data reuse.
BioAyan earned his bachelor's degree in Electrical and Electronic Engineering from Bangladesh University of Engineering and Technology (BUET). He began his PhD in Electrical and Computer Engineering at Purdue University in Fall 2024 under the supervision of Professor Kaushik Roy. His research focuses on memory-centric hardware and architectures for AI and neuromorphic computing.

26

A. Das, V. Raghunathan

COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems

AbstractMultimodal AI (MMAI) systems are becoming integral to safety-critical edge applications, from autonomous driving and robotics to smart healthcare and surveillance, where real-time perception must operate under stringent energy and compute constraints. These systems comprise distinct sensing units, such as cameras, microphones, and inertial sensors, each forming a separate subsystem with unique capabilities and constraints. As environments evolve, modality relevance shifts and redundancy emerges, requiring dynamic adaptation to maintain quality of service while extending system lifetime. Existing methods optimize individual subsystems or static inference paths, lacking responsiveness to contextual changes; coordinating adaptation across these heterogeneous subsystems remains essential yet largely unexplored. We introduce COSMOS, a runtime framework for distributed multimodal edge nodes. Its central contribution is a closed-loop control structure that uses predicted semantic context and measured temporal stability to jointly govern modality activation, inference frequency, and subsystem-level approximation. COSMOS instantiates this control loop through three mechanisms: Scene/State-aware Modality Selection (SMS), Temporal Locality- driven Decision and Regulation (TLDR), and modality-aware Design Space Exploration (DSE). The DSE heuristic coordinates compute, audio preprocessing, sensing, memory, and communication knobs under a unified energy-quality objective. Validated on an FPGA-MCU prototype (Intel Agilex-5 primary + ESP32 auxiliary) for Audio-Visual Event Localization, COSMOS extends system lifetime by 7.73x with <5% accuracy loss, achieving up to 1.77x primary and 3.01x auxiliary energy savings across six transformer backbones.
BioArghadip Das earned his B.E. degree in Electronics and Telecommunication Engineering from Jadavpur University, Kolkata, India, in 2020. Currently, he is pursuing his Ph.D. at the School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA, under the guidance of Prof. Vijay Raghunathan. During his undergraduate studies, his research centered on designing low-power hardware accelerators. At Purdue, Mr. Das's primary research interests lie in deep learning for resource-constrained embedded systems. His current work focuses on creating innovative approximate systems to enable collaborative distributed inference in tiny, resource-limited computing systems and IoT devices. His approach emphasizes energy efficiency, adaptability, and enhanced performance through strategic approximations across various levels of system design. Mr. Das has been recognized as a Distinguished Design Automation Conference Young Fellow twice, in 2022 and 2024. He is also the recipient of the University Gold Medal from Jadavpur University and the prestigious ECE STAR Fellowship from Purdue University.

27

H. Kim, V. Raghunathan

TESSERA: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs

AbstractAI model architectures are diversifying rapidly. Although dense matrix multiplication still underlies today's CNNs and transformers, emerging architectures (state-space models, long convolutions via the fast Fourier transform (FFT), Kolmogorov-Arnold networks, and spiking networks) are not multiply-accumulate (MAC) dominated; they spend much of their computation on vector and non-MAC primitives that homogeneous, MAC-centric neural processing units (NPUs) serve poorly. This has motivated heterogeneous NPUs (HPUs) built from non-identical tiles. Prior heterogeneous designs, however, vary only one or two coarse knobs (typically MAC precision or array size) and are evaluated on narrow workloads, and no existing framework supports HPU design at a fine granularity, where tiles differ across many architectural dimensions at once.

We present TESSERA, an analytical simulator and design-space-exploration (DSE) framework for HPU microarchitecture design. TESSERA searches the joint space of tile-level heterogeneity: beyond array size and precision, it varies the knobs along which tiles can differ, including tile- type composition (large Big, small Little, and non-MAC Special-Function tiles), dataflow, sparsity mode, MAC engine type, and special-function units for non-MAC operators (FFT, spiking-integrate, polynomial). Unlike prior simulators, which model a single homogeneous tile type, TESSERA models non-MAC tiles with their own energy, area, and timing models and maps operators across a mix of tiles with a heterogeneity-aware compiler. A multi-seed pipeline that pairs a stratified sweep with genetic-algorithm refinement returns Pareto-optimal designs, with cost models calibrated to a 7 nm node and cross-validated against NVIDIA's Deep Learning Accelerator (NVDLA). Across a 20-workload suite, the best general-purpose HPU found by TESSERA (200 mm 2 Big+Little+Special-Function) achieves +46.91% mean iso-area energy savings over the best iso-area homogeneous baseline.
BioHoseok Kim is a Ph.D. student in Electrical and Computer Engineering at Purdue University, advised by Professor Vijay Raghunathan. His research focuses on energy-efficient edge computing and NPU architecture. He previously earned his M.S. at Korea University, where his work on DNN inference efficiency won the Best Paper Award at ISLPED 2024.

28

I. Ahmed, S. Gupta

1.58-b FeFET-Based Ternary Neural Networks: Achieving Robust Compute-In-Memory With Weight-Input Transformations

AbstractCompute-in-memory (CiM)-enabled Ternary weight neural networks (TWNs), with weights ∈ {-1, 0, 1}, are attractive for edge-AI as they offer high energy-efficiency while maintaining acceptable inference accuracies. To that end, multi-weight Ferroelectric transistors (FeFETs) can enable highly scalable TWN-CiM acceleration. However, standard 1T-FeFET-based CiM is vulnerable to hardware non-idealities at deeply scaled technology nodes, necessitating custom 2T-differential designs to improve robustness at the cost of area/energy. In this work, we explore three FeFET- based TWN-CiM solutions: (i) 1T (1.58-bit 3-level storage), (ii) 2T-differential (1-bit storage), (iii) 2T- pull-up-pull-down (1-bit storage) designs; and propose static-weight transformation (WT) and static- weight-dynamic-input transformation (WIT) for increasing the computational robustness of 1T- FeFET. Using our rigorous phase-field FeFET models and non-ideality-aware inference simulator, we conduct comparative analyses of these designs, showing that WT and WIT significantly recover inference accuracy for 1T-FeFET (up to 73.61% for ResNet18 TWN-CiM on CIFAR100), making it competitive against 2T-differential-FeFET (76.4%), while offering 1.98x and 1.91x reductions in macro-area and macro-energy, respectively.
BioImtiaz Ahmed is a Ph.D. candidate in Electrical and Computer Engineering at Purdue University, conducting research at the Integrated Circuits & Devices Lab (ICDL). His research revolves around efficient AI hardware design, with a keen focus on emerging memory technologies, in-memory computing, and robust deep neural network accelerator design. His work aims to develop energy-efficient, high-performance, and scalable hardware for next-generation of AI accelerators. Imtiaz is also an SRC Research Scholar affiliated with CoCoSys: Center for the Co-Design of Cognitive Systems, under Joint University Microelectronics Program 2.0 (JUMP2.0)- a Semiconductor Research Corporation (SRC) program sponsored by the Defense Advanced Research Projects Agency (DARPA). He completed his B.Sc. in Electrical and Electronic Engineering from Bangladesh University of Engineering and Technology (BUET), Dhaka, Bangladesh in 2019, with a concentration in Electronics. Prior to joining Purdue University, Imtiaz worked as a Lecturer at BRAC University (Bangladesh) for 2.5 years, where he coordinated multiple undergraduate courses and contributed to curriculum redesign for integrating Outcome- Based Education (OBE) framework, to align courseworks with industry and research advancements.

29

S. Bhattacharjee, S. Gupta

WEBCRAFT: Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework

AbstractThe deployment of deep neural networks (DNNs) on compute-in-memory (CiM) accelerators offers significant energy savings and speed-up by reducing data movement during inference. However, the reliability of CiM-based systems is challenged by stuck-at-faults (SAFs) in memory cells, which corrupt stored weights and lead to accuracy degradation. While closest value mapping (CVM) has been shown to partially mitigate these effects for multibit DNNs deployed on bit-sliced crossbars, its fault tolerance is often insufficient under high SAF rates or for complex tasks. In this work, we propose two training-free weight transformation techniques, sign-flip and bit-flip, that enhance SAF tolerance in multi-bit DNNs deployed on bit-sliced crossbar arrays. Sign-flip operates at the weight-column level by selecting between a weight and its negation, whereas bit-flip provides finer granularity by selectively inverting individual bit slices. Both methods expand the search space for fault-aware mappings, operate synergistically with CVM, and require no retraining or additional memory. To enable scalability, we introduce a look-uptable (LUT)-based framework that accelerates the computation of optimal transformations and supports rapid evaluation across models and fault rates. Extensive experiments on ResNet-18, ResNet-50, and ViT models with CIFAR-100 and ImageNet demonstrate that the proposed techniques recover most of the accuracy lost under SAF injection, often restoring performance to within 1-2% of ideal baselines. Hardware analysis shows that these methods incur negligible overhead, with sign-flip leading to sub-2% energy, latency, and area cost, and bit-flip providing higher fault resilience with modest overheads. These results establish sign-flip and bit-flip as practical and scalable SAFmitigation strategies for CiM-based DNN accelerators.
BioSaptashwa Bhattacharjee received the B.Tech. degree in Electronics and Electrical Communication Engineering from the Indian Institute of Technology Kharagpur, India, in 2025, and is currently pursuing the Ph.D. degree in Electrical and Computer Engineering at Purdue University. His research interests include in-memory computing, AI accelerators, and hardware-aware machine learning.

30

D. Kim, S. Gupta

SCION: A Comprehensive Simulation Framework for Charge-Based In-Memory Computing for Rapid Evaluation of Hardware Non-Idealities and DNN Accuracy

AbstractCharge-based in-memory computing (IMC) has shown great potential in achieving higher computational robustness compared to current-based IMC. However, it suffers from its own non-idealities such as parasitic capacitive coupling. Accurately evaluating these effects requires time-intensive SPICE simulations, making it challenging to conduct cross-layer exploration. To overcome these limitations, we propose SCION, a PyTorch-based framework that rigorously models hardware non-idealities in charge-based IMC and integrates them directly with DNN inference for rapid accuracy evaluation. We show that SCION predicts the IMC output with more than 99% accuracy with respect to SPICE while offering four orders of magnitude speedup. We demonstrate the capability of our framework by analyzing an SRAM-based charge-IMC accelerator deploying ResNet-50 and ViT-small DNNs with CIFAR-100 dataset. We show how inter-column capacitive coupling leads to data-dependent non-idealities, severely impairing the inference accuracy. We also explore techniques to mitigate non-idealities using SCION. To that end, we propose alternate column enablement (ACE) to eliminate inter-column coupling. Our results show that compared to the baseline design and another non-ideality mitigation approach based on ground shielding, ACE achieves significant improvement in sense margin and near-software accuracy under nominal conditions. Further, compared to ground shielding, ACE exhibits a higher tolerance to analog-to-digital converter (ADC) noise and superior voltage scalability.
BioDoug Hyun Kim is currently pursuing the Ph.D. degree in electrical and computer engineering at Purdue University, West Lafayette, IN, USA, where he is a Graduate Research Assistant with the Integrated Circuits and Devices Laboratory under the supervision of Prof. Sumeet Kumar Gupta. His research interests include compute-in-memory architectures, AI hardware accelerators, emerging memory technologies, mixed-signal and charge-domain computing, and hardware-software co-design for efficient artificial intelligence systems. His current research focuses on cross-layer modeling and optimization of compute-in-memory systems, including circuit nonidealities, and device-circuit-algorithm co-design for deep neural-network inference.

31

E. Berscheid, A. Raghunathan

Optimizing Data Placement for Near-Memory Acceleration of Recommendation Systems

AbstractRecommendation systems are widely-used machine learning models that process user data to recommend content. These models heavily rely on looking up and summing embedding vectors from large tables, which is a memory-bound operation that bottlenecks performance. Near-memory processing (NMP) accelerates embedding look- ups by employing processing units within the DRAM system, which compute multiple vector operations in parallel and reduce data volume before transferring the results to the host. We observe that data placement has a profound impact on NMP efficiency, an effect that is not fully addressed by prior work. To that end, we make the following contributions: First, we propose a memory-efficient statistical method for profiling embedding table look-ups to find commonly co-occurring sets of vectors. We recognize that vector replication holds potential for enhancing vector co-locality, and not only load balance as explored in prior works. Accordingly, our second contribution is a replicate resolution policy that enables a principled and tunable trade-off between load balance and vector co-locality to maximize performance. Next, we redistribute non-replicated embedding vectors in physical DRAM to collocate (or separate) vectors that are likely (or unlikely) to belong to the same access set. DRAM simulations demonstrate up to 3.41x speedup and 3.01x energy savings over prior methods.
BioElijah Berscheid is a PhD student in ECE, advised by Professor Anand Raghunathan. His research interests relate to efficient data center scale AI, including systems, architecture, and algorithms. His previous project aimed to optimize near-DRAM computing performance for memory-bound recommendation system workloads (under review at TCASAI). His current work explores performance gaps due to dynamic irregularities in sparse decode attention for long-context workloads to improve performance on chiplet-based accelerators. Elijah has interned at the University of Minnesota (control systems and data processing), Western Digital (HDD servo firmware), Broadcom (3nm ROM verification), and NVIDIA (efficient LLM inference on future MCM-GPUs).

32

D. Kim, S. Hsu, V. Jain

Hardware Efficient VLA Acceleration using Vector Quantization for Physical AI

AbstractVision-Language-Action (VLA) models have emerged as a promising framework for general-purpose robotic control, leveraging the rich representations and generalization capabilities of foundation models such as Vision-Language Models (VLMs). However, their large parameter scale and computational demands pose significant challenges for deployment on resource-constrained robotic platforms, where low-latency inference is critical for real-time interaction. In this work, we investigate several low-bit quantization methods for accelerating VLA inference, focusing on both KV- cache and weight compression. At low-bit precision, quantization quality degrades significantly due to outliers concentrated in certain activation channels. Rotation-based transformations redistribute these outliers across multiple dimensions, reducing their concentration and improving quantization quality. First, we explore TurboQuant for KV-cache quantization and evaluate MSE-optimal and inner-product-optimal quantization schemes using the π0 model on VLABench. We further investigate QuaRot for weight quantization, which applies Hadamard-based rotations to input activations to enable INT4 GEMM in major Transformer components, including gated FFNs and QKVO projection layers. With QuaRot-based INT4 GEMM, we achieve up to a 1.45× inference speedup. Finally, we examine Leech Lattice Vector Quantization (LLVQ), which leverages the sphere-packing properties of the 24-dimensional Leech lattice and uses its structured lattice points as a predefined codebook, eliminating the need for codebook retraining while achieving up to 8× weight compression on Gemma-2B. Together, these approaches explore complementary opportunities to reduce the memory and computational overhead of VLA models when deployed in real-time robotic systems.
BioShun-Hsiu (William) Hsu received his B.S. dual degree in Aeronautics and Astronautics and Electrical Engineering, followed by his M.S. degree in Electrical Engineering, from National Cheng Kung University (NCKU), Tainan, Taiwan. He is currently a Ph.D. student in the Elmore Family School of Electrical and Computer Engineering at Purdue University, West Lafayette, IN, advised by Prof. Vikram Jain. His research focuses on Physical AI, digital integrated circuit (IC) design, and compute-in-memory (CIM) AI accelerators.

Dongyoon Kim received his B.S. degree in Electronic and Electrical Engineering from Sungkyunkwan University (SKKU), Suwon, South Korea in 2026. He is currently a Ph.D. student in the Elmore Family School of Electrical and Computer Engineering at Purdue University, West Lafayette, IN, advised by Prof. Vikram Jain. His research focuses on Domain Specific Accelerators and Embodied AI.

33

H. Cho, Y. Kim

CO-MAC: A Center-Out Ordered Stochastic MAC for Low-Latency Inference

AbstractStochastic computing offers efficient approximate arithmetic that aligns well with error-tolerant machine learning workloads, but its deployment is limited by long bitstream latency in stochastic multiply-accumulate (MAC) units. Prior work reduces MAC latency through deterministic bitstream generation and differential accumu lation, but these methods do not fully exploit the statistical property of convolution weights. This work presents a novel stochastic MAC architecture named CO-MAC, which employs center-out weight ordering and an enhanced convolution engine design to reduce effective computation cycles while maintaining high accuracy. The method sorts weights by magnitude, reuses the incremental differ ences in magnitudes, and applies sign handling after accumulation. This shortens counter activity, maintains accuracy with long ef fective bitstreams, and simplifies the MAC hardware by avoiding bidirectional counters. Across convolutional neural network work loads, CO-MAC decreases MAC latency by up to 54.8% compared to prior stochastic MAC architectures, while preserving accuracy and hardware simplicity.

34

D. Kim, Y. Kim

Rethinking DRAM Protection for Hybrid-Bonded 3D Stacks

AbstractMemory reliability now limits large AI system availability. Hyperscale datacenter operators report memory as the largest identified cause of unplanned interruptions, and bandwidth demand is pushing that memory vertically. Stacking the array on the logic die it feeds runs it hotter, leaves each channel with far fewer banks, and returns a much larger block per access. That larger block changes what the errors look like. Bits that once failed together now land apart, beyond the reach of the symbol codes deployed today. Covering that difference in hardware takes area, and every channel pays it again.

We propose and investigate a protection scheme reorganized around a single controller-side code. The same code corrects the errors that arrive and folds out the cells that stay bad. Freed of standing faults, most accesses need only shallow correction. The rest is handled as it arises. The trade-off moves from area to time, and the worst case is paid only when it occurs.
BioDoh Yon Kim received his B.S. degree in Electrical and Computer Engineering from Sungkyunkwan University, Republic of Korea, in 2026. He is currently pursuing the Ph.D. degree with the Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, under the mentorship of Prof. Younghyun Kim. His research interests include memory reliability, error correction and runtime repair for 3D- stacked DRAM, and memory systems for AI accelerators.

35

M. Chen, H. Li

Cross-Layer Heterogeneous Memory Systems for Scalable AI and Optimization Workloads

AbstractEmerging AI and optimization workloads place increasing pressure on the memory hierarchy through growing capacity, bandwidth, and data- movement demands. At the same time, emerging memory technologies provide new opportunities to redesign how data is stored, moved, and processed across the hierarchy. This work explores cross-layer heterogeneous memory architectures that match the characteristics of different memory technologies with workload requirements. We investigate eDRAM?RRAM heterogeneous memories for hyperdimensional computing and zeroth-order fine-tuning, stochastic MTJ compute-in- memory for Ising optimization, and HBM?HBF systems for large language model inference. Across these systems, we study how device and circuit characteristics translate into architectural tradeoffs in data placement, computation, bandwidth utilization, and energy efficiency. Together, these efforts establish a cross-layer methodology for building scalable memory hierarchies for emerging AI and optimization workloads.
BioMufeng Chen received the B.S. degree in optics and electronic information engineering from Huazhong University of Science and Technology, Wuhan, China, in 2021. From 2021 to 2024, he was a Research Assistant with Zhejiang University, Hangzhou, China, where he worked on hardware/software co-design and system integration for emerging computing systems. He is currently pursuing the Ph.D. degree with the Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, USA. His research interests include heterogeneous memory systems, memory-centric architectures, hardware/software co-design, and system-level integration for data- intensive computing.

36

Z. Wang, H. Li

'CMOS+X' Monolithic 3D-ICs for Heterogeneous Cognitive Computing Systems

AbstractTwo-dimensional CMOS scaling alone can no longer simultaneously deliver dense on-chip memory and new device-level functionality demanded by emerging cognitive computing workloads. The "CMOS+X" approach addresses this by keeping a single foundry Si CMOS platform and sequentially integrating heterogeneous 'X' device tiers in the back-end-of-line (BEOL), forming monolithic 3D integrated circuits (M3D-ICs). This work explores two complementary CMOS+X tracks. First, we develop CMOS+OSFET hybrid structured-ASIC fabrics, where BEOL oxide-semiconductor transistors enable capacitor-less eDRAM and reusable logic/memory tiles on top of a fixed pre-fabricated CMOS base, guided by a device?circuit?system co-design framework that reduces NRE cost and accelerates application-specific deployment. Second, we design a CMOS+3D-MTJ probabilistic computing prototype, where BEOL stochastic magnetic tunnel junctions serve as physics-based true-random entropy sources for FEOL digital compute cores, realizing Boltzmann-machine-based invertible logic that unifies forward arithmetic and reverse inference within one hardware fabric. Across both tracks, we study how BEOL device characteristics translate into circuit- and system-level tradeoffs in density, energy efficiency, and functional heterogeneity. Together, these efforts establish a foundry-compatible M3D design pathway toward heterogeneous cognitive computing systems.
BioZeshu Wang received the B.E. degree with honors in microelectronics from Wuhan University, Wuhan, China, in 2022, and the M.E. degree in electrical engineering from Tsinghua University, Beijing, China, in 2025, where he worked on heterogeneous integration of perovskite-based optoelectronic devices and systems. He is currently pursuing the Ph.D. degree with the Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA, advised by Prof. Haitong Li. His research interests include monolithic 3D integrated circuits, memory-centric computing with BEOL technologies, probabilistic computing hardware, and device-circuit-system co-design for heterogeneous computing systems.

37

C. Baddouh, T. Rogers

Beyond PTX: Generating and Modifying Native GPU Code for AI Systems Research

AbstractEfficient AI kernels depend on precise control over GPU execution, yet NVIDIA's native instruction set, SASS, remains largely undocumented, restricting code generation and optimization below PTX. We present a framework for systematically generating, modifying, and evaluating native GPU code. The framework extracts machine-readable instruction specifications from NVIDIA tooling and uses encoding constraints to generate candidate instructions for validation on real hardware. Its binary-editing interface enables researchers to replace and reorder instructions, modify operands and scheduling controls, and patch existing kernels directly. Across seven architectures spanning Volta through Blackwell, we validate 539-1,338 instruction classes per target. We demonstrate these capabilities through instruction timing and contention experiments, as well as direct RT Core traversal from a standalone compute kernel. This work provides an experimental foundation for developing and evaluating code-generation strategies, exploring instruction sequences beyond those emitted by existing toolchains, and investigating low-level optimizations for AI kernels.

38

Y. Liu, T. Rogers

Characterizing Communication-Computation Interference in LLM Training and Inference

AbstractOverlapping communication with computation is now standard practice in large-scale LLM training and inference. However, the interference between the two has been studied mostly at the application level, lacking a system-wide view and architectural insight. We characterize communication?computation co-execution in both LLM training (Megatron-LM) and inference (vLLM), across parallelism strategies (DP, FSDP, TP, PP, and EP) on NVIDIA H200 GPUs, using hardware counters that capture SM occupancy, L2, HBM, and NVLink behavior and attributing them to individual overlapped kernels. Our in-depth analysis of the resulting interference reveals several significant observations about which kernels are damaged and why, and yields practical guidelines for both software and hardware optimization.

39

J. Pan, T. Rogers

Accel-Sim 2.0: Designing Asynchronous, Distributed GPUs in the AI Era

AbstractThe rapid evolution of machine learning has fundamentally altered GPU design, pushing architectures toward Multi-Chip Module topologies, asynchronous execution, and persistent, multi-phase kernel behaviors. Despite these profound shifts, the architecture community lacks cycle-level simulation frameworks that can natively capture the physical non-uniformity of contemporary GPUs and the massive scale of state-of-the-art AI workloads. To close this gap, this paper introduces a robust simulation infrastructure that accurately models Ampere, Hopper and Blackwell. Specifically, we evaluate the microarchitectural impact of off-core memory bandwidth expansion and asynchronous synchronization overheads. Furthermore, our analysis of intra-GPU NUMA effects reveals that while LLMs inherently tolerate spatial latency, the resulting data replication across partitions diminishes effective cache capacity, severely degrading locality-dependent workloads like graph analytics and CNNs. Finally, we demonstrate that inter-GPU memory prefetching is severely constrained by register file capacity limits. We present Accel-Sim v2.0, a cycle-level simulation framework validated against Ampere, Hopper, and Blackwell architectures. Our model consistently achieves a 99\% Pearson correlation coefficient across contemporary silicon, alongside mean absolute cycle errors of 13.5% on Hopper H100 and 8.9% on Blackwell B200, respectively. Using Accel-Sim v2.0, we conduct a series of architectural investigations to expose new design trajectories surrounding asynchrony, chiplets, increased off-core bandwidth, and inter-GPU prefetching.

40

A. Nallathambi, A. Raghunathan

A3D: Agentic AI flow for Autonomous Accelerator Design

AbstractAccelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite advances, such as high-level synthesis (HLS), designing accelerators for complex applications still remains highly labor-intensive, demanding considerable expertise in understanding workloads to be accelerated, hardware design, micro-architecture, and EDA tool usage, posing challenges for application domain experts. Therefore, most accelerator solutions are limited to applications with a regular predictable dataflow. Advances in AI have enabled agents that perform autonomous planning, reasoning, execution and reflection, leading to unprecedented potential for automation through agentic AI. We present A3D, an Agentic AI flow for end-to-end Automation of hardware Accelerator Design. A3D automates workload analysis, performance bottleneck identification, code refactoring for HLS compatibility and micro-architecture generation. A3D also generates diverse accelerator designs by automatically exploring the speed-area tradeoff space. Recent efforts have explored the use of AI for specific tasks such as design space exploration in HLS, leaving several tasks to still be performed manually. A3D addresses the challenges in applying modern LLMs to accelerator design by judiciously partitioning tasks among specialist agents, orchestrating process loops with specialist and verifier agents, utilizing pre-existing and custom tools, and employing agentic RAG for codebase and proprietary EDA tool documentation exploration.

41

R. Maity, K. Roy

HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads

AbstractAnalog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNN). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from crossbars leads to substantial power and area overhead. Moreover, the high area overhead of ADCs constrains the throughput due to the limited number of ADCs that can be integrated per crossbar. An approach to mitigate this issue involves the adoption of extreme low-precision quantization (binary or ternary) for partial sums. Training based on such an approach eliminates the need for ADCs. While this strategy effectively reduces ADC costs, it introduces the challenge of managing numerous floating-point scale factors, which are trainable parameters like DNN weights. These scale factors must be multiplied with the binary or ternary outputs at the columns of the crossbar to ensure system accuracy. To that effect, we propose an algorithm-hardware co-design approach, where DNNs are first trained with quantization-aware training. Subsequently, we introduce HCiM, an ADC-Less Hybrid Analog-Digital CiM accelerator. HCiM uses analog CiM crossbars for performing Matrix-Vector Multiplication operations coupled with a digital CiM array dedicated to processing scale factors. This digital CiM array can execute both addition and subtraction operations within the memory array, thus enhancing processing speed. Additionally, it exploits the inherent sparsity in ternary quantization to achieve further energy savings. Compared to an analog CiM baseline architecture using 7 and 4-bit ADC, HCiM achieves energy reductions up to 28% and 12%, respectively.
BioRiddhiman received his Bachelor of Technology (Honours) degree in 2025 from the Indian Institute of Technology Kharagpur in Electronics and Electrical Communication Engineering. During his undergraduate years, he worked at the Space Application Centre (SAC), ISRO, as an Undergraduate Research Fellow and contributed to the tape-out of a SERDES IC. He completed his research internship at NRL, Purdue, in the summer of 2024. He was selected for a summer internship at Texas Instruments in Bangalore, India, as an analog design intern. He also contributed to the design of an SRAM-based CiM hardware in FinFET technology. After completing his undergraduate studies, he joined Purdue University in the fall of 2025 as a Ph.D. student in Electrical and Computer Engineering under the supervision of Professor Kaushik Roy. His research interest lies on post-CMOS device-circuit co-integration and hardware-algorithm co-design. Apart from research, Riddhiman is interested in painting and traveling.

Hackathon Finalists

42

E. Yoon

Agents for IC Physical Design

43

X. Dai, S. Mohammadi

Analog Mixed-Signal Layout Agent

44

A. Grabka, R. Oum

AntoRo / Training Neural Networks Robust to Soft Faults in ReRAM Crossbar Arrays

45

S. Cowen, R. Spitzenberger

Carbon Compute

46

J. Balasubramanian, T. Panchagnula, S. Rank

Compute Commons

47

R. Brar, H. Koh, A. Sridhar, K. Zisko

KERA

48

A. Arens, R. Harrison, W. Tao

Laser Drone Detection

49

S. Akella, A. Alam, S. Velmurugan

RouteEdge

50

J. Urgiles, C. Wai

Tiered Wakeup Visual Classification Cascade

Accomodations

  • The Union Club Hotel at Purdue University - Website
  • Hilton Garden Inn West Lafayette Wabash Landing - Website
  • Hampton Inn & Suites West Lafayette - Website

Contact Us

For questions, please email CHIPSandAI@purdue.edu.