| AI for CHIPS |
| Number |
Authors |
Title |
|
1 |
M. Khan, S. Gupta |
Repurposing Batch Normalization for Autonomous and In-Situ Detection and Mitigation of Weight Corruptions in CNN Accelerators
AbstractIn safety-critical deployments, AI hardware must remain reliable against a broad spectrum of threats such as radiation-induced upsets, aging, hard faults, and adversarial attacks (e.g., the progressive bit-flip attack (PBFA)). All of these corrupt stored weights while the chip keeps producing confident (but inaccurate) predictions. Detecting and mitigating such weight perturbations is therefore crucial for safety-critical platforms. To that end, we propose an on-chip, batch normalization (BN)-based technique that continually senses shifts in activation statistics to detect and mitigate a wide variety of weight corruptions on the fly, including random and localized faults as well as adversarial bit flips. The technique operates in two phases: (1) offline pre-characterization of the relationship between activation shifts and inference accuracy drop, and (2) on-chip runtime detection and mitigation of weight corruptions. It is designed so that benign perturbations stay silent while harmful faults are flagged. Upon detection, the flagged layer is re-centered toward its stored clean reference within the same forward pass. The technique is fully autonomous, eliminating the need for host communication, operation halts, or access to fine-tuning data. If the residual shift after mitigation indicates that accuracy has fallen below a user-set floor, a held-out watcher aborts the inference. Evaluated on ResNet-20/50 and MobileNetV2 for CIFAR-10/100, our approach detects harmful corruptions with >99% precision across all fault types. Furthermore, it recovers accuracy from 10% to 85.88% under 0.5% random bit flips (ResNet-50/CIFAR-10), up to 84% for localized faults (MobileNetV2/CIFAR-10), and from random-guess accuracy to 80-83% under PBFA (ResNet-20/CIFAR-10), entirely on chip and within the same inference pass. This comes with <2% latency cost and 0.53% computation overhead.
BioMarzia Khan is a third-year Ph.D. student in the Elmore Family School of Electrical and Computer Engineering at Purdue University, advised by Prof. Sumeet Kumar Gupta. She received her B.Sc. in Electrical Engineering from Bangladesh University of Engineering and Technology (BUET) in 2022. Her research focuses on runtime hardware fault detection and mitigation for quantized DNN accelerators.
|
|
2 |
J. Balasubramanian, A. Raghunathan |
PowerScope: ML-based Intra-Cycle Power Estimation
AbstractPower estimation at sub-clock-cycle temporal resolutions is critical for tasks such as power delivery network (PDN) design, dynamic voltage droop analysis, and pre-silicon power side-channel security evaluation. Designers commonly rely on commercial post-layout gate-level power analysis tools for these tasks, but these flows are
computationally expensive and scale poorly with design size and workload length. Machine learning (ML)-based power estimation frameworks have shown promise in accelerating power estimation, but prior efforts only address average power or per-cycle power estimation. We propose PowerScope, the first ML-based intra-cycle power estimation framework. PowerScope operates purely on RTL simulation traces at inference time, eliminating the
need for post-layout gate-level simulation and power analysis per workload. Across a diverse benchmark suite, PowerScope achieves 3.3% median and 6.7% mean absolute percentage error compared to commercial post-layout gate-level power estimates while running ∼571× faster. We further demonstrate that PowerScope's predictions can be reliably used for the downstream task of pre-silicon power side-channel leakage assessment.
BioJayanth Balasubramanian received his B.Tech and M.Tech in Electrical Engineering from Indian Institute of Technology Madras in 2025. He is currently a PhD student at Purdue University advised by Prof. Anand Raghunathan. His research focuses on the intersection of ML and EDA, with a focus on power estimation for SoCs.
|
|
3 |
G. Yarramneni, A. Raghunathan |
Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs
AbstractModern edge SoCs integrate heterogeneous processing units?including CPUs, GPUs, and
NPUs?with different performance and energy characteristics. Deploying AI inference workloads
on these platforms requires jointly selecting workload mappings and processing-unit
configurations, creating a combinatorial design space that makes exhaustive search impractical.
Existing black-box design-space exploration methods typically guide the search using only
sparse objective values, providing little insight into why a candidate performs as it does.
We present TraceDSE, an agentic design-space exploration flow that uses system execution
traces as rich feedback. An LLM proposer generates candidate mappings and hardware
configurations for evaluation, while an LLM critic analyzes the resulting traces using
programmatic tools to identify bottlenecks and recommend targeted refinements. Across four AI
inference workloads deployed on an Intel Meteor Lake SoC, TraceDSE improves Pareto-front
hypervolume by up to 35% over NSGA-II and 68% over Bayesian optimization while requiring
approximately 6-9× fewer hardware evaluations.
BioGeetha Prasuna Yarramneni is a Ph.D. student in the Elmore Family School of Electrical and
Computer Engineering at Purdue University, advised by Prof. Anand Raghunathan. Her
research focuses on AI for electronic design automation, particularly agentic co-design of
hardware, software, and systems for efficient computing. She received her B.Tech. in Electrical
Engineering from IIT Madras and previously spent two years at NVIDIA designing and verifying
low-power features for GPU SoCs. |
|
4 |
S. Chandra, K. Roy |
2D-ThermAl: TCAD Model Informed Thermal Analysis of Circuits using GenAI
AbstractThermal analysis is increasingly critical in modern integrated circuits, where non-uniform power dissipation and high transistor densities can cause rapid temperature spikes and reliability concerns. Traditional methods such as FEM-based simulations offer high accuracy but computationally prohibitive for early-stage design, often requiring multiple iterative redesign cycles to resolve late-stage thermal failures. To address these challenges, we propose '2D-ThermAl', a physics-informed generative AI framework which effectively identifies heat sources and estimates full-chip transient and steady-state thermal distributions directly from input activity profiles. ThermAl employs a hybrid U-Net architecture enhanced with positional encoding and a Boltzmann regularizer to maintain physical fidelity. Our model is trained on an extensive dataset of heat dissipation maps for more than 200 circuit configurations, ranging from simple logic gates (e.g., inverters, NAND, XOR) to complex designs, generated via COMSOL and Cadence EDA flows. The dataset captures diverse activity patterns, and we note that material-dependent thermal properties may require targeted fine-tuning to ensure accuracy across different fabrication contexts. Experimental results demonstrate that ThermAl delivers precise temperature mappings for large circuits, with a root mean squared error (RMSE) of only 0.71°C and outperforms conventional FEM tools by running up to ~200x faster. We analyze performance across diverse layouts and workloads and discuss its applicability to large-scale EDA workflows. While thermal reliability assessments often extend beyond 85°C for post-layout signoff, our focus here is on early-stage hotspot detection and thermal pattern learning. To ensure generalization beyond the nominal operating range (25-55°C), we additionally performed cross-validation on an extended dataset spanning 25-95°C maintaining a high accuracy (<2.2% full-scale RMSE) even under elevated temperature conditions representative of peak power and stress scenarios.} Limitations such as 2D-only modeling and real-world validation are addressed with concrete future directions, including 3D extension, generalization across technology nodes, and transfer learning strategies.
BioSoumyadeep Chandra received his B.Tech degree in Electrical and Telecommunication Engineering from Jadavpur University, Kolkata, India in 2020. Currently, he is a 6th year Ph.D. student in the School of Electrical and Computer Engineering at Purdue University, West Lafayette, IN, under the mentorship of Prof. Kaushik Roy. His ongoing research focuses on areas such as procedural learning, scene understanding, surgical workflow analysis, and generative AI for hardware.
|
|
5 |
A. Joshi, K. Roy |
SHIRE: Enhancing Sample Efficiency using Human Intuition in REinforcement Learning
AbstractThe ability of neural networks to perform robotic perception and control tasks such as
depth and optical flow estimation, simultaneous localization and mapping (SLAM), and
automatic control has led to their widespread adoption in recent years. Deep
Reinforcement Learning (DeepRL) has been used extensively in these settings, as it
does not have the unsustainable training costs associated with supervised learning.
However, DeepRL suffers from poor sample efficiency, i.e., it requires a large number of
environmental interactions to converge to an acceptable solution. Modern RL algorithms
such as Deep Q Learning and Soft Actor-Critic attempt to remedy this shortcoming but
can not provide the explainability required in applications such as autonomous robotics.
Humans intuitively understand the long-time-horizon sequential tasks common in
robotics. Properly using such intuition can make RL policies more explainable while
enhancing their sample efficiency. In this work, we propose SHIRE, a novel framework
for encoding human intuition using Probabilistic Graphical Models (PGMs) and using it
in the Deep RL training pipeline to enhance sample efficiency. Our framework achieves
25-78% sample efficiency gains across the environments we evaluate at negligible
overhead cost. Additionally, by teaching RL agents the encoded elementary behavior,
SHIRE enhances policy explainability. A real-world demonstration further highlights the
efficacy of policies trained using our framework.
BioAmogh Joshi received his B.Tech. in Electronics Engineering from the University of
Mumbai in 2020. From 2020 to 2021, he was a Project Research Assistant at the Indian
Institute of Technology, Bombay where he worked on safety devices for Lighter-than-Air
systems, and an avionics and control suite for a medium-weight class UAV. Amogh
joined the Nanoelectronics Research Lab in the Spring of 2022, where he is currently
pursuing his PhD under the supervision of Prof. Kaushik Roy. His research interests
include autonomous and resilient robot learning, reinforcement learning, and
neuromorphic computing for autonomous drone applications. He is currently exploring
robot synthesis techniques using constrained co-learning approaches. |
|
6 |
D. Lee, K. Roy |
ViP-VLA: Vision-Internal Token Pruning for Efficient Vision-Language-Action Models
AbstractVision-Language-Action (VLA) models enable general robotic control by leveraging pretrained vision-language representations. However, their high inference cost makes closed-loop deployment expensive. At every decision step, hundreds of visual tokens are repeatedly processed through the vision encoder and the downstream model. Reducing this computation requires identifying which visual tokens are important for action generation and how early the less important tokens can be removed. We introduce ViP-VLA, an action-aware visual-token pruning framework that identifies visual tokens important for action generation and removes unimportant tokens inside the vision encoder. ViP-VLA learns token importance at an intermediate vision layer using the pretrained VLA's own action objective. During inference, the selector retains only the highest-scoring visual tokens after the intermediate vision layer. These tokens continue through the remaining vision-encoder and cross-modal layers, while discarded tokens receive no further processing, reducing computation without modifying the pretrained VLA. We evaluate ViP-VLA across multiple VLA architectures and robot manipulation benchmarks. On LIBERO with OpenVLA-OFT, ViP-VLA uses only 29.3% of the total FLOPs required by the original policy without token pruning and achieves a measured 1.85x inference speedup, while its average success rate is 96.8%, compared with 96.2% for the original policy.
BioDonghun Lee (James) received his B.S. in Electrical and Electronic Engineering from Yonsei University, Korea, in 2022. During his undergraduate years, he worked on Computer Vision. Currently, he is pursuing an M.S. degree in Electrical and Computer Engineering at Purdue University under the supervision of Prof. Kaushik Roy. His primary research interests lie in Neuro-Inspired Algorithms and Spiking Neural Networks. In his free time, he enjoys swimming as a hobby.
|
|
7 |
A. Mukherjee, K. Roy |
DRC-Aid: Design-Rule Violation Correction via Agentic Framework utilizing Inference-Time Large Language Models
AbstractResolving Design Rule Violations (DRVs) in layouts entails an iterative loop of geometric edits and
verification. We present DRC-Aid, a closed-loop agentic framework that automates local DRC repair by
formulating it as verification-in-the-loop search. To constrain the combinatorial geometric repair space, a
deterministic Rule Engine converts physical verification tool-reported violations into a bounded menu of
geometric edits. An off-the-shelf Large Language Model (LLM) evaluates local geometric context to select
edits from this menu, with budgeted depth-first search and backtracking. Immediate feedback from
verification tools enforces geometric compliance and guards against electrical-topology degradation,
while a global Memory Bank prevents cyclic re-exploration. Evaluated on FreePDK45 layouts containing
DRVs, DRC-Aid achieves DRC-clean, electrically-equivalent repairs in ∼92.5% of cases, while residual
cases yield partially repaired candidates. Under an identical search and verification infrastructure, LLM-
based selection outperforms random (54.4%) and deterministic-heuristic (83.3%) policies, with the gap
widening on cases more than six violations.
BioAnushka completed her bachelor's degree from TISL, Kolkata in Electronics and
Communication engineering in 2021. She worked at Cognizant Technology Solutions
from 2021 to 2022. She completed her masters degree from Indian Institute of
Engineering Science and Technology, Shibpur in VLSI Design in 2024, where she
primarily worked on hardware security and was awarded the silver medal for securing
the second position in the M.Tech program. Anushka is currently enrolled in Purdue's
PhD program under the direction of Prof Kaushik Roy. She currently works on
hardware-software co-design and automation in EDA workflows.
|
|
8 |
T. Chen, Y. Kim |
Kernel-Informed Predictive VRM Control for Efficient and Robust GPU Power Delivery
AbstractLarge language model inference on GPUs produces repeated load-current steps at kernel
launches. On sub-1 V rails at kilowatt-class power, these steps reach hundreds of amps,
and the static voltage guardband that absorbs them is increasingly expensive. On-die
responses such as adaptive clocking protect against a droop by stretching the clock, so
every event they catch costs performance. Moving the regulator setpoint ahead of the step
costs voltage instead and is the response that can shrink the guardband. But given the
regulator's limited slew rate, pre-positioning the setpoint takes microseconds of lead, far
more than any on-die signal can provide. Kernel launches offer that lead: decode is
structurally repetitive, so the step each launch will cause is predictable from the identity of
the kernel about to run. This work explores a predictor that exploits this recurrence to pre-
position the setpoint before the step arrives. We investigate what information in the launch
stream identifies a kernel's load step, how to learn step magnitudes online and decide
when a prediction is trustworthy enough to act on. We evaluate the design in AccelWattch
coupled to a PDN and regulator model, separating the ideal mechanism bound from
interface and learning cost.
|
|
9 |
S. Lee, Y. Kim |
Lightweight Early Workload Prediction for Energy-Efficient Audio Event Classification
AbstractAudio event classification (AEC) enables environmental awareness in various real-world
applications, but the deployment of always-on AEC on battery-powered low-power
microcontrollers (MCUs) remains challenging due to the tight energy budget and strict
real-time constraints. Although acoustic machine learning (ML) models are relatively
compact, continuous feature extraction and inference consume substantial energy in
streaming workloads. Early-exit inference combined with dynamic voltage and
frequency scaling (DVFS) has proven effective in energy-constrained real-time systems,
yet significant energy-saving opportunities remain under-explored for streaming
workloads. In this paper, we exploit the observation that adjacent audio segments from
the same event often produce similar feature tensors, while each event class exhibits a
distinct and relatively consistent exit-depth requirement. Together, these two signals
enable early workload prediction and adaptive inference. Based on these, we develop
an energy-efficient inference framework for AEC on MCUs. Through early, lightweight
workload prediction, the proposed framework enables proactive DVFS under real-time
constraints. Experimental results show that the proposed framework reduces average
inference energy by 16.17%-17.55% while preserving accuracy, enabling more
practical always-on AEC on commodity ultra-low-power platforms.
|
|
10 |
H. Vu, Y. Kim |
DynaBreath: Motion-Robust mmWave Respiration Monitoring
AbstractMillimeter-wave (mmWave) radar enables contactless, privacy-preserving respiration monitoring using commodity chipsets, making it attractive for clinical and home-care settings. However, current systems lose accuracy when the subject moves. Existing approaches reduce radar measurements to a one-dimensional phase trace, which collapses the spatial information needed to tell breathing apart from whole-body movement.
We propose DynaBreath that addresses this challenge by processing radar signals in three dimensions. The key insight is that breathing causes localized chest-wall deformation, while body sway translates the torso approximately uniformly. These two types of motion are geometrically distinguishable in a 3D point cloud. The system tracks the subject's torso pose to establish a body-centric reference frame, generates a spatially dense point cloud from the tracked torso region using interferometric elevation estimation and Doppler compensation, and estimates the respiration waveform through anchor-based spatial encoding with temporal modeling.
We evaluate DynaBreath under continuous periodic oscillation and free random motion. The system achieves mean absolute errors of 1.89 and 3.01 breaths per minute with waveform availabilities of 88.6% and 76.8%, representing a 3-4× error reduction and up to 6.4× availability improvement over four reimplemented baselines. Overall, the results demonstrate that 3D spatial processing on commodity radar hardware enables robust respiration monitoring under the tested continuous-motion conditions.
|
|
11 |
A. Mukherjee, K. Roy |
FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding
AbstractFine-grained visual understanding depends on local detail, yet visual encoders face a trade-off between costly full-image high-resolution processing and compact global encoding that can weaken such evidence. Inspired by human active vision, we separate where to look from what to encode. We focus on the latter and introduce FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry. We first isolate this encoding problem using oracle ground-truth crops in a controlled small-object regime. On ImageNet objects with a native maximum side of 96 pixels, FAVE improves Top-1 by 9.4 points over a fixed-resolution ViT on the same crop window with 12.7 times lower FLOPs. Increasing global resolution or backbone capacity does not recover the same operating point. We then integrate FAVE as a complementary local branch in FastVLM. Its local tokens are combined with FastVLM's global visual tokens, while the original global pathway and language model remain frozen. With at most 16 additional local tokens, FAVE improves TextVQA by 1.60 points and achieves a 3.3 times controlled TTFT speedup over SmolVLM2-2.2B. On GQA attribute questions, it improves FastVLM-1.5B by 1.31 points, extending the benefit beyond text while narrowing the gap to FastVLM-7B. Together, these results show that selectively allocating high-acuity local capacity provides an efficient complement to broader global representations and model scaling for fine-grained understanding of small objects, text, and attributes.
BioAmitangshu received his B. Tech in Applied Electronics and Instrumentation Engineering from West Bengal University of Technology in 2016. He graduated from Iowa State University, USA with a Master's in Computer Engineering in Fall 2019 where he worked as a Graduate Research Assistant at the Self-Aware Complex Systems Lab and DICE Lab (now moved to New York University) from Fall 2018 to Spring 2020. His primary research work revolved around topics in Deep Learning such as Domain Adaptation and Adversarial Machine Learning and he published his research works in conferences such as ICCV 2019 (main conference), CVPR and Neurips (Workshops). His M.S thesis project focused on GAN based Data Augmentation techniques for improving performances of Deep Networks in Perception based tasks for Autonomous Driving. In Summer 2020, he joined Purdue University's ECE department to pursue his PhD degree under the guidance of Prof. Kaushik Roy at the Center for Brain-Inspired Computing (C-BRIC). His primary research interests will focus on designing Neuro-inspired Algorithms for Explainable and Robust Learning in Autonomous Systems.
|
|
12 |
M. Chauhan, A. Bera |
HEROIC: Heterogeneous Evidential Reasoning for Open-Vocabulary Identification and Cross-Robot Collaboration
AbstractMulti-agent heterogeneous air-ground robot teams are attractive for open world
search, with applications for reconnaissance, urban search and rescue missions (USAR),
disaster response and recovery, and hazardous environments. These two platforms have
different failure modes: aerial robots cover ground quickly but cannot resolve small or occluded
targets from altitude, while ground robots can identify objects-of-interest, such as people or
hazardous objects, at close range but cover less area. Existing language-tasked teams either
have roles fixed prior, or have a language model assign them from hand-written capability tags,
so the team is unable to know when within a mission an asset is no longer useful. We present
HEROIC, a decentralized heterogeneous multi-agent open-vocabulary search coordination
framework that requires agents to communicate in natural language only. HEROIC's initial agent
role assignment is derived from sensor properties and a scale law to determine whether targets
can be detected with a high confidence. From the mission's natural language prompt alone, this
law assigns aerial flight altitudes and sweep spacing. When this calculated height falls below the
altitude for safe flight, aerial agents re-task themselves from searcher to aerial triage, escort,
and route guide for ground agents. Both robots maintain an evidential belief over the search
area (bearing rays for positive evidence, a log-odds posterior for negative evidence) and gate
any arrival on close-range verification. In full-stack experiments, HEROIC reaches the target
84\% of the time across all 6 scenes, compares to 35-53\% for vision-language frontier
baselines, frontier-based search, lawnmower, and random-walk running the same perception, all
while being 2-4x sooner to arrive at the target.
BioMihir Chauhan is pursuing B.S. degrees in computer
science and mathematics at Purdue University, West
Lafayette, IN, USA. He is an undergraduate researcher in the
IDEAS Lab, advised by Professor Aniket Bera, where he
works on reinforcement learning for autonomous and
collaborative robotics, including adversary-aware navigation,
heterogeneous air-ground robot teams, and structured value
learning for manipulation. His work has been published at
IEEE ICRA and in the IEEE Sensors Journal, where his
paper has been nominated for the Best Paper Award. He
was previously a Research Intern with the Department of
Electrical Engineering and Computer Sciences, University of
California at Berkeley. Mihir has twice interned as a Member of Technical Staff at Nutanix,
building ML-driven tools for hardware performance analysis and LLM inference on CPUs. In
2020, he co-founded the Virtual Robot Simulator, a robotics learning platform used by more than
60,000 students in 35 countries. His work on a portable water harvesting machine was
showcased in the Smithsonian's FUTURES exhibition in Washington D.C. Mihir?s recognitions
include FIRST Dean's List Finalist and FIRST Tech Challenge World Championship Finalist.
|
|
13 |
K. Cheng, A. Bera |
RTG: Reverse Trajectory Generation for Rigid-Body Manipulation Under Sparse Rewards
AbstractSparse-reward reinforcement learning remains a long-standing challenge for robotics, where extensive exploration is often required before meaningful reward signals can guide the propagation of state-value functions. Prior approaches typically rely on expert offline demonstrations or carefully designed curriculum learning strategies to improve exploration efficiency. Leveraging recent advances in differentiable rigid-body dynamics and trajectory optimization, we compute physically plausible trajectories that end at a goal configuration to provide dense learning signals. We propose Reverse Trajectory Generation (RTG), a replay-generation method that starts from sampled successful configurations and constructs predecessor transitions with a trajectory-optimization-based Reverse Rigid-Body Simulator (RRBS). RTG forward-replays generated transitions to reduce simulator inconsistency, uses beam search to obtain diverse reverse trajectories, and mixes the resulting replay with online data in off-policy RL. Across gathering, sorting, and articulated object manipulation tasks, RTG consistently improves learning efficiency and task performance over off-policy baselines and improved sampling strategies like reverse-curriculum and hindsight-relabeling baselines. We further quantify reverse-generation cost and demonstrate that the generated replay remains useful after forward replay in a 3D PyBullet environment. Results show that RTG is a practical way to exploit rigid-body structure for sparse-reward robot learning.
|
|
14 |
A. Kumar, A. Bera |
Fitting Through: Perceptive Posture Reconfiguration for Humanoid Locomotion in Height-Constrained Environments
AbstractHumanoid robots are mechanically capable of entering
spaces that are inaccessible to wheeled and quadrupedal platforms, but
doing so requires abandoning upright bipedal locomotion in favor of
contact modes that recruit the knees and hands. We present a single
perceptive policy that enables a Unitree G1 humanoid to traverse
vertically constrained environments by continuously reconfiguring
between crouched bipedal walking, knee walking, and quadrupedal
crawling, spanning base heights from 0.78 m when standing down to 0.3
m when crawling, with smooth transitions between contact modes rather
than discrete skill switching. Perceiving overhead geometry across these
postures is difficult because the sensor pose relative to the ground
changes drastically with contact mode, and because the underside of an
overhang is only observable from below it. We therefore introduce a
configuration-invariant clearance map, which the policy observes
directly: depth returns from multiple body-mounted cameras are
projected through the kinematic chain into a gravity-aligned,
terrain-anchored grid in which each cell stores the free vertical space
between the floor and the ceiling above it. Because clearance is a
difference of heights in a common frame, the representation is invariant
to the robot's own posture and sensor placement, so a single visual
observation is shared across all contact modes. The map carries an
explicit observability mask, distinguishing unobserved geometry from
free space, so that the policy commits to a posture from the geometry it
has resolved rather than treating occluded regions as passable.
Conditioned on this map, the policy selects its own posture, transition
timing, and forward speed from the geometry ahead, with no external
mode arbitration. We train one teacher per mode and per transition using
privileged information, distill them into a unified student, and fine-tune
with multi-critic PPO, which decouples the conflicting reward structures
of the different contact modes.
BioAyush is a Master's student in Robotics at Purdue University, where
he has been a member of the IDEAS Lab since 2025. His research
focuses on humanoid locomotion control and multi-robot collaboration.
Prior to his Master's, he spent 2.5 years in industry working on controls
and autonomy for legged robots.
|
|
15 |
Z. Li, A. Bera |
PoseShield: Neural Collision Fields for Human Self-Collision Resolution
AbstractSelf-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extreme articulations or stochastic motion synthesis, generated meshes frequently exhibit self-penetrations, leading to physically implausible results. We propose PoseShield, a neural collision constraint defined directly in SMPL pose space. We formulate collision correction as a constrained optimization problem and connect the learned constraint with the Eikonal equation. Enforcing Eikonal regularization ensures non-vanishing gradients near the collision boundary, improving numerical stability and robustness of the optimization process. Unlike prior methods that operate in the mesh space or rely on heuristic penalties, our approach operates directly in the low-dimensional space of human poses and is theoretically grounded. The same learned constraint extends to human motion sequences, providing a generator-agnostic post-hoc collision corrector without retraining the underlying motion model. Experiments on a newly constructed SMPL pose benchmark show that our method achieves a 95.8% success rate and outperforms state-of-the-art baselines.
|
|
16 |
X. Sun, A. Bera |
M3P-R1: Reinforcement Learning for Large Language Model Guided Multi-Modal Motion Planning via MIP Code Generation
AbstractMulti-Modal Motion Planning (M3P) requires joint reasoning over continuous motions and discrete mode transitions, making it difficult to solve efficiently. For instance, a bipedal robot may walk to a target location and then use its arms to grasp an object. This scenario captures both mode transitions and continuous dynamics, yielding feasible paths that neither purely discrete nor continuous planners can handle. While Mixed-Integer Programming (MIP) offers a principled framework, constructing tractable formulations for non-convex problems is typically manual and domain-specific, especially in the approximate, discretization-based MIP regime needed for non-convex robotic tasks.
We propose MP-R1, a reinforcement learning method that fine-tunes large language models (LLMs) to decompose MP tasks into MIP variables, constraints, and objectives. Instead of directly outputting answers—often prone to hallucination—the model generates executable Python code using MIP optimization libraries and constraint interfaces. This enables solver-backed execution for robust and verifiable solutions.
|
|
17 |
A. Vashisth, A. Bera |
CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation
AbstractRobot navigation typically assumes an obstacle-free path exists between start and goal. In real environments, however, clutter may block all routes. We introduce Lifelong Interactive Navigation, where a mobile robot with manipulation capabilities must move objects to forge paths and complete sequential object-placement tasks. Because environment modifications persist, decisions impact future navigability and task difficulty. We propose CoReLIN, an LLM-driven constraint-based reasoning framework with active perception. CoReLIN reasons over a structured scene graph to decide which objects to relocate, where to place them, and where to explore next. A standard motion planner executes reliable navigation and manipulation primitives. To evaluate long-horizon behavior, we introduce 2 new metrics - Long-term Efficiency Score (LES), a unified metric capturing success, execution efficiency, environment optimality, captured by Price of Clutter. In ProcTHOR-10k, CoReLIN outperforms best baseline by 16% under standard metrics and LES, and transfers to real-world hardware.
|
|
18 |
D. Sharma, K. Roy |
Input Loss Curvature as a Predictor of Sample Vulnerability to Hardware Noise
AbstractAnalog in-memory computing (AIMC) accelerators can deliver significant energy efficiency over conventional architectures, but their accuracy is limited by device- and circuit-level non-idealities. While prior work has characterized the effects of these non-idealities at model or layer granularity, their impact on individual samples remains largely unexplored. In this work, we show that an input?s vulnerability to such non-idealities can be strongly predicted by its loss curvature, a metric capturing how sharply the loss changes under small input perturbations. Across multiple models, datasets, and non-idealities, our experiments reveal a strong positive correlation between input loss curvature and analog non-ideality-induced failures, with failure rates increasing substantially for high-curvature samples. These findings uncover a previously overlooked, sample-level dimension of hardware robustness and suggest new opportunities for input-aware strategies for addressing non-idealities.
BioDeepika received her B. Tech in Electrical Engineering from Indian Institute of Technology Roorkee, India. During her undergraduate studies, she worked on digital circuit designing for different applications including wireless communication. She joined Purdue in Fall 2019 and currently she is pursuing her PhD under Prof. Kaushik Roy. Her research interests include developing hardware architectures for neuromorphic computing. She loves to read and travel.
|
|
| CHIPS for AI |
19 |
D. Sharma, K. Roy |
SpiDR: A 65nm 5 TOPS/W Digital CIM Accelerator with Reconfigurable Precision and Temporal Pipelining for Spiking Neural Networks
AbstractSpiking Neural Networks (SNNs) have the inherent ability to process highly sparse event data from dynamic vision sensors (DVS). However, many existing SNN accelerators support only fixed network architectures, limited data precision, and fail to efficiently handle membrane potential dynamics and varying input sparsity. This work presents SpiDR, a reconfigurable digital compute-in-memory (CIM) SNN accelerator that addresses these limitations through: (1) fused weight-Vmem macros for reduced data movement, (2) staggered layout and reconfigurable peripherals for variable-precision support, (3) zero-skipping to exploit unstructured input sparsity, and (4) asynchronous pipelining to maintain throughput despite variable computation times. Fabricated in TSMC's 65 nm CMOS technology, SpiDR achieves up to 5 TOPS/W at 95% input sparsity and supports diverse event-based workloads, including gesture recognition and optical flow estimation. SpiDR achieves 2× to 1000× higher area efficiency and up to 7.6× higher energy efficiency compared to prior designs.
BioDeepika received her B. Tech in Electrical Engineering from Indian Institute of Technology Roorkee, India. During her undergraduate studies, she worked on digital circuit designing for different applications including wireless communication. She joined Purdue in Fall 2019 and currently she is pursuing her PhD under Prof. Kaushik Roy. Her research interests include developing hardware architectures for neuromorphic computing. She loves to read and travel.
|
|
20 |
S. Bose, K. Roy |
HyFPCiM: A 65-nm 417-µW Error-Sensitivity-Aware FP8 Compute-in-Memory Macro
AbstractEdge AI applications such as wearables, health monitors, and IoT sensors require sub-mW inference, but existing floating-point computing-in-memory (FP-CiM) solutions consume 10-460 mW, far exceeding device budgets. Integer quantization (INT8/INT4) reduces bit-width but often demands high-precision scale-factor multiplications and retraining, which erode efficiency. This work presents HyFPCiM, a hybrid FP8 (1S-4E-3M) CiM architecture that enables sub-mW floating-point inference through error-sensitivity-aware FP partitioning (EAP). Exploiting the 8-15x higher error sensitivity of exponents relative to mantissas, EAP maps exponent processing to a digital CiM (DCiM) macro for exact, low-cost computation, while mantissa accumulation is performed using an analog CiM (ACiM) macro, where moderate circuit non-idealities are tolerable and massive parallelism can be exploited. The proposed DCiM employs dual-rail inverter-based sensing for exponent processing, while the ACiM performs pre-analog-to-digital conversion (ADC) charge-domain accumulation using switched-capacitor (SC) MAC circuits, avoiding the power- and area-intensive adder-tree-based accumulation employed in prior FP-CiM designs. By amortizing a single ADC conversion across eight MAC units and employing a duty-cycled, power-gated transimpedance amplifier (TIA)/ADC, HyFPCiM reduces ADC energy by 8x and ADC+TIA power by 75%, and total macro power by 31% compared to prior FP-CiM designs. Fabricated in 65 nm TSMC LP, HyFPCiM achieves 417 µW at 50 MHz (1 V), achieving 40 GFLOPS/W system efficiency and 18.2 TFLOPS/W/mm² area efficiency. After iso-28 nm normalization, HyFPCiM outperforms the best prior work by >2.75x. FP8 inference on ResNet-18 with CIFAR-10/100 and Tiny-ImageNet shows <0.55% accuracy loss versus software FP8. To the best of our knowledge, HyFPCiM achieves the lowest reported power among silicon-measured FP-CiM macros, demonstrating its suitability for energy-constrained edge AI inference.
BioSouhardya received his Bachelor of Technology (Honours) degree from the Indian Institute of Technology Kharagpur, India, in 2025 with a major in Electrical Engineering and a minor in Electronics and Electrical Communication Engineering. He was awarded the Prof. K Venkataratnam Memorial Award for securing the highest GPA in core electrical courses and their associated laboratories. During his undergraduate years, he has worked on mixed-signal circuits, where he contributed to a tape-out of a RHBD SERDES IC, and compute-in-memory architectures. He is currently pursuing his Ph.D. Degree in Electrical and Computer Engineering at Purdue University under the supervision of Professor Kaushik Roy. His current research interests include hardware architecture design for in-memory computing, algorithm-hardware co-design, AI accelerators, and probabilistic computing. Besides research, he enjoys listening to folk music and Rabindra sangeet, gardening, and painting.
|
|
21 |
A. Holla, K. Roy |
LIMO: Low-Power In-Memory-Annealer and Matrix-Multiplication Primitive for Edge Computing
AbstractCombinatorial optimization (CO), exemplified by the NP-hard Traveling Salesman Problem (TSP), is often intractable on von Neumann architectures due to the 'memory wall' and exponential search space growth with problem scale. While compute-in-memory (CiM) architectures using simulated annealing can mitigate these bottlenecks, solution quality degrades as instances scale. To address this, we present LIMO, a programmable mixed-signal macro implementing an in-memory annealing algorithm with reduced search-space complexity. LIMO utilizes stochastic switching of spin-transfer-torque magnetic-tunnel-junctions (STT-MTJs) to escape local minima. For large-scale problems, a refinement-based divide-and-conquer algorithm enables parallel optimization across a spatial architecture. Consequently, LIMO achieves superior quality and faster time-to-solution for instances up to 85,900 cities compared to prior annealers. Furthermore, with annealing functionality implemented as modular peripherals, the compute core can be repurposed for diverse workloads, such as neural network inference. We demonstrate image classification and face detection with software-equivalent accuracy, surpassing baselines in latency and energy efficiency.
BioAmod received his bachelor's degree in Electrical Engineering from the Indian Institute of Technology, New Delhi in 2024. During his undergraduate years, he worked on analog compute-in-memory using spin-based devices and spiking neural network (SNN) architectures for edge-AI applications. His work on crossbar-array-based SNNs for cognitive memory modeling earned him the award for Best Bachelor?s Thesis. His research interests include neuromorphic computing, hardware-algorithm co-design, and circuit design for supporting post-CMOS architectures. Apart from research, he enjoys nature, strength training, endurance running, traveling, gardening, and cooking.
|
|
22 |
R. Koduru, K. Roy |
Integrating BEOL IGZO eDRAM as On-die Scratchpad for GPUs via System Technology Co-optimization
AbstractBack-end-of-line (BEOL) embedded DRAM (eDRAM) alleviates off-chip memory
bottlenecks, limiting front-end-of-line (FEOL) overhead to peripherals. BEOL-compatible
indium-gallium-zinc-oxide (IGZO) transistors offer ultra-low leakage, enabling capacitor-less
high-retention gain-cells (4.3s at 85° C). We deploy these IGZO eDRAM arrays as extra on-die
scratchpads for GPUs, bridging L2 cache and high-bandwidth memory (HBM). Placing them in
intermediate BEOL layers (M6-M10 of in-house 7nm node), directly above their area-matched
peripherals (in M1-M5), we preserve M11 and beyond for global routing. Under iso-GPU-die
constraint, these peripherals displace baseline compute cores (SMs) and L2 cache, forcing a
memory-compute trade-off. We develop a system-technology co-optimization (STCO)
framework that co-designs this reallocation with realizable eDRAM capacity, bandwidth and
evaluates the resulting scratchpad in an A100-like GPU. Utilizing 2T0C and 3T0C eDRAM
macros, we achieve effective densities of 5.30 and 4.05 MB/mm2, yielding up to 800 and 600 MB
scratchpads (15-20× baseline L2). Across LLM workloads, this integration yields up to 35%
energy and 15% latency savings versus unmodified baseline, with eDRAM refresh energy
consuming <0.001% of total system energy. We also identify optimal SM/L2 reductions and IGZO
technology targets.
BioRevanth Koduru is a postdoctoral researcher holding a joint appointment between imec
USA and Purdue University under the guidance of Prof. Kaushik Roy. He earned his
Ph.D. in Electrical and Computer Engineering from Purdue University in December
2025. His doctoral research concentrated on physics-based multi-grain phase-field
models of hafnium-zirconium-oxide ferroelectric devices, providing critical insights into
their fundamental physical mechanisms and device properties. His current research
focuses on Design-Technology Co-Optimization (DTCO) and System-Technology Co-
Optimization (STCO) frameworks. Specifically, he investigates the hardware
acceleration of artificial intelligence workloads and large language models (LLMs). His
approach focuses on optimizing emerging memory technologies and their integration
into advanced computing architectures via 3D bonding and monolithic integration. |
|
23 |
A. Kosta, K. Roy |
SG-CPG: Severity-Gated Central Pattern Generators for Adaptive Quadruped Locomotion under Continuous Actuator Degradation
AbstractLegged robots are poised to carry increasingly capable AI into homes, workplaces and open terrain, where they must keep working for long periods. Over such deployments their actuators can weaken, and an overheating or underpowered motor can no longer deliver the torque the gait needs. Existing fault-tolerant controllers majorly model faults such as locked joints or complete power loss, but not a lowered torque ceiling that stays invisible to proprioception until the gait demands the capped torque. Animals, in contrast, sense an injured or weakened limb and readily adapt their walking pattern to accommodate it. Spinal central pattern generators (CPGs), which produce the basic locomotor rhythm, contribute to this adaptation. Inspired by CPGs, we propose SG-CPG, a framework for estimating and adapting to continuous actuator degradation under both torque-ceiling and gain-scaling faults. SG-CPG freezes a healthy CPG policy and adds a residual policy that acts only under a fault and modulates the actions based on fault-severity. We also perform a controlled study of severity estimators showing that the fault mechanism decides which estimator to use. In simulation, SG-CPG keeps a trot at every tested severity under both fault mechanisms. With a joint reduced to 5% of its strength, it survives ≥97% of straight-line episodes at a forward velocity tracking error of ≤9% of the commanded 1.0 m/s. For omnidirectional commands, it still survives 83-94% of episodes with an error of 9-17%. On a real Unitree Go2 with an emulated calf torque ceiling lowered by up to 93%, SG-CPG survives 28 of 29 forward and turning runs (97%). Across these runs, commanded up to 1.0 m/s and 1.0 rad/s, the forward-speed and yaw-rate errors average 26% and 31% of the command, respectively.
BioAdarsh Kumar Kosta is a Ph.D. candidate in Electrical and Computer Engineering at Purdue University, working with Prof. Kaushik Roy in the Nanoelectronics Research Laboratory. His research explores bio-inspired computation for efficient perception, decision-making, and control, spanning event-based vision, spiking neural networks, and reinforcement learning, with recent work extending to fault-tolerant algorithms for legged robot locomotion. He has published in venues including ICRA, IROS, and Frontiers in Science, and previously interned at Samsung AI Center, where he developed a patented acoustic-based tactile sensing system for robots. Adarsh is broadly interested in building algorithms that draw on biology's efficiency ? sparse, adaptive computation that performs well under real-world constraints.
|
24 |
M. Mukherjee, K. Roy |
MIRAGE:MRAM-Based Near ADC-Less Compute-In-Memory Macro for Deep Learning Acceleration
AbstractNon-volatile memory (NVM) based Compute-in-Memory (CiM) architectures have emerged as a promising compute primitive for accelerating deep neural networks (DNNs) by performing in-situ matrix?vector multiplications (MVMs). Among various NVMs, STT-MRAM (Spin Transfer Torque based Magnetoresistive Random Access Memory) shows potential due to its high endurance, low energy consumption and high density. However, existing STT-MRAM CiM designs typically rely on multi-bit analog-to-digital converters (ADCs) at the peripherals to digitize accumulated bit-line currents. While enabling high-precision computation, ADCs add substantial energy, latency, and area overheads. To alleviate such problems, we propose a system-technology co-design approach to a Near ADC-Less CiM design with ternary partial-sums called MIRAGE. The accuracy is maintained by considering hardware level partial sum quantization in the training loop. Specifically, we develop an STT-MRAM based CiM macro which features differential bitcells and an adaptive threshold sensing that is amenable to the requirements posed by ternary partial-sum quantization. We do a thorough energy, area, latency, and sense margin analysis along with robust benchmarking against conventional 1T-1MTJ (1 transistor-1 Magnetoresistive Tunnel Junction) based MRAM CiM. The proposed CiM macro occupies ∼ 20% less area, consumes 1.8× less MVM energy and shows 5× better latency with improved distinguishability compared to 1T-1MTJ CiM macro while achieving better accuracy.
BioMainakh Mukherjee received his B.Tech. from the West Bengal University of Technology in 2021 and completed his M.Tech. in VLSI Design at the Indian Institute of Engineering Science and Technology (IIEST), Shibpur in 2024. During his master's, he worked on modeling non-linearities and designing mixed-signal circuits, including data converters such as SAR and flash ADCs. He joined Purdue University in Fall 2024 to pursue a Ph.D. under Prof. Kaushik Roy. His research interests lie in device?circuit co-design, mainly circuit designing using emerging memory technologies for compute-in-memory at advanced technology nodes.
|
|
25 |
A. Pranta, K. Roy |
To CiM or Not to CiM? A System-Level Study of Digital and Analog Compute-in-Memory
AbstractThe widening gap between compute throughput and memory bandwidth has made data movement a major
performance and energy bottleneck in modern AI accelerators. Compute-in-memory (CiM) addresses this by
performing computation within or near the memory arrays that store the model weights, reducing weight
movement and exploiting the arrays' inherent parallelism. However, these macro-level gains may not translate
into system-level benefits once peripheral overheads, compute utilization, buffering, and off-chip memory traffic
are included. We examine when CiM is advantageous over conventional compute, and under what conditions
that advantage holds by comparing digital CiM (DCiM), ADC-based analog CiM (ACiM), near-ADC-less ACiM, and
a conventional digital MAC array. All compute primitives are implemented in 22-nm GlobalFoundries FDX
technology, and the CiM primitives are based on SRAM arrays. At the macro level, CiM provides 12.4×-67.2×
higher area efficiency and up to 3.2× higher energy efficiency than the conventional digital MAC baseline.
However, full-precision ACiM remains energy-inefficient due to ADC overhead. Reducing ADC precision from
6-bits to 4-bits improves ACiM's area and energy efficiency by 1.8× and 1.57×, respectively, while near-ADC-less
ACiM achieves the highest area efficiency (27.43 TOPS/W/mm2) and energy efficiency (4.83 TOPS/W).
Integrating these primitives into a common accelerator framework under iso-area constraints shows that CiM's
system-level benefit is governed by the on-chip compute to off-chip memory traffic ratio. Near-ADC-less ACiM
achieves up to 3× speedup on ResNet-18. DCiM and ADC-based ACiM variants achieve 54%-105% speedup
over the digital MAC baseline on compute-intensive, prefill-heavy LLM workloads, shrinking to only 1%-2% on
the most decode-dominated workloads. As off-chip DRAM traffic dominates, the compute-primitive choice has
little effect on system performance. Overall, CiM is most effective when its area efficiency translates into
balanced system resources with sufficient array utilization and data reuse.
BioAyan earned his bachelor's degree in Electrical and Electronic Engineering from
Bangladesh University of Engineering and Technology (BUET). He began his PhD in
Electrical and Computer Engineering at Purdue University in Fall 2024 under the
supervision of Professor Kaushik Roy. His research focuses on memory-centric hardware
and architectures for AI and neuromorphic computing.
|
|
26 |
A. Das, V. Raghunathan |
COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems
AbstractMultimodal AI (MMAI) systems are becoming integral to safety-critical edge applications,
from autonomous driving and robotics to smart healthcare and surveillance, where real-time
perception must operate under stringent energy and compute constraints. These systems
comprise distinct sensing units, such as cameras, microphones, and inertial sensors, each
forming a separate subsystem with unique capabilities and constraints. As environments
evolve, modality relevance shifts and redundancy emerges, requiring dynamic adaptation to
maintain quality of service while extending system lifetime. Existing methods optimize
individual subsystems or static inference paths, lacking responsiveness to contextual changes;
coordinating adaptation across these heterogeneous subsystems remains essential yet largely
unexplored. We introduce COSMOS, a runtime framework for distributed multimodal edge
nodes. Its central contribution is a closed-loop control structure that uses predicted semantic
context and measured temporal stability to jointly govern modality activation, inference
frequency, and subsystem-level approximation. COSMOS instantiates this control loop
through three mechanisms: Scene/State-aware Modality Selection (SMS), Temporal Locality-
driven Decision and Regulation (TLDR), and modality-aware Design Space Exploration
(DSE). The DSE heuristic coordinates compute, audio preprocessing, sensing, memory, and
communication knobs under a unified energy-quality objective. Validated on an FPGA-MCU
prototype (Intel Agilex-5 primary + ESP32 auxiliary) for Audio-Visual Event Localization,
COSMOS extends system lifetime by 7.73x with <5% accuracy loss, achieving up to 1.77x
primary and 3.01x auxiliary energy savings across six transformer backbones.
BioArghadip Das earned his B.E. degree in Electronics and Telecommunication Engineering
from Jadavpur University, Kolkata, India, in 2020. Currently, he is pursuing his Ph.D. at the
School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN,
USA, under the guidance of Prof. Vijay Raghunathan. During his undergraduate studies, his research centered on designing low-power hardware accelerators. At Purdue, Mr. Das's
primary research interests lie in deep learning for resource-constrained embedded systems.
His current work focuses on creating innovative approximate systems to enable collaborative
distributed inference in tiny, resource-limited computing systems and IoT devices. His
approach emphasizes energy efficiency, adaptability, and enhanced performance through
strategic approximations across various levels of system design. Mr. Das has been recognized
as a Distinguished Design Automation Conference Young Fellow twice, in 2022 and 2024.
He is also the recipient of the University Gold Medal from Jadavpur University and the
prestigious ECE STAR Fellowship from Purdue University.
|
|
27 |
H. Kim, V. Raghunathan |
TESSERA: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs
AbstractAI model architectures are diversifying rapidly. Although dense matrix multiplication still
underlies today's CNNs and transformers, emerging architectures (state-space models, long
convolutions via the fast Fourier transform (FFT), Kolmogorov-Arnold networks, and spiking
networks) are not multiply-accumulate (MAC) dominated; they spend much of their computation
on vector and non-MAC primitives that homogeneous, MAC-centric neural processing units
(NPUs) serve poorly. This has motivated heterogeneous NPUs (HPUs) built from non-identical
tiles. Prior heterogeneous designs, however, vary only one or two coarse knobs (typically MAC
precision or array size) and are evaluated on narrow workloads, and no existing framework
supports HPU design at a fine granularity, where tiles differ across many architectural
dimensions at once.
We present TESSERA, an analytical simulator and design-space-exploration (DSE) framework
for HPU microarchitecture design. TESSERA searches the joint space of tile-level heterogeneity:
beyond array size and precision, it varies the knobs along which tiles can differ, including tile-
type composition (large Big, small Little, and non-MAC Special-Function tiles), dataflow,
sparsity mode, MAC engine type, and special-function units for non-MAC operators (FFT,
spiking-integrate, polynomial). Unlike prior simulators, which model a single homogeneous tile
type, TESSERA models non-MAC tiles with their own energy, area, and timing models and
maps operators across a mix of tiles with a heterogeneity-aware compiler. A multi-seed pipeline
that pairs a stratified sweep with genetic-algorithm refinement returns Pareto-optimal designs,
with cost models calibrated to a 7 nm node and cross-validated against NVIDIA's Deep Learning
Accelerator (NVDLA). Across a 20-workload suite, the best general-purpose HPU found by
TESSERA (200 mm 2 Big+Little+Special-Function) achieves +46.91% mean iso-area energy
savings over the best iso-area homogeneous baseline.
BioHoseok Kim is a Ph.D. student in Electrical and Computer Engineering at Purdue University,
advised by Professor Vijay Raghunathan. His research focuses on energy-efficient edge
computing and NPU architecture. He previously earned his M.S. at Korea University, where his
work on DNN inference efficiency won the Best Paper Award at ISLPED 2024.
|
|
28 |
I. Ahmed, S. Gupta |
1.58-b FeFET-Based Ternary Neural Networks: Achieving Robust Compute-In-Memory With Weight-Input Transformations
AbstractCompute-in-memory (CiM)-enabled Ternary weight neural networks (TWNs), with weights ∈ {-1, 0,
1}, are attractive for edge-AI as they offer high energy-efficiency while maintaining acceptable
inference accuracies. To that end, multi-weight Ferroelectric transistors (FeFETs) can enable highly
scalable TWN-CiM acceleration. However, standard 1T-FeFET-based CiM is vulnerable to
hardware non-idealities at deeply scaled technology nodes, necessitating custom 2T-differential
designs to improve robustness at the cost of area/energy. In this work, we explore three FeFET-
based TWN-CiM solutions: (i) 1T (1.58-bit 3-level storage), (ii) 2T-differential (1-bit storage), (iii) 2T-
pull-up-pull-down (1-bit storage) designs; and propose static-weight transformation (WT) and static-
weight-dynamic-input transformation (WIT) for increasing the computational robustness of 1T-
FeFET. Using our rigorous phase-field FeFET models and non-ideality-aware inference simulator,
we conduct comparative analyses of these designs, showing that WT and WIT significantly recover
inference accuracy for 1T-FeFET (up to 73.61% for ResNet18 TWN-CiM on CIFAR100), making it
competitive against 2T-differential-FeFET (76.4%), while offering 1.98x and 1.91x reductions in
macro-area and macro-energy, respectively.
BioImtiaz Ahmed is a Ph.D. candidate in Electrical and Computer Engineering
at Purdue University, conducting research at the Integrated Circuits &
Devices Lab (ICDL). His research revolves around efficient AI hardware
design, with a keen focus on emerging memory technologies, in-memory
computing, and robust deep neural network accelerator design. His work
aims to develop energy-efficient, high-performance, and scalable hardware
for next-generation of AI accelerators. Imtiaz is also an SRC Research
Scholar affiliated with CoCoSys: Center for the Co-Design of Cognitive
Systems, under Joint University Microelectronics Program 2.0 (JUMP2.0)- a
Semiconductor Research Corporation (SRC) program sponsored by the Defense Advanced
Research Projects Agency (DARPA). He completed his B.Sc. in Electrical and Electronic
Engineering from Bangladesh University of Engineering and Technology (BUET), Dhaka,
Bangladesh in 2019, with a concentration in Electronics. Prior to joining Purdue University, Imtiaz
worked as a Lecturer at BRAC University (Bangladesh) for 2.5 years, where he coordinated
multiple undergraduate courses and contributed to curriculum redesign for integrating Outcome-
Based Education (OBE) framework, to align courseworks with industry and research
advancements. |
|
29 |
S. Bhattacharjee, S. Gupta |
WEBCRAFT: Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework
AbstractThe deployment of deep neural networks (DNNs) on compute-in-memory (CiM)
accelerators offers significant energy savings and speed-up by reducing data movement during
inference. However, the reliability of CiM-based systems is challenged by stuck-at-faults (SAFs)
in memory cells, which corrupt stored weights and lead to accuracy degradation. While closest
value mapping (CVM) has been shown to partially mitigate these effects for multibit DNNs
deployed on bit-sliced crossbars, its fault tolerance is often insufficient under high SAF rates or
for complex tasks. In this work, we propose two training-free weight transformation techniques,
sign-flip and bit-flip, that enhance SAF tolerance in multi-bit DNNs deployed on bit-sliced
crossbar arrays. Sign-flip operates at the weight-column level by selecting between a weight
and its negation, whereas bit-flip provides finer granularity by selectively inverting individual bit
slices. Both methods expand the search space for fault-aware mappings, operate synergistically
with CVM, and require no retraining or additional memory. To enable scalability, we introduce a
look-uptable (LUT)-based framework that accelerates the computation of optimal
transformations and supports rapid evaluation across models and fault rates. Extensive
experiments on ResNet-18, ResNet-50, and ViT models with CIFAR-100 and ImageNet
demonstrate that the proposed techniques recover most of the accuracy lost under SAF
injection, often restoring performance to within 1-2% of ideal baselines. Hardware analysis
shows that these methods incur negligible overhead, with sign-flip leading to sub-2% energy,
latency, and area cost, and bit-flip providing higher fault resilience with modest overheads.
These results establish sign-flip and bit-flip as practical and scalable SAFmitigation strategies
for CiM-based DNN accelerators.
BioSaptashwa Bhattacharjee received the B.Tech. degree in
Electronics and Electrical Communication Engineering from the Indian
Institute of Technology Kharagpur, India, in 2025, and is currently
pursuing the Ph.D. degree in Electrical and Computer Engineering at
Purdue University. His research interests include in-memory
computing, AI accelerators, and hardware-aware machine learning.
|
|
30 |
D. Kim, S. Gupta |
SCION: A Comprehensive Simulation Framework for Charge-Based In-Memory Computing for Rapid Evaluation of Hardware Non-Idealities and DNN Accuracy
AbstractCharge-based in-memory computing (IMC) has shown great potential in achieving higher
computational robustness compared to current-based IMC. However, it suffers from its own
non-idealities such as parasitic capacitive coupling. Accurately evaluating these effects requires
time-intensive SPICE simulations, making it challenging to conduct cross-layer exploration. To
overcome these limitations, we propose SCION, a PyTorch-based framework that rigorously
models hardware non-idealities in charge-based IMC and integrates them directly with DNN
inference for rapid accuracy evaluation. We show that SCION predicts the IMC output with more
than 99% accuracy with respect to SPICE while offering four orders of magnitude speedup. We
demonstrate the capability of our framework by analyzing an SRAM-based charge-IMC
accelerator deploying ResNet-50 and ViT-small DNNs with CIFAR-100 dataset. We show how
inter-column capacitive coupling leads to data-dependent non-idealities, severely impairing the
inference accuracy. We also explore techniques to mitigate non-idealities using SCION. To that
end, we propose alternate column enablement (ACE) to eliminate inter-column coupling. Our
results show that compared to the baseline design and another non-ideality mitigation approach based on ground shielding, ACE achieves significant improvement in sense margin and
near-software accuracy under nominal conditions. Further, compared to ground shielding, ACE
exhibits a higher tolerance to analog-to-digital converter (ADC) noise and superior voltage
scalability.
BioDoug Hyun Kim is currently pursuing the Ph.D. degree in electrical and computer engineering at
Purdue University, West Lafayette, IN, USA, where he is a Graduate Research Assistant with
the Integrated Circuits and Devices Laboratory under the supervision of Prof. Sumeet Kumar
Gupta. His research interests include compute-in-memory architectures, AI hardware
accelerators, emerging memory technologies, mixed-signal and charge-domain computing, and
hardware-software co-design for efficient artificial intelligence systems. His current research
focuses on cross-layer modeling and optimization of compute-in-memory systems, including
circuit nonidealities, and device-circuit-algorithm co-design for deep neural-network inference.
|
|
31 |
E. Berscheid, A. Raghunathan |
Optimizing Data Placement for Near-Memory Acceleration of Recommendation Systems
AbstractRecommendation systems are widely-used machine learning models that process user
data to recommend content. These models heavily rely on looking up and summing
embedding vectors from large tables, which is a memory-bound operation that
bottlenecks performance. Near-memory processing (NMP) accelerates embedding look-
ups by employing processing units within the DRAM system, which compute multiple
vector operations in parallel and reduce data volume before transferring the results to
the host. We observe that data placement has a profound impact on NMP efficiency, an
effect that is not fully addressed by prior work. To that end, we make the following
contributions: First, we propose a memory-efficient statistical method for profiling
embedding table look-ups to find commonly co-occurring sets of vectors. We recognize
that vector replication holds potential for enhancing vector co-locality, and not only load
balance as explored in prior works. Accordingly, our second contribution is a replicate
resolution policy that enables a principled and tunable trade-off between load balance
and vector co-locality to maximize performance. Next, we redistribute non-replicated
embedding vectors in physical DRAM to collocate (or separate) vectors that are likely
(or unlikely) to belong to the same access set. DRAM simulations demonstrate up to
3.41x speedup and 3.01x energy savings over prior methods.
BioElijah Berscheid is a PhD student in ECE, advised by Professor Anand
Raghunathan. His research interests relate to efficient data center scale AI, including systems,
architecture, and algorithms. His previous project aimed to optimize near-DRAM
computing performance for memory-bound recommendation system workloads (under
review at TCASAI). His current work explores performance gaps due to dynamic
irregularities in sparse decode attention for long-context workloads to improve
performance on chiplet-based accelerators. Elijah has interned at the University of
Minnesota (control systems and data processing), Western Digital (HDD servo
firmware), Broadcom (3nm ROM verification), and NVIDIA (efficient LLM inference on
future MCM-GPUs). |
|
32 |
D. Kim, S. Hsu, V. Jain |
Hardware Efficient VLA Acceleration using Vector Quantization for Physical AI
AbstractVision-Language-Action (VLA) models have emerged as a
promising framework for general-purpose robotic control, leveraging the rich
representations and generalization capabilities of foundation models such as
Vision-Language Models (VLMs). However, their large parameter scale and
computational demands pose significant challenges for deployment on
resource-constrained robotic platforms, where low-latency inference is critical
for real-time interaction. In this work, we investigate several low-bit
quantization methods for accelerating VLA inference, focusing on both KV-
cache and weight compression. At low-bit precision, quantization quality
degrades significantly due to outliers concentrated in certain activation
channels. Rotation-based transformations redistribute these outliers across
multiple dimensions, reducing their concentration and improving quantization
quality. First, we explore TurboQuant for KV-cache quantization and evaluate
MSE-optimal and inner-product-optimal quantization schemes using the π0
model on VLABench. We further investigate QuaRot for weight quantization,
which applies Hadamard-based rotations to input activations to enable INT4
GEMM in major Transformer components, including gated FFNs and QKVO
projection layers. With QuaRot-based INT4 GEMM, we achieve up to a 1.45×
inference speedup. Finally, we examine Leech Lattice Vector Quantization
(LLVQ), which leverages the sphere-packing properties of the 24-dimensional
Leech lattice and uses its structured lattice points as a predefined codebook,
eliminating the need for codebook retraining while achieving up to 8× weight
compression on Gemma-2B. Together, these approaches explore
complementary opportunities to reduce the memory and computational
overhead of VLA models when deployed in real-time robotic systems.
BioShun-Hsiu (William) Hsu received his B.S. dual degree
in Aeronautics and Astronautics and Electrical
Engineering, followed by his M.S. degree in Electrical
Engineering, from National Cheng Kung University
(NCKU), Tainan, Taiwan. He is currently a Ph.D.
student in the Elmore Family School of Electrical and
Computer Engineering at Purdue University, West
Lafayette, IN, advised by Prof. Vikram Jain. His
research focuses on Physical AI, digital integrated
circuit (IC) design, and compute-in-memory (CIM) AI
accelerators.
Dongyoon Kim received his B.S. degree in Electronic
and Electrical Engineering from Sungkyunkwan
University (SKKU), Suwon, South Korea in 2026. He is
currently a Ph.D. student in the Elmore Family School
of Electrical and Computer Engineering at Purdue
University, West Lafayette, IN, advised by Prof. Vikram
Jain. His research focuses on Domain Specific
Accelerators and Embodied AI.
|
|
33 |
H. Cho, Y. Kim |
CO-MAC: A Center-Out Ordered Stochastic MAC for Low-Latency Inference
AbstractStochastic computing offers efficient approximate arithmetic that aligns well with error-tolerant machine learning workloads, but its deployment is limited by long bitstream latency in stochastic multiply-accumulate (MAC) units. Prior work reduces MAC latency through deterministic bitstream generation and differential accumu lation, but these methods do not fully exploit the statistical property of convolution weights. This work presents a novel stochastic MAC architecture named CO-MAC, which employs center-out weight ordering and an enhanced convolution engine design to reduce effective computation cycles while maintaining high accuracy. The method sorts weights by magnitude, reuses the incremental differ ences in magnitudes, and applies sign handling after accumulation. This shortens counter activity, maintains accuracy with long ef fective bitstreams, and simplifies the MAC hardware by avoiding bidirectional counters. Across convolutional neural network work loads, CO-MAC decreases MAC latency by up to 54.8% compared to prior stochastic MAC architectures, while preserving accuracy and hardware simplicity.
|
|
34 |
D. Kim, Y. Kim |
Rethinking DRAM Protection for Hybrid-Bonded 3D Stacks
AbstractMemory reliability now limits large AI system availability. Hyperscale
datacenter operators report memory as the largest identified cause of unplanned
interruptions, and bandwidth demand is pushing that memory vertically. Stacking the
array on the logic die it feeds runs it hotter, leaves each channel with far fewer banks,
and returns a much larger block per access. That larger block changes what the errors
look like. Bits that once failed together now land apart, beyond the reach of the symbol
codes deployed today. Covering that difference in hardware takes area, and every
channel pays it again.
We propose and investigate a protection scheme reorganized around a single
controller-side code. The same code corrects the errors that arrive and folds out the
cells that stay bad. Freed of standing faults, most accesses need only shallow
correction. The rest is handled as it arises. The trade-off moves from area to time, and
the worst case is paid only when it occurs.
BioDoh Yon Kim received his B.S. degree in Electrical and Computer Engineering
from Sungkyunkwan University, Republic of Korea, in 2026. He is currently pursuing the
Ph.D. degree with the Elmore Family School of Electrical and Computer Engineering,
Purdue University, West Lafayette, under the mentorship of Prof. Younghyun Kim. His
research interests include memory reliability, error correction and runtime repair for 3D-
stacked DRAM, and memory systems for AI accelerators. |
|
35 |
M. Chen, H. Li |
Cross-Layer Heterogeneous Memory Systems for Scalable AI and Optimization Workloads
AbstractEmerging AI and optimization workloads place increasing pressure on the memory hierarchy through growing capacity, bandwidth, and data-
movement demands. At the same time, emerging memory technologies
provide new opportunities to redesign how data is stored, moved, and
processed across the hierarchy. This work explores cross-layer
heterogeneous memory architectures that match the characteristics of
different memory technologies with workload requirements. We
investigate eDRAM?RRAM heterogeneous memories for hyperdimensional
computing and zeroth-order fine-tuning, stochastic MTJ compute-in-
memory for Ising optimization, and HBM?HBF systems for large
language model inference. Across these systems, we study how device
and circuit characteristics translate into architectural tradeoffs in
data placement, computation, bandwidth utilization, and energy
efficiency. Together, these efforts establish a cross-layer
methodology for building scalable memory hierarchies for emerging AI
and optimization workloads.
BioMufeng Chen received the B.S. degree in optics and electronic information
engineering from Huazhong University of Science and Technology, Wuhan, China, in
2021. From 2021 to 2024, he was a Research Assistant with Zhejiang University,
Hangzhou, China, where he worked on hardware/software co-design and system
integration for emerging computing systems. He is currently pursuing the Ph.D.
degree with the Elmore Family School of Electrical and Computer Engineering,
Purdue University, West Lafayette, USA.
His research interests include heterogeneous memory systems, memory-centric
architectures, hardware/software co-design, and system-level integration for data-
intensive computing.
|
|
36 |
Z. Wang, H. Li |
'CMOS+X' Monolithic 3D-ICs for Heterogeneous Cognitive Computing Systems
AbstractTwo-dimensional CMOS scaling alone can no longer simultaneously deliver dense on-chip memory and new
device-level functionality demanded by emerging cognitive computing workloads. The "CMOS+X" approach
addresses this by keeping a single foundry Si CMOS platform and sequentially integrating heterogeneous
'X' device tiers in the back-end-of-line (BEOL), forming monolithic 3D integrated circuits (M3D-ICs). This
work explores two complementary CMOS+X tracks. First, we develop CMOS+OSFET hybrid structured-ASIC
fabrics, where BEOL oxide-semiconductor transistors enable capacitor-less eDRAM and reusable logic/memory
tiles on top of a fixed pre-fabricated CMOS base, guided by a device?circuit?system co-design framework that
reduces NRE cost and accelerates application-specific deployment. Second, we design a CMOS+3D-MTJ
probabilistic computing prototype, where BEOL stochastic magnetic tunnel junctions serve as physics-based
true-random entropy sources for FEOL digital compute cores, realizing Boltzmann-machine-based invertible
logic that unifies forward arithmetic and reverse inference within one hardware fabric. Across both tracks,
we study how BEOL device characteristics translate into circuit- and system-level tradeoffs in density, energy
efficiency, and functional heterogeneity. Together, these efforts establish a foundry-compatible M3D design
pathway toward heterogeneous cognitive computing systems.
BioZeshu Wang received the B.E. degree with honors in microelectronics from Wuhan University, Wuhan, China, in 2022, and the M.E. degree in electrical engineering from Tsinghua University, Beijing, China, in 2025, where he worked on heterogeneous integration of perovskite-based optoelectronic devices and systems. He is currently pursuing the Ph.D. degree with the Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA, advised by Prof. Haitong Li. His research interests include monolithic 3D integrated circuits, memory-centric computing with BEOL technologies, probabilistic computing hardware, and device-circuit-system co-design for heterogeneous computing systems.
|
|
37 |
C. Baddouh, T. Rogers |
Beyond PTX: Generating and Modifying Native GPU Code for AI Systems Research
AbstractEfficient AI kernels depend on precise control over GPU execution, yet NVIDIA's native instruction set, SASS, remains largely undocumented, restricting code generation and optimization below PTX. We present a framework for systematically generating, modifying, and evaluating native GPU code. The framework extracts machine-readable instruction specifications from NVIDIA tooling and uses encoding constraints to generate candidate instructions for validation on real hardware. Its binary-editing interface enables researchers to replace and reorder instructions, modify operands and scheduling controls, and patch existing kernels directly. Across seven architectures spanning Volta through Blackwell, we validate 539-1,338 instruction classes per target. We demonstrate these capabilities through instruction timing and contention experiments, as well as direct RT Core traversal from a standalone compute kernel. This work provides an experimental foundation for developing and evaluating code-generation strategies, exploring instruction sequences beyond those emitted by existing toolchains, and investigating low-level optimizations for AI kernels.
|
|
38 |
Y. Liu, T. Rogers |
Characterizing Communication-Computation Interference in LLM Training and Inference
AbstractOverlapping communication with computation is now standard practice in large-scale LLM training and inference. However, the interference between the two has been studied mostly at the application level, lacking a system-wide view and architectural insight. We characterize communication?computation co-execution in both LLM training (Megatron-LM) and inference (vLLM), across parallelism strategies (DP, FSDP, TP, PP, and EP) on NVIDIA H200 GPUs, using hardware counters that capture SM occupancy, L2, HBM, and NVLink behavior and attributing them to individual overlapped kernels. Our in-depth analysis of the resulting interference reveals several significant observations about which kernels are damaged and why, and yields practical guidelines for both software and hardware optimization.
|
|
39 |
J. Pan, T. Rogers |
Accel-Sim 2.0: Designing Asynchronous, Distributed GPUs in the AI Era
AbstractThe rapid evolution of machine learning has fundamentally altered GPU design, pushing architectures toward Multi-Chip Module topologies, asynchronous execution, and persistent, multi-phase kernel behaviors. Despite these profound shifts, the architecture community lacks cycle-level simulation frameworks that can natively capture the physical non-uniformity of contemporary GPUs and the massive scale of state-of-the-art AI workloads. To close this gap, this paper introduces a robust simulation infrastructure that accurately models Ampere, Hopper and Blackwell. Specifically, we evaluate the microarchitectural impact of off-core memory bandwidth expansion and asynchronous synchronization overheads. Furthermore, our analysis of intra-GPU NUMA effects reveals that while LLMs inherently tolerate spatial latency, the resulting data replication across partitions diminishes effective cache capacity, severely degrading locality-dependent workloads like graph analytics and CNNs. Finally, we demonstrate that inter-GPU memory prefetching is severely constrained by register file capacity limits.
We present Accel-Sim v2.0, a cycle-level simulation framework validated against Ampere, Hopper, and Blackwell architectures. Our model consistently achieves a 99\% Pearson correlation coefficient across contemporary silicon, alongside mean absolute cycle errors of 13.5% on Hopper H100 and 8.9% on Blackwell B200, respectively. Using Accel-Sim v2.0, we conduct a series of architectural investigations to expose new design trajectories surrounding asynchrony, chiplets, increased off-core bandwidth, and inter-GPU prefetching.
|
|
40 |
A. Nallathambi, A. Raghunathan |
A3D: Agentic AI flow for Autonomous Accelerator Design
AbstractAccelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite advances, such as high-level synthesis (HLS), designing accelerators for complex applications still remains highly labor-intensive, demanding considerable expertise in understanding workloads to be accelerated, hardware design, micro-architecture, and EDA tool usage, posing challenges for application domain experts. Therefore, most accelerator solutions are limited to applications with a regular predictable dataflow. Advances in AI have enabled agents that perform autonomous planning, reasoning, execution and reflection, leading to unprecedented potential for automation through agentic AI. We present A3D, an Agentic AI flow for end-to-end Automation of hardware Accelerator Design. A3D automates workload analysis, performance bottleneck identification, code refactoring for HLS compatibility and micro-architecture generation. A3D also generates diverse accelerator designs by automatically exploring the speed-area tradeoff space. Recent efforts have explored the use of AI for specific tasks such as design space exploration in HLS, leaving several tasks to still be performed manually. A3D addresses the challenges in applying modern LLMs to accelerator design by judiciously partitioning tasks among specialist agents, orchestrating process loops with specialist and verifier agents, utilizing pre-existing and custom tools, and employing agentic RAG for codebase and proprietary EDA tool documentation exploration.
|
|
41 |
R. Maity, K. Roy |
HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
AbstractAnalog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNN). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from crossbars leads to substantial power and area overhead. Moreover, the high area overhead of ADCs constrains the throughput due to the limited number of ADCs that can be integrated per crossbar. An approach to mitigate this issue involves the adoption of extreme low-precision quantization (binary or ternary) for partial sums. Training based on such an approach eliminates the need for ADCs. While this strategy effectively reduces ADC costs, it introduces the challenge of managing numerous floating-point scale factors, which are trainable parameters like DNN weights. These scale factors must be multiplied with the binary or ternary outputs at the columns of the crossbar to ensure system accuracy. To that effect, we propose an algorithm-hardware co-design approach, where DNNs are first trained with quantization-aware training. Subsequently, we introduce HCiM, an ADC-Less Hybrid Analog-Digital CiM accelerator. HCiM uses analog CiM crossbars for performing Matrix-Vector Multiplication operations coupled with a digital CiM array dedicated to processing scale factors. This digital CiM array can execute both addition and subtraction operations within the memory array, thus enhancing processing speed. Additionally, it exploits the inherent sparsity in ternary quantization to achieve further energy savings. Compared to an analog CiM baseline architecture using 7 and 4-bit ADC, HCiM achieves energy reductions up to 28% and 12%, respectively.
BioRiddhiman received his Bachelor of Technology (Honours) degree in 2025 from the Indian Institute of Technology Kharagpur in Electronics and Electrical Communication Engineering. During his undergraduate years, he worked at the Space Application Centre (SAC), ISRO, as an Undergraduate Research Fellow and contributed to the tape-out of a SERDES IC. He completed his research internship at NRL, Purdue, in the summer of 2024. He was selected for a summer internship at Texas Instruments in Bangalore, India, as an analog design intern. He also contributed to the design of an SRAM-based CiM hardware in FinFET technology. After completing his undergraduate studies, he joined Purdue University in the fall of 2025 as a Ph.D. student in Electrical and Computer Engineering under the supervision of Professor Kaushik Roy. His research interest lies on post-CMOS device-circuit co-integration and hardware-algorithm co-design. Apart from research, Riddhiman is interested in painting and traveling.
|
|
| Hackathon Finalists |
42 |
E. Yoon |
Agents for IC Physical Design
|
|
43 |
X. Dai, S. Mohammadi |
Analog Mixed-Signal Layout Agent
|
|
44 |
A. Grabka, R. Oum |
AntoRo / Training Neural Networks Robust to Soft Faults in ReRAM Crossbar Arrays
|
|
45 |
S. Cowen, R. Spitzenberger |
Carbon Compute
|
|
46 |
J. Balasubramanian, T. Panchagnula, S. Rank |
Compute Commons
|
|
47 |
R. Brar, H. Koh, A. Sridhar, K. Zisko |
KERA
|
|
48 |
A. Arens, R. Harrison, W. Tao |
Laser Drone Detection
|
|
49 |
S. Akella, A. Alam, S. Velmurugan |
RouteEdge
|
|
50 |
J. Urgiles, C. Wai |
Tiered Wakeup Visual Classification Cascade
|
|