Academically, I am a computer architect. I study the design of the hardware components that efficiently solve problems written in software. My research has been recognized with an NSF Career Award in 2020, an NVIDIA graduate fellowship in 2013, best paper nominations at MICRO 2012, PPoPP 2017, ISPASS 2021, and IISWC 2024, an IEEE Micro Top Picks in Computer Architecture in 2013, as a Communications of the ACM Research Highlight in 2014, and my PhD thesis was nominated for the Governor General of Canada’s Gold Medal and the 2016 ACM Doctoral Dissertation Award. In 2024, I was inducted into the MICRO Hall of Fame, and in 2026 I was named an NVIDIA Research Faculty Fellow. My teaching has been recognized with multiple Outstanding Engineering Teacher citations, the 2018 Ruth and Joel Spira Award for Excellence in Teaching, the 2020 Hesselberth Award for Teaching Excellence, and the 2022 College of Engineering Excellence in Early Career Teaching Award. I completed my PhD in Computer Architecture at the University of British Columbia in 2015. During the course of my PhD, I interned for the research divisions of both AMD and NVIDIA, where I worked on the design of future GPU computing microarchitectures. Prior to entering graduate school, I worked as a software engineer at Electronic Arts where I gained insight into how industrial software was made before AI.
PhD in Computer Engineering, 2015
University of British Columbia
BEng in Honours Electrical Engineering, 2005
McGill University
NVIDIA’s native SASS instruction set is undocumented, leaving each research effort to rediscover it from scratch. EoSS is an end-to-end framework for exploring SASS systematically: it generates instruction encodings, executes them on real hardware, and characterizes how they behave across seven architectures spanning Volta through Blackwell.
Warp-specialized attention kernels lose performance when their pipeline phases fall out of balance. Progress-Aware Warp Scheduling (PAWS) uses pre-existing machine code fields to pass compiler-generated priority information to the warp scheduler, which then prioritizes warps in slower phases, delivering 28% speedup on highly skewed phases and 15% on QWen3, Llama4, and Grok 1.0 configurations.
Ray tracing units execute one ray per thread while the GPU manages resources at warp granularity. Ray-by-Ray decomposes the ray tracing unit’s on-chip buffers to per-ray granularity, letting more warps run concurrently without growing buffer capacity, and improves performance by 1.91x for a 0.19% area cost.
Expert GPU kernels apply Hopper’s Tensor Memory Accelerator to some operands but not others. We catalog the TMA, multicast, DSM, and cp.async choices made inside FlashAttention-3, isolate the features that discriminate between them, and produce predictive rules of thumb for the load-mechanism decision on H100.