A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal AttentionDavis WertheimerHaochen Shenet al.2026COLM 2026Conference paper
DropKV: Decoupling Residual-Output Perturbation for Near-Optimal KV-Cache EvictionAozhong ZhangSelcuk Gurseset al.2026ICML 2026Workshop paper
Universal Position Interpolation: Unified Context Scaling for Hybrid Mamba-Transformer ModelsHaochen ShenDavis Wertheimeret al.2026ICLR 2026Conference paper
Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea FasoliMonodeep Karet al.2026ICLR 2026Conference paper
Frayed RoPE and Long Inputs: A Geometric PerspectiveDavis WertheimerAozhong Zhanget al.2026ICLR 2026Conference paper
Advancing Fluorescence Light Detection and Ranging in Scattering Media with a Physics-Guided Mixture-of-Experts and Evidential CriticsIsmail ErbasFerhat Demikiranet al.2025NeurIPS 2025Workshop paper
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory SystemYunhua FangRui Xieet al.2025IEEE Computer Architecture LettersPaper
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA InitializationYanxia DengAozhong Zhanget al.2025TMLRPaper
Generative AI Through CAS Lens: An Integrated Overview of Algorithmic Optimizations, Architectural Advances, and Automated DesignsChuan ZhangYou Youet al.2025IEEE JESTCSPaper
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex OptimizationWei LiuAnweshit Pandaet al.2025TMLRPaper
08 Mar 2021US10943732Magnetic Material Stack And Magnetic Inductor Structure Fabricated With Surface Roughness Control
KEKaoutar El MaghraouiTechnical Assistant to Priya Nagpurkar, VP AI Native Systems | Principal Research Scientist