Tao Lu and five coauthors submitted a paper to arXiv titled "Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices." The paper, categorized under cs.LG, was submitted on June 13, 2026, and addresses methods for improving GPU inference speed for large language models through sparse weight matrices.
No score is assigned. Sources and their independence are shown in the citation chain below.