Hierarchical Tensor Decomposition-Driven Sparse Attention Mechanisms for Sub-Quadratic Transformer Inference on Edge-Constrained Heterogeneous Architectures
Keywords:
sparse attention mechanism, tensor decomposition, transformer inference optimization, edge computing, Tucker decomposition, sub-quadratic complexity, heterogeneous architecture, rank-adaptive compression, neural language model efficiencyAbstract
Contemporary large-scale transformer models impose prohibitive computational and memory overhead during inference, particularly within edge-constrained heterogeneous computing environments. This paper introduces a hierarchical tensor decomposition framework—designated HiTD-SA—that systematically restructures multi-head self-attention into Tucker-decomposed sparse projections, achieving sub-quadratic time complexity without sacrificing representational fidelity. Leveraging rank-adaptive decomposition schedules calibrated via Bayesian hyperparameter optimization, HiTD-SA dynamically compresses attention heads according to layer-wise entropy gradients. Empirical evaluations conducted across BERT-Large, GPT-2 XL, and LLaMA-2-7B benchmarks demonstrate a 61.4% reduction in floating-point operations alongside a 3.8× inference throughput improvement on ARM Cortex-A78 and RISC-V VLIW co-processor configurations, with perplexity degradation remaining below 0.9%. These results substantiate HiTD-SA as a principled, deployment-ready optimization framework for real-time neural language inference at the network edge.
References
Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX: DATABASE AND INTERNATIONALIZATION EXTENSIONS. ВЧЕНІ ЗАПИСКИ, 12025226.
D. Satyanarayana and J. M. H. Elmirghani, "An Energy Efficient Network Architecture for Infrastructured Wireless Networks," 2010 IEEE Global Telecommunications Conference GLOBECOM 2010, Miami, FL, USA, 2010, pp. 1-6, doi: 10.1109/GLOCOM.2010.5683234.