高级检索

    面向Transformer的端到端推理加速硬件平台

    End-to-End Inference Acceleration Hardware Platform for Transformers

    • 摘要: Transformer架构已成为人工智能模型的基石,但其庞大的参数量给高效推理带来了严峻挑战。利用FPGA的可重构特性构建领域专用加速器,是突破传统指令集架构瓶颈的有效途径。然而,现有方案多局限于特定算子的加速,导致端到端推理面临频繁的片外数据交互与通信延迟,且其与特定硬件的高度耦合严重制约了可移植性。为此,提出一种基于INT8/FP16混合精度的端到端Transformer硬件加速器架构TFETA。该架构通过片上互连总线实现多粒度算子的动态统一调度,有效规避了片外通信开销。并且针对非线性计算因算术强度低导致的流水线阻塞问题,创新性地将exp函数重构为以2为底的幂分解与动态可重构浮点查找表架构,并结合片上坐标变换与自适应零填充策略优化数据流。该非线性计算单元的相对误差稳定在0.1%以内,算子执行延迟最高降低27.97%。系统验证显示,TFETA成功部署了ViT-B与Qwen2这2类典型模型。在ViT-B测试中,加速器实现了41.8 FPS的推理速度,吞吐量为1467.0 GOPS,能效比为41.8 GOPS/W。在不依赖结构剪枝且不损失模型精度的前提下,性能优于现有整型量化方案;同时,凭借全硬件无关的RTL设计,为Transformer模型在异构FPGA平台上的高效部署与移植提供了可行方案。

       

      Abstract: Although the Transformer architecture has become the cornerstone of artificial intelligence models, its massive parameter scale poses significant challenges to efficient inference. Leveraging the reconfigurability of FPGAs to build Domain-Specific Accelerators (DSAs) offers an effective pathway to overcome the bottlenecks of traditional instruction set architectures. However, existing solutions predominantly focus on accelerating specific operators, leading to frequent off-chip data movement and communication latency during end-to-end inference. Furthermore, their tight coupling with specific hardware severely limits portability. To address these limitations, this paper proposes TFETA, an end-to-end Transformer accelerator architecture based on INT8/FP16 mixed precision. The architecture implements dynamic and unified scheduling of multi-granularity operators via an on-chip interconnect bus, effectively eliminating off-chip communication overhead. To resolve pipeline stalls caused by the low arithmetic intensity of nonlinear operations, this work innovatively reconstructs the exp function into a “power-of-2 decomposition + dynamic reconfigurable floating-point look-up table” architecture. This is further optimized with on-chip coordinate transformation and adaptive zero-padding strategies to refine data flow. Experimental results demonstrate that the relative error of the proposed nonlinear unit remains stable within 0.1%, while operator latency is reduced by up to 27.97%. System validation shows that TFETA successfully deploys two representative models: ViT-B and Qwen2. In ViT-B testing, the accelerator achieves an inference speed of 41.8 FPS, a throughput of 1467.0 GOPS, and an energy efficiency of 41.8 GOPS/W. Without relying on structural pruning and without compromising model accuracy, this study outperforms state-of-the-art integer quantization schemes. Moreover, the hardware-agnostic RTL design provides a viable solution for the efficient deployment and seamless migration of Transformer models across heterogeneous FPGA platforms.

       

    /

    返回文章
    返回