Abstract:
Although the Transformer architecture has become the cornerstone of artificial intelligence models, its massive parameter scale poses significant challenges to efficient inference. Leveraging the reconfigurability of FPGAs to build Domain-Specific Accelerators (DSAs) offers an effective pathway to overcome the bottlenecks of traditional instruction set architectures. However, existing solutions predominantly focus on accelerating specific operators, leading to frequent off-chip data movement and communication latency during end-to-end inference. Furthermore, their tight coupling with specific hardware severely limits portability. To address these limitations, this paper proposes TFETA, an end-to-end Transformer accelerator architecture based on INT8/FP16 mixed precision. The architecture implements dynamic and unified scheduling of multi-granularity operators via an on-chip interconnect bus, effectively eliminating off-chip communication overhead. To resolve pipeline stalls caused by the low arithmetic intensity of nonlinear operations, this work innovatively reconstructs the exp function into a “power-of-2 decomposition + dynamic reconfigurable floating-point look-up table” architecture. This is further optimized with on-chip coordinate transformation and adaptive zero-padding strategies to refine data flow. Experimental results demonstrate that the relative error of the proposed nonlinear unit remains stable within 0.1%, while operator latency is reduced by up to 27.97%. System validation shows that TFETA successfully deploys two representative models: ViT-B and Qwen2. In ViT-B testing, the accelerator achieves an inference speed of 41.8 FPS, a throughput of
1467.0 GOPS, and an energy efficiency of 41.8 GOPS/W. Without relying on structural pruning and without compromising model accuracy, this study outperforms state-of-the-art integer quantization schemes. Moreover, the hardware-agnostic RTL design provides a viable solution for the efficient deployment and seamless migration of Transformer models across heterogeneous FPGA platforms.