EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
原文英文,约100词,阅读约需1分钟。
📝
内容提要
本研究提出了EDIT(编码-解码图像变换器)架构,旨在解决视觉变换器模型中的注意力下沉问题。该方法通过层对齐的结构优化特征提取,提升了在ImageNet数据集上的性能。
🏷️