原文英文,约400词,阅读约需2分钟。
📝
内容提要
We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention...
🏷️