对于一个 Transformer Decoder-only 模型,在处理一个长度为 L 的序列时,其自注意力(Self-Attention)机制的计算复杂度(按乘加运算次数计)大致为:
( O(L) )
( O(L^2) )
( O(L^3) )
( O(L^4) )