Estimate the floating-point operations of a single attention layer and of the whole model from sequence length, head count and hidden size.
注意力和序列长度平方相关,长序列代价高。