What is KV Cache? Why does it always mention when talking about large model reasoning acceleration and the cost of long dialogue?
KV Cache is a very important layer of caching mechanism in the inference stage of Transformers. To put it simply, it will save some of the keys and va...