GLM 4.7 Flash: Huge performance improvement with -kvu

Posted by TokenRingAI@reddit | LocalLLaMA | View on Reddit | 72 comments

TLDR; Try passing -kvu to llama.cpp when running GLM 4.7 Flash. On RTX 6000, my tokens per second on a 8K token output rose from 17.7t/s to 100t/s Also, check out the one shot zelda game it made, pretty good for a 30B: [https://talented-fox-j27z.pagedrop.io](https://talented-fox-j27z.pagedrop.io)