2-bit VPTQ: 6.5x Smaller LLMs whereas Preserving 95% Accuracy
Very correct 2-bit quantization for operating 70B LLMs on a 24 GB GPUGenerated with ChatGPTLatest developments in low-bit quantization for LLMs, like AQLM and AutoRound, at the moment are displaying...











