observes that salient weights are always tied to high-magnitude activation channels. Rather than keeping these weights in higher, memory-heavy precision, AWQ uses a mathematically equivalent transformation to scale up the salient weight channels prior to quantization
