Shrinking the Footprint, Boosting the Speed: Post-Training Quantization for Production Text Classifiers
Strategically applying post-training quantization can drastically reduce the memory footprint and inference latency of text classification models…