Estimate the parameter and memory compression ratio when quantizing a model, for example from FP32 down to INT8, to plan deployment.
FP32→INT8 约 4× 压缩,需校准保精度。