聚类 K-Means 代价

📊 K-Means Cost (Inertia) Estimator

Estimate K-Means inertia, the within-cluster sum of squares, from sample points and centroids, to compare clustering configurations.

📐 计算公式 / 原理
Inertia 估算 ≈ n·dim·spread²/k;每簇样本 ≈ ⌈n/k⌉
在样本近似均匀分布的假设下,簇内平方和随 K 增大按 1/k 衰减,可用于快速预判需要多少簇。实际数据有簇密度差异时,应结合业务含义定 K,而不是只看指标拐点。

📐 计算公式与说明

Inertia ≈ ΣΣ||x − μ||²

Inertia 越小簇内越紧密。

📚 深度解析:K-Means 代价

💡 常见使用场景

代价比较
k=2 时 inertia=500, k=3 时=300, k=4 时=220。inertia 持续下降,k=3→4 降幅(80)明显小于 2→3(200),结合业务选 k=3 或 4。

❓ 常见问题(FAQ)

inertia 越小越好吗?
不是。k 越大 inertia 越小,k=样本数时为 0 但失去聚类意义;应兼顾紧致性与簇数简洁,用手肘/轮廓系数平衡。
异常值怎么处理?
异常值会大幅拉高 inertia 并偏移质心,可先做离群检测或用 K-Medoids(基于距离中位数)更稳健。