Compute the number of training steps per epoch from dataset size and batch size, for training-loop and sharding checks.
分布式下每卡步数不变,总样本覆盖相同。