Estimate inference throughput in tokens per second from the generated token count and end-to-end latency, for capacity planning.
实际受批大小与排队影响。