Latency is how long one request takes; throughput is how many requests you complete per second. Improving one often costs the other, as batching does.
You have done this if
You streamed tokens to the browser so the first words appeared quickly even though the full answer took longer.
Say it in a review
We optimise time to first token for chat and throughput for batch jobs; they get different settings.