Users forgive a lot, but not silence. Past roughly a second without visible progress, an interface reads as broken — and that budget has to cover retrieval, model reasoning, and the first token of output.
A workable split
Retrieval gets about 400ms. Model time-to-first-token takes another 300–500ms depending on context size — which is one more reason trimming context pays off twice. That leaves almost nothing for your own orchestration, so the planning step needs to be cheap or skipped entirely for simple queries.
Two techniques buy real headroom. First, classify intent before retrieving: most queries don't need a deep multi-source fan-out, and routing them to a quick lookup saves hundreds of milliseconds. Second, stream the retrieval status to the UI. Showing which sources are being read converts dead time into visible progress, and the perceived budget roughly doubles.
When to break the budget deliberately
Research workloads are different. If a user has asked for a report, they'll wait minutes — provided the interface says so up front and shows intermediate work. The mistake is applying interactive-chat latency assumptions to a research task, or the reverse.