The reply can begin quickly while the useful thing still arrives late.

A different clock

Voice products commonly measure how quickly sound starts coming back. That matters, but it can reward an immediate sentence that merely narrates work still waiting to happen.

We now look at the full interval: approval to executed action, background job to delivered result, open loop to closed loop and tool call to usable answer.

In the private pilot, one conversation can now place up to six independent jobs into the durable work queue. The dialogue stays available while they advance; each remains in progress until its own provider or browser receipt arrives, and that receipt closes the matching open work only when the match is unambiguous.

The tail is where trust goes

The operating view uses the median and the slow tail rather than one flattering average. A result that usually arrives in seconds but occasionally disappears for a day is not dependable.

This is an internal measure, not a public performance claim. Its purpose is to make the uncomfortable waits visible to the people building the product.

Short answers should start when they are ready

We have updated the private-pilot media call path so a short answer can start as soon as it passes the speech checks. We also combine sentences that are already waiting for speech synthesis to reduce repeated starts within a reply.

We also found that speech could be marked complete before all its audio arrived. The private-pilot v3 adapter now waits for the final event belonging to each speech request. Its readiness check requires two complete replies; live listening still needs to verify continuity between phrases.

A slow reply may give one brief thinking sound. Your speech or the answer can stop it, and it does not mean an action has completed. We still require the provider receipt before confirming an action.

The operations view now separates time spent waiting for the previous turn, preparing reply text and starting speech. Telephone playback confirmation remains an upper bound, and older calls do not gain measurements we never recorded. We will use live calls to evaluate the improvement; we make no response-time guarantee.

Language is part of latency

The shortest route to meaning is sometimes a systems improvement and sometimes an edited sentence. “Two things need you” can be more useful than a fast inventory of everything SAINT checked.

That is why voice presence, comfortable silence and operational latency belong to the same design problem.

Private-pilot calls now measure from the end of the person's speech to the beginning of audible reply, rather than beginning the clock after transcription has already finished. Personal English calls use a bounded, more responsive turn endpoint; supported listening words no longer stop the reply, and a slow answer can give one preloaded thinking breath without claiming that any work has completed. These are evaluation settings, not a public latency guarantee.