// TOPIC
#streaming
3 articles
◆◆◆AdvancedOpenAIDeepgram
01Voice AI latency budget: hitting sub-500ms in production
A systematic breakdown of every millisecond in the voice AI pipeline and the specific techniques that compound to sub-500ms time-to-first-audio in production.
#voice#latency#production
18 min◆◆IntermediateOpenAIAnthropic
02Streaming LLM Responses: The Engineering Nobody Talks About
SSE vs WebSockets, partial-JSON parsing for streamed tool calls, backpressure, cancellation, and rendering streamed markdown without flicker — the complete engineering guide.
#streaming#latency#production
20 min◆◆IntermediateOpenAILiveKit
03VAD and Turn Detection: The Hardest UX Problem in Voice AI
How voice activity detection and turn-taking work in real-time AI agents — energy vs semantic VAD, barge-in handling, and why getting this wrong is the top user complaint.
#voice#latency#streaming
17 min