Why use one transcription stream per speaker?
Live captions, meeting assistants, and call analytics need to answer one question: who said what? In this tutorial, we will combine pyannoteAI live diarization with one transcription stream per speaker. Each stream is assigned a pyannoteAI speaker label, so all text returned by that stream can be attributed directly to that speaker. The client combines timestamped results instead of matching transcript events to speaker turns. Sending each speaker’s speech to a separate transcription session may improve accuracy compared with sending all speakers to one session. This approach requires a known upper bound on the number of speakers. The client opens one transcription stream for each possible speaker, so transcription cost scales with that upper bound. It also buffers audio while waiting for speaker events, which adds transcription latency. Approach 1: single STT stream uses one transcription session for all speakers, but requires reconciliation logic. Evaluate both approaches with your audio, cost, and latency requirements.Read Combining real-time diarization and transcription for the concepts behind both designs. This tutorial implements Approach 2.
Build a live speaker-attributed transcript
We will build a command-line program that listens to your microphone and prints a live transcript:max_speakers to 2. It supports up to two speakers and opens two AssemblyAI sessions before capture starts. Both sessions incur transcription costs for the full call duration, even if the conversation has only one speaker.
Prerequisites
- A pyannoteAI API key from the dashboard
- An AssemblyAI API key from the AssemblyAI dashboard
- Python 3.10+
- uv
- Microphone access
uv run installs them in an isolated environment automatically.
Create a .env file:
Architecture overview
The script captures each microphone frame once. It sends the original float32 frame to pyannoteAI and stores the frame in a short buffer. The router uses speaker events to send PCM16 speech to each active speaker’s AssemblyAI stream and silence to the other streams. Speaker events arrive after the corresponding audio. The delayed audio buffer holds audio for one second to give those events time to arrive before routing. If the router processes speech before its speaker-start event arrives, it can send silence instead of speech. This fixed delay does not guarantee that all events arrive in time; adjust it for your network and processing latency. The diagram shows selected streams for a generalN; this example opens two. All N AssemblyAI streams receive the same amount of audio. Silence keeps each stream’s clock moving when its speaker is inactive. Word timestamps from all streams therefore use one common timeline.
1. Choose the speaker upper bound and open a pyannoteAI stream
Setmax_speakers to the maximum number of speakers that can join the conversation. This value is N, the speaker-count upper bound. The script must know it before opening the transcription sessions. The router assigns each new pyannoteAI speaker label to one of the N slots. pyannoteAI live diarization tracks up to eight speakers, so use a value from 1 to 8:
2. Open N AssemblyAI streams
Create one AssemblyAI WebSocket for each of theN speaker slots before audio capture starts:
universal-3-5-pro uses turn detection. min_turn_silence starts an end-of-turn check after a short pause, and max_turn_silence forces the turn to end after one second of silence. AssemblyAI speaker labels stay disabled because pyannoteAI supplies the speaker identity.
3. Route audio to each speaker
Store the pyannoteAI events with their timestamps:SPEAKER_00 talks, stream 0 receives the microphone samples and the other N - 1 streams receive the same number of zero samples.
4. Handle partials, finals, and the shared timeline
AssemblyAI emits aTurn message more than once as a turn develops. Its transcript field contains the current text for the full turn, so each partial replaces the prior partial from that stream:
end_of_turn is true, store the completed turn and remove that speaker’s partial:
start value places the turn on the shared timeline. The terminal renderer sorts completed and partial turns by this timestamp. Results can arrive from the N AssemblyAI sessions in a different order, but their display order remains chronological.
Full code
Save this complete script aslive_diarized_transcription_per_speaker.py:
Run the script
Run the file from the directory that contains.env:
Ctrl+C to stop. The script stops microphone capture, waits for pyannoteAI to finalize its speaker events, flushes the delayed audio, and waits for all N AssemblyAI sessions to terminate with their final turns.
Next steps
- To compare both designs, read Combining real-time diarization and transcription.
- To use one transcription stream with reconciliation logic and avoid paying for one session per speaker slot, see Approach 1: single STT stream.
- Set
max_speakersto the smallest reliable upper bound for your conversation. The script opens and pays for that many AssemblyAI sessions for the full call, including silent sessions. If more speakers join, the router mutes the extra speakers. - pyannoteAI live diarization supports up to eight speakers. If more than eight speakers join, it may merge speakers under the same labels.
- For production use, measure pyannoteAI event latency on your network and set
ROUTING_DELAYabove a high percentile of that measurement.