Skip to main content

Why use one transcription stream per speaker?

Live captions, meeting assistants, and call analytics need to answer one question: who said what? In this tutorial, we will combine pyannoteAI live diarization with one transcription stream per speaker. Each stream is assigned a pyannoteAI speaker label, so all text returned by that stream can be attributed directly to that speaker. The client combines timestamped results instead of matching transcript events to speaker turns. Sending each speaker’s speech to a separate transcription session may improve accuracy compared with sending all speakers to one session. This approach requires a known upper bound on the number of speakers. The client opens one transcription stream for each possible speaker, so transcription cost scales with that upper bound. It also buffers audio while waiting for speaker events, which adds transcription latency. Approach 1: single STT stream uses one transcription session for all speakers, but requires reconciliation logic. Evaluate both approaches with your audio, cost, and latency requirements.
Read Combining real-time diarization and transcription for the concepts behind both designs. This tutorial implements Approach 2.

Build a live speaker-attributed transcript

We will build a command-line program that listens to your microphone and prints a live transcript:
This example sets max_speakers to 2. It supports up to two speakers and opens two AssemblyAI sessions before capture starts. Both sessions incur transcription costs for the full call duration, even if the conversation has only one speaker.

Prerequisites

The full script declares its dependencies inline. uv run installs them in an isolated environment automatically. Create a .env file:

Architecture overview

The script captures each microphone frame once. It sends the original float32 frame to pyannoteAI and stores the frame in a short buffer. The router uses speaker events to send PCM16 speech to each active speaker’s AssemblyAI stream and silence to the other streams. Speaker events arrive after the corresponding audio. The delayed audio buffer holds audio for one second to give those events time to arrive before routing. If the router processes speech before its speaker-start event arrives, it can send silence instead of speech. This fixed delay does not guarantee that all events arrive in time; adjust it for your network and processing latency. The diagram shows selected streams for a general N; this example opens two. All N AssemblyAI streams receive the same amount of audio. Silence keeps each stream’s clock moving when its speaker is inactive. Word timestamps from all streams therefore use one common timeline.
Routing is not source separation. If multiple speakers talk into one microphone at the same time, each active stream receives the same mixed frame. Use separate input channels or a source-separation model when you must isolate overlapping voices.

1. Choose the speaker upper bound and open a pyannoteAI stream

Set max_speakers to the maximum number of speakers that can join the conversation. This value is N, the speaker-count upper bound. The script must know it before opening the transcription sessions. The router assigns each new pyannoteAI speaker label to one of the N slots. pyannoteAI live diarization tracks up to eight speakers, so use a value from 1 to 8:
The response provides a stream ID and a single-use WebSocket URL. pyannoteAI accepts 16 kHz mono float32 PCM in 100 ms frames and returns stable speaker labels with start and end timestamps. The script holds audio for one second before routing it:
Speaker events describe audio that the microphone already captured. This delay gives those events time to arrive before the matching buffered frame is sent to AssemblyAI. Increase it if network or processing latency causes clipped turn starts.

2. Open N AssemblyAI streams

Create one AssemblyAI WebSocket for each of the N speaker slots before audio capture starts:
universal-3-5-pro uses turn detection. min_turn_silence starts an end-of-turn check after a short pause, and max_turn_silence forces the turn to end after one second of silence. AssemblyAI speaker labels stay disabled because pyannoteAI supplies the speaker identity.
AssemblyAI bills each open streaming session by its duration, including periods that contain silence. Each session is billed independently, so N speaker slots produce N times the billed session duration of one stream. Always terminate every session when capture stops.

3. Route audio to each speaker

Store the pyannoteAI events with their timestamps:
For each buffered frame, allocate one row per speaker in a two-dimensional array. Copy samples to every row whose speaker is active. Leave the other rows at zero:
Convert each routed float32 array to little-endian PCM16 before sending it to AssemblyAI:
Each stream receives every frame. For example, while only SPEAKER_00 talks, stream 0 receives the microphone samples and the other N - 1 streams receive the same number of zero samples.

4. Handle partials, finals, and the shared timeline

AssemblyAI emits a Turn message more than once as a turn develops. Its transcript field contains the current text for the full turn, so each partial replaces the prior partial from that stream:
When end_of_turn is true, store the completed turn and remove that speaker’s partial:
The first word’s start value places the turn on the shared timeline. The terminal renderer sorts completed and partial turns by this timestamp. Results can arrive from the N AssemblyAI sessions in a different order, but their display order remains chronological.

Full code

Save this complete script as live_diarized_transcription_per_speaker.py:

Run the script

Run the file from the directory that contains .env:
Speak into your microphone. Partial text updates in place. Completed AssemblyAI turns remain on screen and sort by their first word timestamp. Press Ctrl+C to stop. The script stops microphone capture, waits for pyannoteAI to finalize its speaker events, flushes the delayed audio, and waits for all N AssemblyAI sessions to terminate with their final turns.

Next steps

  • To compare both designs, read Combining real-time diarization and transcription.
  • To use one transcription stream with reconciliation logic and avoid paying for one session per speaker slot, see Approach 1: single STT stream.
  • Set max_speakers to the smallest reliable upper bound for your conversation. The script opens and pays for that many AssemblyAI sessions for the full call, including silent sessions. If more speakers join, the router mutes the extra speakers.
  • pyannoteAI live diarization supports up to eight speakers. If more than eight speakers join, it may merge speakers under the same labels.
  • For production use, measure pyannoteAI event latency on your network and set ROUTING_DELAY above a high percentile of that measurement.
You can use another streaming transcription provider if it accepts continuous audio or explicit silence for each speaker stream. Change the audio conversion, connection setup, and event handler to match that provider. Partials may append text or replace the complete partial, turn-final signals use provider-specific fields, and timestamps may use a session clock or start at the first speech frame. Keep every stream on one common audio clock before merging its results.