Why speaker attribution is hard in real time
Live captions, meeting assistants, and call analytics need both the words being spoken and a speaker label for each turn before the conversation ends. pyannoteAI streaming diarization provides speaker labels and turn boundaries as audio arrives. A streaming speech-to-text service provides the words. Because both APIs emit events independently, the client must combine their outputs into one live, speaker-attributed transcript. This approach opens one diarization stream and one transcription stream. The transcription stream receives audio from all speakers. One transcription session keeps session costs independent of the speaker count, but requires reconciliation logic to associate transcript events with speaker turns. The per-speaker approach uses one transcription session per speaker slot and buffers audio before routing it. Evaluate both approaches with your audio, cost, and latency requirements.Build a live speaker-attributed transcript
In this tutorial, we will build a Python script that listens to your microphone and prints a live, speaker-attributed transcript:Read Combining real-time diarization and transcription for the concepts behind both designs. This tutorial implements Approach 1. To give each known speaker a separate transcription stream, see Approach 2: per-speaker transcription streams.
Prerequisites
The full script declares its dependencies inline.uv run installs them in an isolated environment automatically.
Create a .env file:
Architecture overview
The script captures each microphone frame once and sends it to both streaming APIs: pyannoteAI receives 16 kHz mono float32 PCM and emits speaker start and end events. OpenAI receives the same audio converted to 24 kHz PCM16 and emits transcript deltas and completed segments. Shared state joins these two event streams.1. Open two streaming sessions
First, create a pyannoteAI stream. The response contains an ID and a single-use WebSocket URL:gpt-realtime-whisper with turn detection disabled to transcribe those turns.
2. Capture the microphone once and fan out
The microphone callback puts each 100 ms frame into one asynchronous queue. Capturing once keeps both APIs on the same audio timeline:3. Build the reconciler
The reconciler tracks four pieces of shared state:activecontains the speakers pyannoteAI currently hears. The most recently active speaker labels live transcript deltas.pendingis a queue of pyannoteAI speaker labels waiting for OpenAI to acknowledge a committed audio buffer.pyannote_speaker_by_itemassociates each OpenAI transcript item with its pyannoteAI speaker label.uncommittedrecords whether the OpenAI buffer contains audio that can be committed.
4. Attribute completed text
When pyannoteAI reports the end of a speaker turn, add that pyannoteAI speaker label topending and commit the OpenAI audio buffer. Committing the buffer instructs OpenAI to complete the transcript segment for that turn:
item_id. Associate that ID with the pending pyannoteAI label. Completion events can arrive out of order, so use item_id to recover the correct pyannoteAI speaker:
5. Update partial attribution as text arrives
OpenAI sends partial transcript deltas before a segment is complete. Store partial text byitem_id, then display it with the associated pyannoteAI label:
item_id fixes the completed text to the queued pyannoteAI label.
Full code
Here is the complete script:Run the script
Save the full code aslive_diarized_transcription.py, then run it:
[SPEAKER_00].
Example output:
Next steps
- Compare this reconciliation logic with per-speaker transcription streams (Approach 2).
- Read Combining real-time diarization and transcription for the tradeoffs between both approaches.
- Replace terminal output with WebSocket broadcasts to your frontend.
- Store completed
{speaker, text}turns in your database. - Add timestamps from pyannote events if your UI needs time-aligned captions.