Skip to main content

Why speaker attribution is hard in real time

Live captions, meeting assistants, and call analytics need both the words being spoken and a speaker label for each turn before the conversation ends. pyannoteAI streaming diarization provides speaker labels and turn boundaries as audio arrives. A streaming speech-to-text service provides the words. Because both APIs emit events independently, the client must combine their outputs into one live, speaker-attributed transcript. This approach opens one diarization stream and one transcription stream. The transcription stream receives audio from all speakers. One transcription session keeps session costs independent of the speaker count, but requires reconciliation logic to associate transcript events with speaker turns. The per-speaker approach uses one transcription session per speaker slot and buffers audio before routing it. Evaluate both approaches with your audio, cost, and latency requirements.

Build a live speaker-attributed transcript

In this tutorial, we will build a Python script that listens to your microphone and prints a live, speaker-attributed transcript:
Read Combining real-time diarization and transcription for the concepts behind both designs. This tutorial implements Approach 1. To give each known speaker a separate transcription stream, see Approach 2: per-speaker transcription streams.

Prerequisites

  • A pyannoteAI API key from the dashboard
  • An OpenAI API key
  • Python 3.10+
  • uv
  • Microphone access
The full script declares its dependencies inline. uv run installs them in an isolated environment automatically. Create a .env file:

Architecture overview

The script captures each microphone frame once and sends it to both streaming APIs: pyannoteAI receives 16 kHz mono float32 PCM and emits speaker start and end events. OpenAI receives the same audio converted to 24 kHz PCM16 and emits transcript deltas and completed segments. Shared state joins these two event streams.
This approach sends mixed microphone audio to one transcription stream. It does not separate simultaneous speakers, so overlapping speech can share a transcript segment.

1. Open two streaming sessions

First, create a pyannoteAI stream. The response contains an ID and a single-use WebSocket URL:
Both connections now exist before microphone capture starts. The pyannoteAI WebSocket accepts 16 kHz mono float32 PCM chunks every 100 ms. See Streaming Diarization for the full stream format and event reference. pyannoteAI controls the speaker boundaries and triggers each transcription commit. OpenAI uses gpt-realtime-whisper with turn detection disabled to transcribe those turns.

2. Capture the microphone once and fan out

The microphone callback puts each 100 ms frame into one asynchronous queue. Capturing once keeps both APIs on the same audio timeline:
Send the original float32 frame to pyannoteAI. Convert that same frame to the format OpenAI expects, then append it to the OpenAI input buffer:

3. Build the reconciler

The reconciler tracks four pieces of shared state:
  • active contains the speakers pyannoteAI currently hears. The most recently active speaker labels live transcript deltas.
  • pending is a queue of pyannoteAI speaker labels waiting for OpenAI to acknowledge a committed audio buffer.
  • pyannote_speaker_by_item associates each OpenAI transcript item with its pyannoteAI speaker label.
  • uncommitted records whether the OpenAI buffer contains audio that can be committed.

4. Attribute completed text

When pyannoteAI reports the end of a speaker turn, add that pyannoteAI speaker label to pending and commit the OpenAI audio buffer. Committing the buffer instructs OpenAI to complete the transcript segment for that turn:
OpenAI acknowledges the commit with an item_id. Associate that ID with the pending pyannoteAI label. Completion events can arrive out of order, so use item_id to recover the correct pyannoteAI speaker:
The script uses OpenAI’s item ID only to match transcription events. Every speaker label comes from pyannoteAI.

5. Update partial attribution as text arrives

OpenAI sends partial transcript deltas before a segment is complete. Store partial text by item_id, then display it with the associated pyannoteAI label:
Each delta uses the latest pyannoteAI attribution available for that transcript item. Once OpenAI acknowledges the commit, item_id fixes the completed text to the queued pyannoteAI label.

Full code

Here is the complete script:

Run the script

Save the full code as live_diarized_transcription.py, then run it:
Speak into your microphone. Partial text updates in place while OpenAI streams deltas. Completed turns print as separate lines with speaker labels like [SPEAKER_00]. Example output:

Next steps