> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyannote.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Live diarized transcription: single transcription stream (Approach 1)

> Learn how to send mixed audio to one transcription stream and match transcript events with pyannoteAI speaker turns.

## Why speaker attribution is hard in real time

Live captions, meeting assistants, and call analytics need both the words being spoken and a speaker label for each turn before the conversation ends.

pyannoteAI streaming diarization provides speaker labels and turn boundaries as audio arrives. A streaming speech-to-text service provides the words. Because both APIs emit events independently, the client must combine their outputs into one live, speaker-attributed transcript.

This approach opens one diarization stream and one transcription stream. The transcription stream receives audio from all speakers.

One transcription session keeps session costs independent of the speaker count, but requires reconciliation logic to associate transcript events with speaker turns. The [per-speaker approach](/tutorials/live-diarized-transcription-per-speaker-stream) uses one transcription session per speaker slot and buffers audio before routing it. Evaluate both approaches with your audio, cost, and latency requirements.

## Build a live speaker-attributed transcript

In this tutorial, we will build a Python script that listens to your microphone and prints a live, speaker-attributed transcript:

```text theme={null}
[SPEAKER_00] Hello
[SPEAKER_00] This is transcribing while diarization is running
[SPEAKER_01] Now another person is speaking, and the transcript switches speakers.
[SPEAKER_00] And now the first speaker is back.
```

<Note>
  Read [Combining real-time diarization and transcription](https://pyannote.ai/blog/combining-real-time-diarization-and-transcription) for the concepts behind both designs. This tutorial implements Approach 1. To give each known speaker a separate transcription stream, see [Approach 2: per-speaker transcription streams](/tutorials/live-diarized-transcription-per-speaker-stream).
</Note>

## Prerequisites

* A pyannoteAI API key from the [dashboard](https://dashboard.pyannote.ai)
* An OpenAI API key
* Python 3.10+
* [uv](https://docs.astral.sh/uv/getting-started/installation/)
* Microphone access

The full script declares its dependencies inline. `uv run` installs them in an isolated environment automatically.

Create a `.env` file:

```bash theme={null}
PYANNOTEAI_API_KEY=sk_xxx
OPENAI_API_KEY=sk-xxx
```

## Architecture overview

The script captures each microphone frame once and sends it to both streaming APIs:

```mermaid theme={null}
%%{init: {'themeVariables': {'fontSize': '12px'}, 'flowchart': {'nodeSpacing': 20, 'rankSpacing': 24}}}%%
flowchart TB
    microphone[Microphone] --> queue[Audio queue]
    queue --> pyannote[pyannoteAI WebSocket]
    queue --> convert[Convert to 24 kHz PCM16]
    convert --> openai[OpenAI Realtime]
    pyannote -->|Speaker events| reconciler[Reconciler]
    openai -->|Transcript events| reconciler
    reconciler --> terminal[Terminal]
```

pyannoteAI receives 16 kHz mono float32 PCM and emits speaker start and end events. OpenAI receives the same audio converted to 24 kHz PCM16 and emits transcript deltas and completed segments. Shared state joins these two event streams.

<Warning>
  This approach sends mixed microphone audio to one transcription stream. It does not separate simultaneous speakers, so overlapping speech can share a transcript segment.
</Warning>

## 1. Open two streaming sessions

First, create a pyannoteAI stream. The response contains an ID and a single-use WebSocket URL:

```python theme={null}
def create_stream() -> tuple[str, str]:
    r = requests.post(
        "https://api.pyannote.ai/v1/live",
        headers={"Authorization": f"Bearer {PYANNOTEAI_API_KEY}"},
    )
    r.raise_for_status()
    data = r.json()
    return data["id"], data["url"]


stream_id, url = create_stream()
client = AsyncOpenAI(api_key=OPENAI_API_KEY)

async with websockets.connect(url) as ws:
    async with client.realtime.connect(extra_query={"intent": "transcription"}) as conn:
        await conn.session.update(session={
            "type": "transcription",
            "audio": {"input": {
                "format": {"type": "audio/pcm", "rate": OPENAI_RATE},
                "transcription": {"model": "gpt-realtime-whisper", "language": "en"},
                "turn_detection": None,
            }},
        })
        ...
```

Both connections now exist before microphone capture starts. The pyannoteAI WebSocket accepts 16 kHz mono float32 PCM chunks every 100 ms. See [Streaming Diarization](/tutorials/streaming-real-time) for the full stream format and event reference.

pyannoteAI controls the speaker boundaries and triggers each transcription commit. OpenAI uses `gpt-realtime-whisper` with turn detection disabled to transcribe those turns.

## 2. Capture the microphone once and fan out

The microphone callback puts each 100 ms frame into one asynchronous queue. Capturing once keeps both APIs on the same audio timeline:

```python theme={null}
def microphone(loop, queue: "asyncio.Queue[np.ndarray]") -> sd.InputStream:
    def enqueue(frame: np.ndarray) -> None:
        try:
            queue.put_nowait(frame)
        except asyncio.QueueFull:
            pass

    def callback(indata, frames, time, status):
        loop.call_soon_threadsafe(enqueue, indata[:, 0].copy())

    stream = sd.InputStream(
        samplerate=SAMPLE_RATE,
        channels=1,
        dtype="float32",
        blocksize=CHUNK,
        callback=callback,
    )
    stream.start()
    return stream
```

Send the original float32 frame to pyannoteAI. Convert that same frame to the format OpenAI expects, then append it to the OpenAI input buffer:

```python theme={null}
async def pump(queue: "asyncio.Queue[np.ndarray]", ws, convo: Conversation, conn) -> None:
    while True:
        frame = await queue.get()
        await ws.send(frame.astype("<f4").tobytes())
        pcm16 = (np.clip(upsample(frame), -1, 1) * 32767).astype("<i2").tobytes()
        await conn.input_audio_buffer.append(audio=base64.b64encode(pcm16).decode())
        convo.uncommitted = True
```

## 3. Build the reconciler

The reconciler tracks four pieces of shared state:

* `active` contains the speakers pyannoteAI currently hears. The most recently active speaker labels live transcript deltas.
* `pending` is a queue of pyannoteAI speaker labels waiting for OpenAI to acknowledge a committed audio buffer.
* `pyannote_speaker_by_item` associates each OpenAI transcript item with its pyannoteAI speaker label.
* `uncommitted` records whether the OpenAI buffer contains audio that can be committed.

```python theme={null}
class Conversation:
    def __init__(self) -> None:
        self.active: list[str] = []
        self.pending: deque[str] = deque()
        self.pyannote_speaker_by_item: dict[str, str] = {}
        self.uncommitted = False

    def started(self, speaker: str) -> None:
        if speaker not in self.active:
            self.active.append(speaker)

    def ended(self, speaker: str) -> None:
        if speaker in self.active:
            self.active.remove(speaker)

    @property
    def speaker(self) -> str | None:
        return self.active[-1] if self.active else None

    def committed(self, item_id: str) -> None:
        if self.pending:
            self.pyannote_speaker_by_item[item_id] = self.pending.popleft()

    def speaker_for(self, item_id: str) -> str | None:
        return self.pyannote_speaker_by_item.get(item_id, self.speaker)

    def completed(self, item_id: str) -> str | None:
        return self.pyannote_speaker_by_item.pop(item_id, self.speaker)
```

## 4. Attribute completed text

When pyannoteAI reports the end of a speaker turn, add that pyannoteAI speaker label to `pending` and commit the OpenAI audio buffer. Committing the buffer instructs OpenAI to complete the transcript segment for that turn:

```python theme={null}
elif msg.get("type") == "diarization_speaker_end":
    convo.ended(data["speaker"])
    if convo.uncommitted:
        convo.pending.append(data["speaker"])
        convo.uncommitted = False
        await conn.input_audio_buffer.commit()
```

OpenAI acknowledges the commit with an `item_id`. Associate that ID with the pending pyannoteAI label. Completion events can arrive out of order, so use `item_id` to recover the correct pyannoteAI speaker:

```python theme={null}
elif et.endswith("input_audio_buffer.committed"):
    convo.committed(event.item_id)
elif et.endswith("transcription.completed"):
    sp = convo.completed(event.item_id)
    text = event.transcript.strip()
    if text:
        tag = f"{color(sp)}[{sp}]{_RESET} " if sp else ""
        print(f"\r{tag}{text}\033[K")
```

The script uses OpenAI's item ID only to match transcription events. Every speaker label comes from pyannoteAI.

## 5. Update partial attribution as text arrives

OpenAI sends partial transcript deltas before a segment is complete. Store partial text by `item_id`, then display it with the associated pyannoteAI label:

```python theme={null}
elif et.endswith("transcription.delta"):
    partial = partials.get(event.item_id, "") + event.delta
    partials[event.item_id] = partial
    sp = convo.speaker_for(event.item_id)
    tag = f"{color(sp)}[{sp}]{_RESET} " if sp else ""
    print(f"\r{tag}{partial}\033[K", end="", flush=True)
```

Each delta uses the latest pyannoteAI attribution available for that transcript item. Once OpenAI acknowledges the commit, `item_id` fixes the completed text to the queued pyannoteAI label.

## Full code

Here is the complete script:

```python theme={null}
#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
#     "sounddevice",
#     "numpy",
#     "websockets>=13",
#     "openai[realtime]",
#     "python-dotenv",
#     "requests",
# ]
# ///

import asyncio
import base64
import json
import os
import signal
from collections import deque

import numpy as np
import requests
import sounddevice as sd
import websockets
from dotenv import load_dotenv
from openai import AsyncOpenAI

load_dotenv()
PYANNOTEAI_API_KEY = os.environ["PYANNOTEAI_API_KEY"]
OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]

SAMPLE_RATE = 16_000
OPENAI_RATE = 24_000
CHUNK = 1600

_COLORS = ["\033[32m", "\033[33m", "\033[34m", "\033[35m", "\033[36m", "\033[31m", "\033[37m", "\033[93m"]
_RESET = "\033[0m"
_colors: dict[str, str] = {}


def color(speaker: str) -> str:
    _colors.setdefault(speaker, _COLORS[len(_colors) % len(_COLORS)])
    return _colors[speaker]


class Conversation:
    def __init__(self) -> None:
        self.active: list[str] = []
        self.pending: deque[str] = deque()
        self.pyannote_speaker_by_item: dict[str, str] = {}
        self.uncommitted = False

    def started(self, speaker: str) -> None:
        if speaker not in self.active:
            self.active.append(speaker)

    def ended(self, speaker: str) -> None:
        if speaker in self.active:
            self.active.remove(speaker)

    @property
    def speaker(self) -> str | None:
        return self.active[-1] if self.active else None

    def committed(self, item_id: str) -> None:
        if self.pending:
            self.pyannote_speaker_by_item[item_id] = self.pending.popleft()

    def speaker_for(self, item_id: str) -> str | None:
        return self.pyannote_speaker_by_item.get(item_id, self.speaker)

    def completed(self, item_id: str) -> str | None:
        return self.pyannote_speaker_by_item.pop(item_id, self.speaker)


def upsample(x: np.ndarray) -> np.ndarray:
    n = len(x) * OPENAI_RATE // SAMPLE_RATE
    return np.interp(np.linspace(0, len(x) - 1, n), np.arange(len(x)), x).astype(np.float32)


def create_stream() -> tuple[str, str]:
    r = requests.post(
        "https://api.pyannote.ai/v1/live",
        headers={"Authorization": f"Bearer {PYANNOTEAI_API_KEY}"},
    )
    r.raise_for_status()
    data = r.json()
    return data["id"], data["url"]


async def pyannote_task(ws, convo: Conversation, conn) -> None:
    async for raw in ws:
        msg = json.loads(raw)
        data = msg.get("data", {})
        if msg.get("type") == "diarization_speaker_start":
            convo.started(data["speaker"])
        elif msg.get("type") == "diarization_speaker_end":
            convo.ended(data["speaker"])
            if convo.uncommitted:
                convo.pending.append(data["speaker"])
                convo.uncommitted = False
                await conn.input_audio_buffer.commit()
        elif msg.get("type") == "error":
            print(f"\npyannote error: {msg.get('message')}")


async def openai_task(conn, convo: Conversation) -> None:
    partials: dict[str, str] = {}
    async for event in conn:
        et = event.type
        if et.endswith("transcription_session.created"):
            print(f"OpenAI session:  {event.session.id}")
        elif et.endswith("input_audio_buffer.committed"):
            convo.committed(event.item_id)
        elif et.endswith("transcription.delta"):
            partial = partials.get(event.item_id, "") + event.delta
            partials[event.item_id] = partial
            sp = convo.speaker_for(event.item_id)
            tag = f"{color(sp)}[{sp}]{_RESET} " if sp else ""
            print(f"\r{tag}{partial}\033[K", end="", flush=True)
        elif et.endswith("transcription.completed"):
            partials.pop(event.item_id, None)
            sp = convo.completed(event.item_id)
            text = event.transcript.strip()
            if text:
                tag = f"{color(sp)}[{sp}]{_RESET} " if sp else ""
                print(f"\r{tag}{text}\033[K")


def microphone(loop, queue: "asyncio.Queue[np.ndarray]") -> sd.InputStream:
    def enqueue(frame: np.ndarray) -> None:
        try:
            queue.put_nowait(frame)
        except asyncio.QueueFull:
            pass

    def callback(indata, frames, time, status):
        loop.call_soon_threadsafe(enqueue, indata[:, 0].copy())

    stream = sd.InputStream(samplerate=SAMPLE_RATE, channels=1, dtype="float32", blocksize=CHUNK, callback=callback)
    stream.start()
    return stream


async def pump(queue: "asyncio.Queue[np.ndarray]", ws, convo: Conversation, conn) -> None:
    while True:
        frame = await queue.get()
        await ws.send(frame.astype("<f4").tobytes())
        pcm16 = (np.clip(upsample(frame), -1, 1) * 32767).astype("<i2").tobytes()
        await conn.input_audio_buffer.append(audio=base64.b64encode(pcm16).decode())
        convo.uncommitted = True


async def main() -> None:
    loop = asyncio.get_running_loop()
    convo = Conversation()
    queue: "asyncio.Queue[np.ndarray]" = asyncio.Queue(maxsize=50)

    stream_id, url = create_stream()
    print(f"pyannoteAI stream: {stream_id}")

    client = AsyncOpenAI(api_key=OPENAI_API_KEY)
    async with websockets.connect(url) as ws:
        async with client.realtime.connect(extra_query={"intent": "transcription"}) as conn:
            await conn.session.update(session={
                "type": "transcription",
                "audio": {"input": {
                    "format": {"type": "audio/pcm", "rate": OPENAI_RATE},
                    "transcription": {"model": "gpt-realtime-whisper", "language": "en"},
                    "turn_detection": None,
                }},
            })

            mic = microphone(loop, queue)
            print("\nListening — speak now (Ctrl+C to stop)\n")

            stop = asyncio.Event()
            loop.add_signal_handler(signal.SIGINT, stop.set)
            tasks = [
                asyncio.create_task(pyannote_task(ws, convo, conn)),
                asyncio.create_task(openai_task(conn, convo)),
                asyncio.create_task(pump(queue, ws, convo, conn)),
            ]
            await asyncio.wait({*tasks, asyncio.create_task(stop.wait())}, return_when=asyncio.FIRST_COMPLETED)

            print("\nStopping...")
            mic.stop()
            mic.close()
            for task in tasks:
                task.cancel()
            await asyncio.gather(*tasks, return_exceptions=True)
            await ws.send(json.dumps({"type": "end_of_stream"}))


if __name__ == "__main__":
    asyncio.run(main())
```

## Run the script

Save the full code as `live_diarized_transcription.py`, then run it:

```bash theme={null}
uv run live_diarized_transcription.py
```

Speak into your microphone. Partial text updates in place while OpenAI streams deltas. Completed turns print as separate lines with speaker labels like `[SPEAKER_00]`.

Example output:

```bash theme={null}
uv run live_diarized_transcription.py
pyannoteAI stream: 7d3f4a21-9b6c-4f8a-8f12-3c9d8b6e2a44

Listening — speak now (Ctrl+C to stop)

[SPEAKER_00] Hello
[SPEAKER_00] This is transcribing while diarization is running
[SPEAKER_00] The speaker label stays attached to this turn
[SPEAKER_01] Now another person is speaking, and the live transcript switches speakers.
[SPEAKER_01] It keeps printing completed turns as they arrive.
[SPEAKER_00] And now the first speaker is back.
^C
Stopping...
```

## Next steps

* Compare this reconciliation logic with [per-speaker transcription streams (Approach 2)](/tutorials/live-diarized-transcription-per-speaker-stream).
* Read [Combining real-time diarization and transcription](https://pyannote.ai/blog/combining-real-time-diarization-and-transcription) for the tradeoffs between both approaches.
* Replace terminal output with WebSocket broadcasts to your frontend.
* Store completed `{speaker, text}` turns in your database.
* Add timestamps from pyannote events if your UI needs time-aligned captions.
