Whisper knowledge graph

Explore Whisper as a knowledge graph: 2,722 files, symbols and docs including 1YOmY-Vjy-o, 3P_XnxdlXu0, 3elIlQzJEQ0, 43P4q1KGKEU — mapped with Lumvise.

What the graph contains

2,722 elements connected by 2,811 relationships.

  • 1,739 property
  • 286 parameter
  • 187 function
  • 157 field
  • 79 markdown_section
  • 64 object
  • 55 heading
  • 42 file

Reports and notes

load_model: checkpoint selection and model construction

Note

whisper/__init__.py

`load_model` accepts a registered model name or an existing checkpoint path. Named models use the download/cache path with SHA-256 verification. The function reads model dimensions and weights, constructs `Whisper`, loads its state dictionary, optionally configures alignment heads, and moves the model to the chosen device. Without an explicit device it chooses CUDA when available, otherwise CPU. Indexing and graph browsing do not invoke this loader or download checkpoints. Evidence: [whisper/__init__.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2F__init__.py%3Afile%3A__init__.py%3A).

transcribe: long-audio orchestration and fallback

Note

whisper/transcribe.py

`transcribe` prepares Mel features, resolves language/task tokenization, and processes audio windows. It returns complete text, segment details, and the detected or selected language. It can condition a window on previous text and add word timing from attention alignment. Its nested `decode_with_fallback` retries configured temperatures when repetition/compression or low average log probability indicates poor decoding. Silence rules can suppress retry for a no-speech segment. CPU execution changes unsupported FP16 decoding to FP32. These controls affect output quality; a graph of static calls cannot guarantee speech-recognition accuracy. Evidence: [whisper/transcribe.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftranscribe.py%3Afile%3Atranscribe.py%3A), [whisper/decoding.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fdecoding.py%3Afile%3Adecoding.py%3A), [whisper/tokenizer.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftokenizer.py%3Afile%3Atokenizer.py%3A).

Audio features: waveform to normalized log-Mel values

Definition

whisper/audio.py

`log_mel_spectrogram` accepts an audio path, NumPy array, or tensor. A path is decoded by `load_audio`; waveform inputs are expected at 16 kHz. It optionally pads audio, computes an STFT with a Hann window, applies Mel filters, clamps before the logarithm, limits the dynamic range, and scales the result. The Mel-filter count must match the model dimensions. The function returns a tensor shaped `(n_mels, n_frames)` for the audio encoder. Feature normalization is part of model input semantics, not presentation. Evidence: [whisper/audio.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Faudio.py%3Afile%3Aaudio.py%3A).

Whisper model: audio encoder and text decoder

Definition

whisper/model.py

The `Whisper` module builds an `AudioEncoder` and `TextDecoder` from `ModelDimensions`. Its forward path encodes Mel features and feeds the resulting audio features to the decoder alongside text tokens. The model exposes audio embedding, token logits, and key/value cache hooks. Cache hooks reuse decoder attention projections across token steps. Language detection, transcription, and decoding are attached from their owning modules, so orchestration rules stay outside the network class. Evidence: [whisper/model.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fmodel.py%3Afile%3Amodel.py%3A).

Start here: whisper knowledge graph

Guide

README.md

# whisper source tour This demo combines the complete published semantic index with selected explanations attached to real files, folders, classes, and functions. Start with architecture, then follow the core concepts: 1. [Whisper architecture: audio to timestamped text](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aarchitecture) 2. [load_model: checkpoint selection and model construction](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aload) 3. [Audio features: waveform to normalized log-Mel values](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aaudio) 4. [Whisper model: audio encoder and text decoder](lumvise://artifact/popular-demo-20260928%3Awhisper%3Amodel) 5. [transcribe: long-audio orchestration and fallback](lumvise://artifact/popular-demo-20260928%3Awhisper%3Atranscribe) 6. [Whisper verification map](lumvise://artifact/popular-demo-20260928%3Awhisper%3Atests) 7. [Source snapshot, index coverage, and validation scope](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aprovenance) Select an artifact to inspect its owning semantic element. Evidence links point to indexed source. The source snapshot and coverage report records the exact scope and parser limitations.

Source snapshot, index coverage, and validation scope

Report

README.md

# Export provenance Upstream: [openai/whisper](https://github.com/openai/whisper). This graph was generated on 2026-09-28 from the existing local source folder. The folder has no Git metadata, so an exact upstream commit is unknown; no branch or commit is guessed. It was not updated from upstream during export. Source snapshot fingerprint: `b9b46c21bcc07b7fd05848e7d69d7c40d85ba0a2b2b8964ac3a255cd5c366c72` (SHA-256 over sorted relative paths, NUL separators, and raw file SHA-256 digests; excludes Git/runtime/generated cache directories and symlinks). Regular source files: 45. Indexed semantic elements: 2722. File/text parser records: 42 (11 plain_text, 28 parsed, 2 unsupported, 1 binary); images have separate semantic kinds. Coverage details: - `whisper/assets/gpt2.tiktoken`: unsupported. - `whisper/assets/mel_filters.npz`: binary. - `whisper/assets/multilingual.tiktoken`: unsupported. Static extraction is best effort. Unresolved dynamic calls are not evidence that dependencies are absent. Knowledge explanations were checked against selected local source; upstream test suites, notebooks, model inference, and model downloads were not run. The task validates index/export contents and readability.

Whisper architecture: audio to timestamped text

Report

whisper

# Whisper inference pipeline Whisper separates audio preparation, the encoder-decoder model, short-window decoding, and long-audio orchestration. ```text audio → log-Mel features → audio encoder ↓ text decoder + tokenizer ↓ transcribe → text, segments, language ``` `audio.py` prepares model inputs. `model.py` owns the PyTorch network. `decoding.py` chooses tokens and exposes decoding options/results. `transcribe.py` coordinates windows, optional language detection, retry temperatures, context prompts, and timestamps. `load_model` assembles a model from a named checkpoint or local checkpoint file. The repository code is small; model weights are separate and are not included in this graph export. Evidence: [whisper/__init__.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2F__init__.py%3Afile%3A__init__.py%3A), [whisper/audio.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Faudio.py%3Afile%3Aaudio.py%3A), [whisper/model.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fmodel.py%3Afile%3Amodel.py%3A), [whisper/decoding.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fdecoding.py%3Afile%3Adecoding.py%3A), [whisper/transcribe.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftranscribe.py%3Afile%3Atranscribe.py%3A).

Documentation topics

  • CHANGELOG — CHANGELOG.md
  • [v20230117](https://github.com/openai/whisper/releases/tag/v20230117) — CHANGELOG.md
  • [v20230124](https://github.com/openai/whisper/releases/tag/v20230124) — CHANGELOG.md
  • [v20230306](https://github.com/openai/whisper/releases/tag/v20230306) — CHANGELOG.md
  • [v20230307](https://github.com/openai/whisper/releases/tag/v20230307) — CHANGELOG.md
  • [v20230308](https://github.com/openai/whisper/releases/tag/v20230308) — CHANGELOG.md
  • [v20230314](https://github.com/openai/whisper/releases/tag/v20230314) — CHANGELOG.md
  • [v20230918](https://github.com/openai/whisper/releases/tag/v20230918) — CHANGELOG.md
  • [v20231105](https://github.com/openai/whisper/releases/tag/v20231105) — CHANGELOG.md
  • [v20231106](https://github.com/openai/whisper/releases/tag/v20231106) — CHANGELOG.md
  • [v20231117](https://github.com/openai/whisper/releases/tag/v20231117) — CHANGELOG.md
  • [v20240927](https://github.com/openai/whisper/releases/tag/v20240927) — CHANGELOG.md
  • [v20240930](https://github.com/openai/whisper/releases/tag/v20240930) — CHANGELOG.md
  • [v20250625](https://github.com/openai/whisper/releases/tag/v20250625) — CHANGELOG.md
  • Approach — README.md
  • Available models and languages — README.md
  • Command-line usage — README.md
  • License — README.md
  • More examples — README.md
  • Python usage — README.md
  • Setup — README.md
  • Whisper — README.md
  • AMI-IHM, AMI-SDM1 — data/README.md
  • Artie — data/README.md
  • CHiME-6 — data/README.md
  • CORAAL — data/README.md
  • CallHome & Switchboard — data/README.md
  • CoVOST 2 — data/README.md
  • Common Voice 5.1 — data/README.md
  • Common Voice 9 — data/README.md

Types and modules

  • 1YOmY-Vjy-o — data/meanwhile.json
  • 3P_XnxdlXu0 — data/meanwhile.json
  • 3elIlQzJEQ0 — data/meanwhile.json
  • 43P4q1KGKEU — data/meanwhile.json
  • 4ktyaJkLMfo — data/meanwhile.json
  • 5Dsh9AgqRG0 — data/meanwhile.json
  • 748OyesQy84 — data/meanwhile.json
  • 8prs9Pq5Xhk — data/meanwhile.json
  • 9gX4kdFajqE — data/meanwhile.json
  • 9ssGpE9oem8 — data/meanwhile.json
  • ARw4K9BRCAE — data/meanwhile.json
  • B1DRmrOlKtY — data/meanwhile.json
  • BT950jqCCUY — data/meanwhile.json
  • C0e8XM30tQI — data/meanwhile.json
  • CKsASCGr_4A — data/meanwhile.json
  • DSc26qAJp_g — data/meanwhile.json
  • DhuCyncmFgM — data/meanwhile.json
  • EnGHyZS4f-8 — data/meanwhile.json
  • G8ajua4Mb5I — data/meanwhile.json
  • I4s-44cPYVE — data/meanwhile.json
  • JAfAApqOeFU — data/meanwhile.json
  • JfT59wBSQME — data/meanwhile.json
  • KT8pCZ5Xw9I — data/meanwhile.json
  • L-kR7UCzhTU — data/meanwhile.json
  • Lf-LkJhKVhk — data/meanwhile.json
  • P72uFdrkaVA — data/meanwhile.json
  • PT5_00Bld_8 — data/meanwhile.json
  • QPDZbNEhsuw — data/meanwhile.json
  • QjQbQlN9Ev8 — data/meanwhile.json
  • R6JV_I36It8 — data/meanwhile.json
  • RFVggCw58lo — data/meanwhile.json
  • RHDQpOVLKeM — data/meanwhile.json
  • TZSw9iRk03E — data/meanwhile.json
  • VV3UJmb8kHw — data/meanwhile.json
  • VYVbTzoggKc — data/meanwhile.json
  • WWWeV8xVNtI — data/meanwhile.json
  • XzJAtzdrY_w — data/meanwhile.json
  • YyV6l8HPmdQ — data/meanwhile.json
  • a8DD__mRtPk — data/meanwhile.json
  • cHhomJMwY1I — data/meanwhile.json

Functions

  • available_models — whisper/__init__.py
  • load_model — whisper/__init__.py
  • load_audio — whisper/audio.py
  • log_mel_spectrogram — whisper/audio.py
  • mel_filters — whisper/audio.py
  • pad_or_trim — whisper/audio.py
  • apply — whisper/decoding.py
  • cleanup_caching — whisper/decoding.py
  • decode — whisper/decoding.py
  • detect_language — whisper/decoding.py
  • finalize — whisper/decoding.py
  • logits — whisper/decoding.py
  • rank — whisper/decoding.py
  • rearrange_kv_cache — whisper/decoding.py
  • reset — whisper/decoding.py
  • run — whisper/decoding.py
  • scores — whisper/decoding.py
  • update — whisper/decoding.py
  • device — whisper/model.py
  • disable_sdpa — whisper/model.py
  • embed_audio — whisper/model.py
  • forward — whisper/model.py
  • install_hooks — whisper/model.py
  • install_kv_cache_hooks — whisper/model.py
  • is_multilingual — whisper/model.py
  • logits — whisper/model.py
  • num_languages — whisper/model.py
  • qkv_attention — whisper/model.py
  • save_to_cache — whisper/model.py
  • set_alignment_heads — whisper/model.py
  • sinusoids — whisper/model.py
  • remove_symbols — whisper/normalizers/basic.py
  • remove_symbols_and_diacritics — whisper/normalizers/basic.py
  • combine_cents — whisper/normalizers/english.py
  • extract_cents — whisper/normalizers/english.py
  • output — whisper/normalizers/english.py
  • postprocess — whisper/normalizers/english.py
  • preprocess — whisper/normalizers/english.py
  • process_words — whisper/normalizers/english.py
  • to_fraction — whisper/normalizers/english.py

How things connect

  • load_model calls _download
  • load_model instantiates Whisper
  • load_model calls available_models
  • __main__.py calls cli
  • FRAMES_PER_SECOND calls exact_div
  • N_FRAMES calls exact_div
  • TOKENS_PER_SECOND calls exact_div
  • log_mel_spectrogram calls load_audio
  • ApplyTimestampRules uses_type LogitFilter
  • BeamSearchDecoder uses_type TokenDecoder
  • GreedyDecoder uses_type TokenDecoder
  • MaximumLikelihoodRanker uses_type SequenceRanker
  • PyTorchInference uses_type Inference
  • SuppressBlank uses_type LogitFilter
  • SuppressTokens uses_type LogitFilter
  • __init__ instantiates GreedyDecoder
  • __init__ instantiates SuppressTokens
  • __init__ instantiates ApplyTimestampRules
  • __init__ instantiates MaximumLikelihoodRanker
  • __init__ calls _get_initial_tokens
  • __init__ instantiates BeamSearchDecoder
  • __init__ calls _verify_options
  • __init__ calls _get_suppress_tokens
  • __init__ instantiates SuppressBlank

Folders

  • data
  • .github
  • whisper
  • notebooks
  • whisper/assets
  • whisper/normalizers

Files

  • .flake8
  • README.md
  • .gitignore
  • MANIFEST.in
  • CHANGELOG.md
  • model-card.md
  • .gitattributes
  • data/README.md
  • pyproject.toml
  • requirements.txt
  • whisper/audio.py
  • whisper/model.py
  • whisper/utils.py
  • whisper/timing.py
  • whisper/version.py
  • data/meanwhile.json
  • whisper/__init__.py
  • whisper/__main__.py
  • whisper/decoding.py
  • whisper/tokenizer.py
  • whisper/transcribe.py
  • whisper/triton_ops.py
  • .pre-commit-config.yaml
  • notebooks/LibriSpeech.ipynb
  • whisper/assets/gpt2.tiktoken
  • whisper/normalizers/basic.py
  • whisper/assets/mel_filters.npz
  • whisper/normalizers/english.py
  • whisper/normalizers/__init__.py
  • notebooks/Multilingual_ASR.ipynb
  • whisper/normalizers/english.json
  • whisper/assets/multilingual.tiktoken