load_model: checkpoint selection and model construction
Notewhisper/__init__.py
`load_model` accepts a registered model name or an existing checkpoint path. Named models use the download/cache path with SHA-256 verification. The function reads model dimensions and weights, constructs `Whisper`, loads its state dictionary, optionally configures alignment heads, and moves the model to the chosen device.
Without an explicit device it chooses CUDA when available, otherwise CPU. Indexing and graph browsing do not invoke this loader or download checkpoints.
Evidence: [whisper/__init__.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2F__init__.py%3Afile%3A__init__.py%3A).
transcribe: long-audio orchestration and fallback
Notewhisper/transcribe.py
`transcribe` prepares Mel features, resolves language/task tokenization, and processes audio windows. It returns complete text, segment details, and the detected or selected language. It can condition a window on previous text and add word timing from attention alignment.
Its nested `decode_with_fallback` retries configured temperatures when repetition/compression or low average log probability indicates poor decoding. Silence rules can suppress retry for a no-speech segment. CPU execution changes unsupported FP16 decoding to FP32.
These controls affect output quality; a graph of static calls cannot guarantee speech-recognition accuracy.
Evidence: [whisper/transcribe.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftranscribe.py%3Afile%3Atranscribe.py%3A), [whisper/decoding.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fdecoding.py%3Afile%3Adecoding.py%3A), [whisper/tokenizer.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftokenizer.py%3Afile%3Atokenizer.py%3A).
Audio features: waveform to normalized log-Mel values
Definitionwhisper/audio.py
`log_mel_spectrogram` accepts an audio path, NumPy array, or tensor. A path is decoded by `load_audio`; waveform inputs are expected at 16 kHz. It optionally pads audio, computes an STFT with a Hann window, applies Mel filters, clamps before the logarithm, limits the dynamic range, and scales the result.
The Mel-filter count must match the model dimensions. The function returns a tensor shaped `(n_mels, n_frames)` for the audio encoder. Feature normalization is part of model input semantics, not presentation.
Evidence: [whisper/audio.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Faudio.py%3Afile%3Aaudio.py%3A).
Whisper model: audio encoder and text decoder
Definitionwhisper/model.py
The `Whisper` module builds an `AudioEncoder` and `TextDecoder` from `ModelDimensions`. Its forward path encodes Mel features and feeds the resulting audio features to the decoder alongside text tokens.
The model exposes audio embedding, token logits, and key/value cache hooks. Cache hooks reuse decoder attention projections across token steps. Language detection, transcription, and decoding are attached from their owning modules, so orchestration rules stay outside the network class.
Evidence: [whisper/model.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fmodel.py%3Afile%3Amodel.py%3A).
Start here: whisper knowledge graph
GuideREADME.md
# whisper source tour
This demo combines the complete published semantic index with selected explanations attached to real files, folders, classes, and functions. Start with architecture, then follow the core concepts:
1. [Whisper architecture: audio to timestamped text](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aarchitecture)
2. [load_model: checkpoint selection and model construction](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aload)
3. [Audio features: waveform to normalized log-Mel values](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aaudio)
4. [Whisper model: audio encoder and text decoder](lumvise://artifact/popular-demo-20260928%3Awhisper%3Amodel)
5. [transcribe: long-audio orchestration and fallback](lumvise://artifact/popular-demo-20260928%3Awhisper%3Atranscribe)
6. [Whisper verification map](lumvise://artifact/popular-demo-20260928%3Awhisper%3Atests)
7. [Source snapshot, index coverage, and validation scope](lumvise://artifact/popular-demo-20260928%3Awhisper%3Aprovenance)
Select an artifact to inspect its owning semantic element. Evidence links point to indexed source. The source snapshot and coverage report records the exact scope and parser limitations.
Source snapshot, index coverage, and validation scope
ReportREADME.md
# Export provenance
Upstream: [openai/whisper](https://github.com/openai/whisper).
This graph was generated on 2026-09-28 from the existing local source folder. The folder has no Git metadata, so an exact upstream commit is unknown; no branch or commit is guessed. It was not updated from upstream during export.
Source snapshot fingerprint: `b9b46c21bcc07b7fd05848e7d69d7c40d85ba0a2b2b8964ac3a255cd5c366c72` (SHA-256 over sorted relative paths, NUL separators, and raw file SHA-256 digests; excludes Git/runtime/generated cache directories and symlinks). Regular source files: 45. Indexed semantic elements: 2722. File/text parser records: 42 (11 plain_text, 28 parsed, 2 unsupported, 1 binary); images have separate semantic kinds.
Coverage details:
- `whisper/assets/gpt2.tiktoken`: unsupported.
- `whisper/assets/mel_filters.npz`: binary.
- `whisper/assets/multilingual.tiktoken`: unsupported.
Static extraction is best effort. Unresolved dynamic calls are not evidence that dependencies are absent. Knowledge explanations were checked against selected local source; upstream test suites, notebooks, model inference, and model downloads were not run. The task validates index/export contents and readability.
Whisper architecture: audio to timestamped text
Reportwhisper
# Whisper inference pipeline
Whisper separates audio preparation, the encoder-decoder model, short-window decoding, and long-audio orchestration.
```text
audio → log-Mel features → audio encoder
↓
text decoder + tokenizer
↓
transcribe → text, segments, language
```
`audio.py` prepares model inputs. `model.py` owns the PyTorch network. `decoding.py` chooses tokens and exposes decoding options/results. `transcribe.py` coordinates windows, optional language detection, retry temperatures, context prompts, and timestamps. `load_model` assembles a model from a named checkpoint or local checkpoint file.
The repository code is small; model weights are separate and are not included in this graph export.
Evidence: [whisper/__init__.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2F__init__.py%3Afile%3A__init__.py%3A), [whisper/audio.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Faudio.py%3Afile%3Aaudio.py%3A), [whisper/model.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fmodel.py%3Afile%3Amodel.py%3A), [whisper/decoding.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Fdecoding.py%3Afile%3Adecoding.py%3A), [whisper/transcribe.py](lumvise://element/filesystem%3Aab1fe783b628fdc3%3Awhisper%2Ftranscribe.py%3Afile%3Atranscribe.py%3A).