Set the image once, reuse its embedding
Notesegment_anything/predictor.py
`set_torch_image` validates the BCHW input and resized long side, resets previous image state, stores original/input dimensions, preprocesses pixels, and runs the image encoder. The resulting features remain on the predictor for subsequent prompt queries.
The higher-level `set_image` accepts HWC uint8 pixels, handles RGB/BGR ordering, resizes with `ResizeLongestSide`, and delegates here. Changing the image must replace the cached embedding; a prediction without a set image is rejected.
Evidence: [segment_anything/predictor.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fpredictor.py%3Afile%3Apredictor.py%3A).
Prompt coordinates, mask scores, and refinement
Definitionsegment_anything/predictor.py
`predict_torch` consumes batched prompts already transformed into the model's input frame. Points carry foreground/background labels; boxes use XYXY coordinates; a previous low-resolution mask can guide refinement.
It encodes prompts, combines them with cached image features in the mask decoder, restores masks to original image dimensions, and optionally thresholds logits. Results include masks, predicted quality scores, and low-resolution logits. With ambiguous prompts, multimask output provides alternatives; scores help the caller choose a mask.
The NumPy-facing `predict` performs coordinate conversion before delegating to this method.
Evidence: [segment_anything/predictor.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fpredictor.py%3Afile%3Apredictor.py%3A).
Sam owns three learned components
Definitionsegment_anything/modeling/sam.py
`Sam` composes `ImageEncoderViT`, `PromptEncoder`, and `MaskDecoder`. Its batched forward method preprocesses images, creates image embeddings, encodes each image's prompts, decodes masks, and postprocesses them to their original size.
It stores pixel mean/std as buffers and uses an RGB image convention. Output records include binary masks, mask-quality predictions, and low-resolution logits. Direct batched forward is useful when prompts are known in advance; the predictor supports repeated prompt interactions with cached image features.
Evidence: [segment_anything/modeling/sam.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fmodeling%2Fsam.py%3Afile%3Asam.py%3A).
Automatic masks: sample, filter, deduplicate
Summarysegment_anything/automatic_mask_generator.py
`SamAutomaticMaskGenerator` uses a `SamPredictor` plus point grids to generate masks without a user supplying every prompt. Options control sampling density, batches, optional crop layers, predicted-IoU and stability thresholds, and non-maximum suppression.
`generate` can remove small disconnected regions/holes and returns records with segmentation, area, XYWH box, predicted IoU, sample points, stability score, and crop box. Output can be binary masks or run-length encodings. More samples and crop layers increase inference work; the code remains separate from the learned model.
Evidence: [segment_anything/automatic_mask_generator.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fautomatic_mask_generator.py%3Afile%3Aautomatic_mask_generator.py%3A).
Interactive segmentation walkthrough
Guidenotebooks/predictor_example.ipynb
Follow the predictor example alongside `SamPredictor`: load a model/checkpoint, prepare an image, call `set_image`, supply point or box prompts to `predict`, and inspect the returned masks and scores. Feed selected low-resolution logits into a later query to refine an ambiguous result.
Keep coordinate spaces explicit: the high-level predictor accepts prompts in the original image frame and transforms them; `predict_torch` expects already transformed inputs.
The notebook was indexed through converted text. No checkpoint-dependent notebook or image segmentation was executed during export.
Evidence: [notebooks/predictor_example.ipynb](lumvise://element/filesystem%3A5149158d1c46f576%3Anotebooks%2Fpredictor_example.ipynb%3Afile%3Apredictor_example.ipynb%3A), [segment_anything/predictor.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fpredictor.py%3Afile%3Apredictor.py%3A).
Start here: segment-anything knowledge graph
GuideREADME.md
# segment-anything source tour
This demo combines the complete published semantic index with selected explanations attached to real files, folders, classes, and functions. Start with architecture, then follow the core concepts:
1. [Segment Anything architecture: reusable image features and prompts](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Aarchitecture)
2. [Set the image once, reuse its embedding](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Acache)
3. [Prompt coordinates, mask scores, and refinement](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Aprompts)
4. [Automatic masks: sample, filter, deduplicate](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Aautomatic)
5. [Sam owns three learned components](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Amodel)
6. [Interactive segmentation walkthrough](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Awalkthrough)
7. [Source snapshot, index coverage, and validation scope](lumvise://artifact/popular-demo-20260928%3Asegment-anything%3Aprovenance)
Select an artifact to inspect its owning semantic element. Evidence links point to indexed source. The source snapshot and coverage report records the exact scope and parser limitations.
Segment Anything architecture: reusable image features and prompts
Reportsegment_anything
# Segment Anything architecture
SAM predicts object masks from an image and prompts. The `Sam` model combines an image encoder, prompt encoder, and mask decoder. `SamPredictor` provides an interactive API that caches image features; `SamAutomaticMaskGenerator` samples prompts across an image and filters the resulting masks.
```text
image → image encoder → cached image features
points / boxes / prior mask → prompt encoder
both → mask decoder → full-size masks + scores
```
The model builder assembles the encoder/decoder configuration and optionally loads a checkpoint. The predictor owns image-specific state; the model owns learned components. This separation allows many prompt queries against one expensive image embedding.
Checkpoint weights are separate from this source/index export.
Evidence: [segment_anything/modeling/sam.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fmodeling%2Fsam.py%3Afile%3Asam.py%3A), [segment_anything/predictor.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fpredictor.py%3Afile%3Apredictor.py%3A), [segment_anything/automatic_mask_generator.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fautomatic_mask_generator.py%3Afile%3Aautomatic_mask_generator.py%3A), [segment_anything/build_sam.py](lumvise://element/filesystem%3A5149158d1c46f576%3Asegment_anything%2Fbuild_sam.py%3Afile%3Abuild_sam.py%3A).
Source snapshot, index coverage, and validation scope
ReportREADME.md
# Export provenance
Upstream: [facebookresearch/segment-anything](https://github.com/facebookresearch/segment-anything).
This graph was generated on 2026-09-28 from the existing local source folder. The folder has no Git metadata, so an exact upstream commit is unknown; no branch or commit is guessed. It was not updated from upstream during export.
Source snapshot fingerprint: `65fc31ee018f7731d4d005d59a904691232ba789ddf4e5e1fa2b4cc4df4ddb96` (SHA-256 over sorted relative paths, NUL separators, and raw file SHA-256 digests; excludes Git/runtime/generated cache directories and symlinks). Regular source files: 59. Indexed semantic elements: 905. File/text parser records: 49 (6 plain_text, 42 parsed, 1 unsupported); images have separate semantic kinds.
Coverage details:
- `demo/src/assets/index.html`: unsupported; project scan `demo/src/assets/index.html` failed: expected non-empty converted Markdown.
Static extraction is best effort. Unresolved dynamic calls are not evidence that dependencies are absent. Knowledge explanations were checked against selected local source; upstream test suites, notebooks, model inference, and model downloads were not run. The task validates index/export contents and readability.