Micrograd knowledge graph

Explore Micrograd as a knowledge graph: 138 files, symbols and docs including Value, Layer, MLP, Module — mapped with Lumvise.

What the graph contains

138 elements connected by 142 relationships.

  • 36 markdown_section
  • 31 function
  • 22 parameter
  • 11 field
  • 10 file
  • 8 heading
  • 6 block
  • 5 class

Reports and notes

Decision-boundary image: inspect the learned regions

Note

moon_mlp.png

This upstream image illustrates the two-moons classifier described in the training notebook. Select this artifact to highlight the central plot region. Colored background regions represent the model's decisions; overlaid samples show their relation to the training set. This is a retained upstream result, not a newly trained model.

Notebook tour: learn the two-moons boundary

Note

demo.ipynb

The notebook generates 100 two-moons samples, maps labels to -1/+1, and builds `MLP(2,[16,16,1])`. Training uses a margin loss plus L2 regularization, clears gradients, calls backward, and updates parameters with SGD for 100 steps. The final cells plot predictions across a grid. Lumvise converted this notebook through AnyToMD; semantic coordinates refer to that Markdown representation. The notebook itself was inspected, not executed.

Trace notebook: connect the graph to its visualization

Note

trace_graph.ipynb

`trace` collects Value nodes and parent edges. `draw_dot` renders node values and gradients using Graphviz. The notebook backpropagates through a small expression and a two-input neuron, then exports `gout.svg`. Compare this graph with Value.backward to connect the visual graph with reverse traversal. Notebook contents were converted to Markdown; Graphviz execution was not performed.

Why backward accumulates gradients

Note

micrograd/engine.py

`backward()` visits each graph node once in topological order, seeds the output gradient to 1, then runs local derivative closures in reverse order. Closures use `+=`: a reused input can contribute through multiple paths. For `x=3; y=x*x+x`, the result is 12 and dy/dx is 7. Gradients persist on nodes; the training loop resets parameter gradients before each backward pass.

Value: a scalar and its derivative

Definition

micrograd/engine.py

`Value` holds a scalar, an accumulated gradient, parent nodes, and a local backward closure. Arithmetic constructs a directed acyclic computation graph. Each closure propagates the output gradient to its inputs using the chain rule. This intentionally scalar implementation makes automatic differentiation easy to inspect.

micrograd project overview

Definition

README.md

{"edges":[],"layers":[{"id":"neuron","label":"1. single neuron (engine.py)","visible":false},{"id":"network","label":"2. full network (nn.py)","visible":false},{"id":"loop","label":"3. training loop (demo.ipynb)","visible":true}],"nodes":[{"height":70,"id":"x1","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"#a5d8ff\",\"boundElements\":[],\"fillStyle\":\"solid\",\"groupIds\":[],\"height\":70,\"id\":\"x1\",\"opacity\":100,\"roughness\":1,\"seed\":11,\"strokeColor\":\"#1971c2\",\"strokeStyle\":\"solid\",\"strokeWidth\":2,\"type\":\"ellipse\",\"version\":2,\"versionNonce\":1614597638,\"width\":70,\"x\":40,\"y\":200,\"link\":null,\"index\":\"a0\",\"isDeleted\":false,\"frameId\":null,\"roundness\":null,\"updated\":1790440028226,\"locked\":false}","width":70,"x":40,"y":200},{"height":24,"id":"x1t","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"transparent\",\"boundElements\":[],\"fillStyle\":\"solid\",\"fontFamily\":1,\"fontSize\":20,\"groupIds\":[],\"height\":24,\"id\":\"x1t\",\"opacity\":100,\"originalText\":\"x1\",\"roughness\":1,\"seed\":12,\"strokeColor\":\"#1971c2\",\"strokeStyle\":\"solid\",\"strokeWidth\":1,\"text\":\"x1\",\"textAlign\":\"center\",\"type\":\"text\",\"version\":2,\"versionNonce\":1766030682,\"verticalAlign\":\"middle\",\"width\":30,\"x\":60,\"y\":223,\"link\":null,\"autoResize\":false,\"index\":\"a1\",\"isDeleted\":false,\"frameId\":null,\"roundness\":null,\"updated\":1790440028226,\"locked\":false,\"containerId\":null,\"lineHeight\":1.2}","width":30,"x":60,"y":223},{"height":70,"id":"x2","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"#a5d8ff\",\"boundElements\":[],\"fillStyle\":\"solid\",\"groupIds\":[],\"height\":70,\"id\":\"x2\",\"opacity\":100,\"roughness\":1,\"seed\":13,\"strokeColor\":\"#1971c2\",\"strokeStyle\":\"solid\",\"strokeWidth\":2,\"type\":\"ellipse\",\"version\":2,\"versionNonce\":1556159814,\"width\":70,\"x\":40,\"y\":330,\"link\":null,\"index\":\"a2\",\"isDeleted\":false,\"frameId\":null,\"roundness\":null,\"updated\":1790440028226,\"locked\":false}","width":70,"x":40,"y":330},{"height":24,"id":"x2t","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"transparent\",\"boundElements\":[],\"fillStyle\":\"solid\",\"fontFamily\":1,\"fontSize\":20,\"groupIds\":[],\"height\":24,\"id\":\"x2t\",\"opacity\":100,\"originalText\":\"x2\",\"roughness\":1,\"seed\":14,\"strokeColor\":\"#1971c2\",\"strokeStyle\":\"solid\",\"strokeWidth\":1,\"text\":\"x2\",\"textAlign\":\"center\",\"type\":\"text\",\"version\":2,\"versionNonce\":559858202,\"verticalAlign\":\"middle\",\"width\":30,\"x\":60,\"y\":353,\"link\":null,\"autoResize\":false,\"index\":\"a3\",\"isDeleted\":false,\"frameId\":null,\"roundness\":null,\"updated\":1790440028226,\"locked\":false,\"containerId\":null,\"lineHeight\":1.2}","width":30,"x":60,"y":353},{"height":20,"id":"w1a","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"#495057\",\"boundElements\":[],\"fillStyle\":\"solid\",\"groupIds\":[],\"height\":20,\"id\":\"w1a\",\"opacity\":100,\"roughness\":1,\"seed\":15,\"strokeColor\":\"#495057\",\"strokeStyle\":\"solid\",\"strokeWidth\":2,\"type\":\"arrow\",\"version\":2,\"versionNonce\":981042310,\"width\":110,\"x\":110,\"y\":235,\"link\":null,\"points\":[[0,0],[110,20]],\"index\":\"a4\",\"isDeleted\":false,\"frameId\":null,\"roundness\":null,\"updated\":1790440028226,\"locked\":false,\"startBinding\":null,\"endBinding\":null,\"lastCommittedPoint\":null,\"startArrowhead\":null,\"endArrowhead\":\"arrow\"}","width":110,"x":110,"y":235},{"height":22,"id":"w1t","kind":"group","text":"{\"angle\":0,\"backgroundColor\":\"transparent\",\"boundElements\":[],\"fillStyle\":\"solid\",\"fontFamily\":1,\"fontSize\":18,\"groupIds\":[],\"height\":22,\"id\":\"w1t\",\"opacity\":100,\"originalText\":\"w1\",\"roughness\":1,\"seed\":16,\"strokeColor\":\"#e8590c\",\"strokeStyle\":\"solid\",\"strokeWidth\":1,\"text\":\"w1\",\"textAlign\":\"center\",\"type\":\"text\",\"version\":2,\"versionNonce\":1731423962,\"verticalAlign\":\"middle\",\"width\":30,\"x\":145,\"y\":195,\"link\":null,\"autoResize\":false,\"index\":\"a5\",

Empty initializer: import engine and network modules explicitly

Summary

micrograd/__init__.py

`micrograd/__init__.py` is an empty package initializer. It defines no functions, classes, variables, or re-exports. Import the scalar engine explicitly with `from micrograd.engine import Value`; the network module can be imported with `from micrograd import nn` or its classes from `micrograd.nn`. Source reviewed on 2026-09-28: the file is zero bytes. This corrects the earlier generated description, which incorrectly claimed Value and the network classes were re-exported here.

From scalars to a 337-parameter network

Summary

micrograd/nn.py

`Neuron` combines weighted inputs plus bias, optionally applying ReLU. `Layer` groups neurons; `MLP` composes layers with a linear final layer. The notebook's `MLP(2, [16,16,1])` has 337 parameters: 48 + 272 + 17. A single-output layer returns a scalar Value. All differentiation is delegated to the same Value engine.

Functional meaning: __init__.py

Summary

micrograd/__init__.py

## Job Package initializer for micrograd; the span is empty, so the file exists only to mark the micrograd directory as an importable Python package. ## Source Interface Defines no classes, functions, or variables; importers get the bare micrograd package namespace with no public members. ## Receives - None; module body contains no statements or inputs ## Outcome Importing micrograd creates an empty package namespace module; no definitions, exports, or values are introduced. ## Notable Effects - None; empty module body performs no statements, imports, or I/O at import time

Functional meaning: engine.py

Summary

micrograd/engine.py

## Job Scalar reverse-mode autograd engine: Value wraps a scalar plus its gradient, records arithmetic ops in a DAG, and propagates gradients via backward(). ## Source Interface class Value: __init__(data, _children=(), _op=''), __add__, __mul__, __pow__, relu, backward, __neg__, __radd__, __sub__, __rsub__, __rmul__, __truediv__, __rtruediv__, __repr__ ## Receives - __init__: data (scalar), optional _children (iterable of Value parents), _op (str label) - __add__/__mul__/__sub__/__truediv__/etc.: a Value or numeric scalar (auto-wrapped into Value) - __pow__/__truediv__: exponent constrained by assert to int or float - backward(): no arguments; caller invokes it on a graph output ## Outcome Operators return a new Value with numeric result, parent links, op label, and a _backward closure; backward() seeds self.grad=1 and applies chain rule over all reachable nodes in reverse topological order, filling their .grad; __repr__ returns 'Value(data=..., grad=...)'; __pow__/div support int/float exponents, relu zeroes negatives. ## Notable Effects - Mutates .grad on operand nodes during backward passes (grads accumulate via +=) - Attaches a per-op _backward closure to each output node; initializes grad=0 and _backward=no-op - Stores _prev as a set of parent nodes and _op label per node for graph/debug use - assert rejects non-int/float exponents in __pow__ - Pure in-memory computation; no I/O or external side effects

Functional meaning: nn.py

Summary

micrograd/nn.py

## Job Defines a minimal neural-network stack (Module, Neuron, Layer, MLP) over micrograd Value scalars, providing parameter collection, gradient zeroing, and forward passes. ## Source Interface Classes: Module (zero_grad, parameters), Neuron(nin, nonlin=True) with __call__/parameters/__repr__, Layer(nin, nout, **kwargs), MLP(nin, nouts). Imports: random, micrograd.engine.Value. ## Receives - nin: int count of inputs per Neuron - nonlin: bool choosing ReLU vs linear activation - nout: int neuron count per Layer (via **kwargs) - nouts: list of layer widths for MLP - x: forward-pass input (scalar or sequence zipped with neuron weights) ## Outcome Neuron stores nin random uniform (-1,1) Value weights and a zero Value bias; calling it returns relu(b + sum(wi*xi)) if nonlin else the raw sum. Layer calls each neuron and returns a single output when it has one neuron, else a list. MLP stacks len(nouts) layers with ReLU on all but the last layer and returns the final x. parameters() flattens per-neuron/per-layer Value lists; Module.zero_grad resets each p.grad to 0; Module.parameters defaults to []. ## Notable Effects - Writes p.grad = 0 on each parameter Value during zero_grad - Draws from global RNG (random.uniform) when constructing Neuron weights - Allocates Value nodes and extends the autograd graph on each forward __call__ - No file, network, or stdout effects

Start here: micrograd across code, notebooks, and images

Guide

README.md

# Micrograd demo A small public project showing how Lumvise connects source code, converted notebooks, images, and attached explanations in one knowledge database. Follow the tour: 1. [Scalar values and derivatives](lumvise://artifact/micrograd-demo%3Avalue) 2. [Gradient accumulation](lumvise://artifact/micrograd-demo%3Abackward) 3. [A 337-parameter neural network](lumvise://artifact/micrograd-demo%3Anetwork) 4. [Training notebook](lumvise://artifact/micrograd-demo%3Atraining) 5. [Decision-boundary image region](lumvise://artifact/micrograd-demo%3Aimage) 6. [Computation graph notebook](lumvise://artifact/micrograd-demo%3Atrace) 7. [Verification](lumvise://artifact/micrograd-demo%3Avalidation) 8. [Source and license](lumvise://artifact/micrograd-demo%3Aprovenance) The source checkout is pinned to `7bc720e951fe422b8f8814aa5aa1b64121d26b4c`. Import: 13 files and 138 semantic elements. Both notebooks use AnyToMD-derived Markdown. Select the image annotation to inspect its highlighted region. Knowledge lives in the Lumvise database; upstream source files are unchanged.

Five-project export: Micrograd snapshot and coverage

Report

README.md

Refreshed the source index on 2026-09-28 for the five-project demo collection. The existing 17 Knowledge records were preserved, including source tours, scalar differentiation notes, neural-network explanations, notebook/image notes, and a project canvas. Source: https://github.com/karpathy/micrograd at commit `7bc720e951fe422b8f8814aa5aa1b64121d26b4c`. Regular source files: 13. Semantic elements: 138. Text-file coverage: 8 parsed and 2 plain text; 3 image elements are represented separately. No syntax or conversion errors reported. `.git` and `.lumvise` were intentionally ignored. Source-tree SHA-256: `124d40d099e0abb124cf3e54bcb293e6f6fd1bc08aa1b365df2ab6b91e042495` (sorted regular file paths, NUL, and raw file SHA-256 digests, excluding Git/runtime/cache directories). The upstream source was not modified. This run validates indexing and export contents; it does not rerun training or PyTorch tests. Earlier verification records retain their original dates.

Project structure

Report

micrograd

# Micrograd project structure and component interactions ## Purpose Micrograd is a deliberately small automatic-differentiation engine plus a compact neural-network library. The repository separates the mechanism that computes derivatives from the model components that use it, then demonstrates both through notebooks, tests, and a retained training result. For a guided tour of the existing project knowledge, see [Start here: micrograd across code, notebooks, and images](lumvise://artifact/micrograd-demo%3Atour). ## Project map | Area | Responsibility | Main entry | | ---------------------- | ------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | | Autodiff engine | Represents scalar values, records operations as a computation graph, and propagates gradients backward | [Value class](lumvise://element/filesystem%3Ab20ff338c283b44b%3Amicrograd%2Fengine.py%3Aclass%3AValue%3A) | | Neural-network library | Builds neurons, layers, and multilayer perceptrons from scalar `Value` operations | [Neural-network module](lumvise://element/filesystem%3Ab20ff338c283b44b%3Amicrograd%2Fnn.py%3Afile%3Ann.py%3A) | | Training demonstration | Creates a two-moons dataset, trains a classifier, and plots its decision regions | [MicroGrad demo](lumvise://element/filesystem%3Ab20ff338c283b44b%3Ademo.ipynb%3Amarkdown_section%3AMicroGrad%20demo%3A) | | Graph visualization | Traces a computation graph and renders values and gradients with Graphviz | [Trace notebook](lumvise://artifact/micrograd-demo%3Atrace) | | Verification | Compares engine values and gradients with PyTorch and covers shared-input gradient behavior | [Verification report](lumvise://artifact/micrograd-demo%3Avalidation) | | Output artifact | Shows the learned decision boundary from the two-moons example | [Decision-boundary image](lumvise://artifact/micrograd-demo%3Aimage) | ## Core components ### Scalar automatic differentiation The [Value class](lumvise://artifact/micrograd-demo%3Avalue) is the foundation. Each instance stores a scalar value, its accumulated gradient, its parent nodes, and a local backward function. Arithmetic between values creates new nodes and records the relationships needed for the chain rule. Calling `backward()` orders the graph from inputs to output, seeds the output gradient with one, and visits the nodes in reverse order. Local backward functions add their contributions to parent gradients, so a value reused along several paths receives every contribution. See [Why backward accumulates gradients](lumvise://artifact/micrograd-demo%3Abackward). ### Neural-network building blocks The neural-network module builds progressively larger objects: 1. [Neuron](lumvise://element/filesystem%3Ab20ff338c283b44b%3Amicrograd%2Fnn.py%3Aclass%3ANeuron%3A) combines weighted inputs and a bias, then optionally applies ReLU. 2. [Layer](lumvise://element/filesystem%3Ab20ff338c283b44b%3Amicrograd%2Fnn.py%3Aclass%3ALayer%3A) groups neurons that share the same input. 3. [MLP](lumvise://element/filesystem%3Ab20ff338c283b44b%3Amicrograd%2Fnn.py%3Aclass%3AMLP%3A) chains layers into a complete feed-forward network. These objects do not implement a second differentiation system. Their calculations use `Value` arithmetic, so the computation graph and gradients come from the engine automatically. The demo network, `MLP(2, [16, 16, 1])`, contains 337 parameters; see [F

micrograd

Report

micrograd

## Functional Component Map ```mermaid flowchart LR %% C4Container c0["<a class='internal-link is-unresolved' href='Lumvise Knowledge/karpathy__micrograd/Nuclei/micrograd/__init__.py/C4 Architecture'>__init__.py</a><br/>Job: Package initializer for micrograd; the span is empty; so the file exists only to mark the micrograd direc"] c1["<a class='internal-link is-unresolved' href='Lumvise Knowledge/karpathy__micrograd/Nuclei/micrograd/engine.py/C4 Architecture'>engine.py</a><br/>Job: Scalar reverse-mode autograd engine: Value wraps a scalar plus its gradient; records arithmetic ops in a "] c2["<a class='internal-link is-unresolved' href='Lumvise Knowledge/karpathy__micrograd/Nuclei/micrograd/nn.py/C4 Architecture'>nn.py</a><br/>Job: Defines a minimal neural-network stack (Module; Neuron; Layer; MLP) over micrograd Value scalars; providi"] c2 -->|instantiates x1| c1 ``` ## Component Jobs ### [__init__.py](<Lumvise Knowledge/karpathy__micrograd/Project/micrograd/__init__.py>) **Job:** Package initializer for micrograd; the span is empty, so the file exists only to mark the micrograd directory as an importable Python package. ### [engine.py](<Lumvise Knowledge/karpathy__micrograd/Project/micrograd/engine.py>) **Job:** Scalar reverse-mode autograd engine: Value wraps a scalar plus its gradient, records arithmetic ops in a DAG, and propagates gradients via backward(). ### [nn.py](<Lumvise Knowledge/karpathy__micrograd/Project/micrograd/nn.py>) **Job:** Defines a minimal neural-network stack (Module, Neuron, Layer, MLP) over micrograd Value scalars, providing parameter collection, gradient zeroing, and forward passes.

Documentation topics

  • Example usage — README.md
  • Installation — README.md
  • License — README.md
  • Running tests — README.md
  • Tracing / visualization — README.md
  • Training a GPT — README.md
  • Training a neural net — README.md
  • micrograd — README.md

Types and modules

  • Value — micrograd/engine.py
  • Layer — micrograd/nn.py
  • MLP — micrograd/nn.py
  • Module — micrograd/nn.py
  • Neuron — micrograd/nn.py

Functions

  • backward — micrograd/engine.py
  • build_topo — micrograd/engine.py
  • relu — micrograd/engine.py
  • parameters — micrograd/nn.py
  • zero_grad — micrograd/nn.py

How things connect

  • __add__ instantiates Value
  • __mul__ instantiates Value
  • __pow__ instantiates Value
  • backward calls build_topo
  • relu instantiates Value
  • Layer uses_type Module
  • MLP uses_type Module
  • Neuron uses_type Module
  • __init__ instantiates Value
  • __init__ instantiates Layer
  • __init__ instantiates Neuron
  • zero_grad calls parameters

Folders

  • micrograd

Files

  • setup.py
  • README.md
  • .gitignore
  • demo.ipynb
  • micrograd/nn.py
  • trace_graph.ipynb
  • micrograd/engine.py
  • micrograd/__init__.py