Neither Chain Nor Hairball: On the Shape of Memorable Knowledge
An edge in a knowledge graph carries information only insofar as it is selective. This sounds like a triviality. It is not. It is the reason the two most obvious ways to structure knowledge — a list and a web-of-everything — both fail, and fail for the same reason, at opposite ends of the same scale.
I want to make an argument here about the shape of knowledge structures and what that shape does to memory. The claim, briefly: there is an optimal topology for remembering, it is neither a chain nor a tree nor a complete graph, and — this is the part that gets missed — it is not a matter of how many edges you draw. It is a matter of where you draw them.
Every graph below has twenty nodes. Only the wiring changes.
Two ways to fail
Failure one: the chain. A—B—C—…—Z. Every node has exactly one predecessor and one successor. This is how a textbook is organized, how a lecture proceeds, how most notes are taken.
Its virtue is real: every edge is maximally diagnostic. If you are at K, the cue “next” points to exactly one thing. Nothing competes.
Its vices are structural. Path length is O(n) — reaching Z from A requires traversing everything between. Redundancy is zero — lose one link and everything downstream is orphaned. And retrieval is possible in essentially one order, which means the knowledge is available for exactly one purpose. Ebbinghaus built the hardest paradigm in the history of memory research on serial lists, and the difficulty was not an accident of the nonsense syllables. It was the shape.
Failure two: the complete graph. Every node connected to every other.
Twenty nodes, 190 edges, every pair joined. Maximal connectivity — and each node’s cue diagnosticity is 1/19.
It looks like abundance. It is the opposite, and the mechanism is well documented.
Anderson’s fan effect (1974) is the canonical demonstration. Participants learn facts pairing people with locations — “the doctor is in the bank.” When the doctor appears in three facts rather than one, verifying any single one is slower and less accurate. The explanation in Anderson’s ACT-R architecture is not vague: activation spreads from a concept to its associates, and the amount reaching each is divided among them, inversely proportional to their number. Each new edge you add to a node does not add retrieval strength to that node. It subtracts strength from every edge already there.
Watkins & Watkins (1975) named the same phenomenon from the retrieval side: the cue overload principle, that the probability of recalling an item declines with the number of items subsumed by its retrieval cue. A cue that points to everything points to nothing.
So the hairball is not a rendering problem you can fix with better layout. It has maximal connectivity and zero information per edge, and the layout is merely reporting that.
Notice the symmetry. The chain maximizes per-edge diagnosticity and destroys reachability. The complete graph maximizes reachability and destroys per-edge diagnosticity. Both are extremal, and both are bad, for reciprocal reasons.
The tree is better, and the evidence is unusually clean
Between the extremes sits the hierarchy.
Nineteen edges, four branches, depth three. One path between any two nodes.
Here the empirical result is one of the more striking in the memory literature. Bower, Clark, Lesgold & Winzenz (1969) presented participants with nested category lists — minerals divided into metals and stones, metals into rare and common — either arranged in their hierarchy or scrambled. Recall was two to three times better with the organized presentation. Not fifteen percent. Two to three times, from an intervention that changed nothing about the words themselves and nothing about study time.
Their analysis is worth noting: the hierarchy functioned as a retrieval plan. Participants used the superordinate categories to generate candidates, then checked those candidates against list membership. The structure was not decoration on the content; it was the algorithm by which the content was searched.
So trees work. Why isn’t this the end of the argument?
Because a tree still admits exactly one path between any two nodes. Cross-cutting relationships — the same idea appearing in two branches, the analogy between distant subfields, the concept that is genuinely a prerequisite for three different things — either get discarded or get duplicated. And a tree remains brittle in the chain’s way: forget a superordinate and everything below it becomes unreachable, because it had no other route in.
The real axis is topology, not density
Here the intuition “somewhere between sparse and dense” needs to be made precise, because as usually stated it is wrong.
There is no optimal average degree. Density is not the variable. The next three graphs prove it: each has exactly forty edges, and every node in all three has exactly four neighbours. Identical density, identical degree distribution. Only placement differs.
Order. Each node joined to its two nearest neighbours on each side. High clustering — and an empty interior, which is what “long paths” looks like.
Every edge hugs the rim. To get from one side to the other you walk the circumference. Local neighbourhoods are dense and redundant; global distance is terrible.
Chaos. Same forty edges, placed arbitrarily. Short paths — and no local structure at all.
Now distance is small: any node reaches any other in a couple of hops. But the neighbourhoods have dissolved. No triangles, no redundancy, no sense that related things sit near each other. Nothing is near anything.
Watts & Strogatz (1998) gave us the parameter that connects these two. Start from the lattice and rewire each edge, with probability p, to a random target. At p=0 you have the first picture; at p=1 the second. The finding that made the paper famous is what happens in between: over a broad interval of small p, path length collapses toward the random-graph value while the clustering coefficient stays near its lattice value.
Four edges rewired — shown in orange. Still forty edges, still degree four everywhere. The rim survives; the interior is now crossed.
Four edges. That is the entire difference between this picture and the lattice, and it is enough to cut average path length by more than half while leaving local clustering essentially intact.
“Order plus chaos” is not a metaphor here. It is a published curve with a parameter on the x-axis, and the four orange lines are what that parameter looks like at p ≈ 0.1.
Translated into knowledge-structure terms, what lives in that interval:
- High local clustering — the neighbours of a concept are themselves related, so any region has redundant internal routes. Forgetting one link orphans nothing.
- Short global paths — a few long-range links put distant regions a few hops apart, not a full traversal away.
- Degree heterogeneity — a handful of hub concepts serve as entry points, which is what makes the whole thing navigable.
None of these is a density claim. All three are placement claims.
What memory itself looks like
The suggestive part is that human semantic memory appears to have exactly this shape.
Steyvers & Tenenbaum (2005) analysed three large semantic networks — free-association norms, WordNet, and Roget’s Thesaurus — and found small-world structure in all three: sparse connectivity, short average path lengths, strong local clustering. They also found power-law degree distributions: most concepts with few connections, a small number of hubs with very many. Not a tree. Not a random graph. Not a hairball.
And retrieval from these networks looks like traversal. Abbott, Austerweil & Griffiths (2015) showed that a censored random walk over a network built from free-association data reproduces the clustering-and-switching pattern people show in category fluency — you name several birds, pause, switch to fish. This remains genuinely contested: Hills, Jones & Todd argue for strategic search in a high-dimensional space rather than a walk on a network, and the debate is live. But the relevant point survives either way: retrieval is modelled as movement through a structure, and the structure’s shape determines what movement is possible.
Most directly on point: Nematzadeh, Steyvers & Griffiths (2016) found that networks with small-world connectivity were particularly conducive to producing realistic retrieval patterns under simple random-walk search. Topology, not density, doing the work.
Three things I should concede
I build software for making knowledge graphs. An argument from me that knowledge graphs aid memory is structurally suspect, so let me put the strongest objections in myself.
The method of loci is a chain. The most powerful mnemonic technique ever documented is a linear walk through a sequence of locations. If chains are so bad, why does the best mnemonic in history use one? The answer, I think, is that the loci chain is a scaffold of retrieval cues imposed on unstructured material — arbitrary items with no intrinsic relationships to exploit. That is a different problem from representing a body of knowledge that has real internal structure. But it is a genuine complication, not a footnote.
The literature here is descriptive, not prescriptive. “Human semantic memory has small-world structure” does not entail “material presented as a small-world graph is learned better.” That inference is plausible and, as far as I know, not directly tested. The Nematzadeh result concerns model retrieval dynamics, not human learning outcomes. Bower is the only clean experimental result in this essay that directly compares structured against unstructured presentation, and it tested a hierarchy, not a web. The honest summary: the extremes are demonstrably bad, the tree is demonstrably good, and the claim that the small-world middle beats the tree is an inference from converging evidence rather than a finding.
Linear text still wins for arguments. Prose carries entailment, qualification, and rhetorical order. This essay is not a graph, and shouldn’t be. Graphs are for bodies of knowledge — things with many parts and many relations among them. They are not for arguments, which have one thread and depend on the reader following it.
Structure is necessary and not sufficient
One more thing, and it undercuts the whole essay if you leave it out.
A graph you look at is a picture. A graph you walk is a route system. Bower’s hierarchy worked because participants used it as a retrieval plan — they executed it. Everything we know about the testing effect and the generation effect points the same way: the act of traversal, of retrieving under mild difficulty, is what converts structure into memory. Studying a beautifully organized diagram and never navigating it is close to studying nothing.
So the design question is not only “what shape should the graph have” but “what does the shape make it natural to do.”
A practical synthesis
If the argument is right, it suggests something I have come to call the gradient rule, and it is the most useful thing I have taken from building a dozen of these:
Cross-link density should be inversely proportional to height in the hierarchy.
Near the root, keep it a tree. The top levels are a map, and a map with lateral edges everywhere is unreadable. One parent per node; lateral links only when strongly justified.
Deep in the graph, actively hunt for cross-links. A leaf with a single edge is a wasted learning surface — reachable one way and forgettable the same way. Every honest second link is a second route in.
The result is a structure that is hierarchical at the top and web-like at the bottom: order where you need orientation, chaos where you need association. It is not a compromise between the extremes. It is a different thing from either, and it is roughly the shape memory already has.
Further reading
- Bower, G. H., Clark, M. C., Lesgold, A. M., & Winzenz, D. (1969). Hierarchical retrieval schemes in recall of categorized word lists. Journal of Verbal Learning and Verbal Behavior, 8, 323–343.
- Anderson, J. R. (1974). Retrieval of propositional information from long-term memory. Cognitive Psychology, 6(4), 451–474.
- Anderson, J. R., & Reder, L. M. (1999). The fan effect: New results and new theories. Journal of Experimental Psychology: General, 128(2), 186–197.
- Watkins, O. C., & Watkins, M. J. (1975). Buildup of proactive inhibition as a cue-overload effect. Journal of Experimental Psychology: Human Learning and Memory, 1(4), 442–452.
- Collins, A. M., & Loftus, E. F. (1975). A spreading-activation theory of semantic processing. Psychological Review, 82(6), 407–428.
- Watts, D. J., & Strogatz, S. H. (1998). Collective dynamics of ‘small-world’ networks. Nature, 393, 440–442.
- Steyvers, M., & Tenenbaum, J. B. (2005). The large-scale structure of semantic networks. Cognitive Science, 29(1), 41–78.
- Abbott, J. T., Austerweil, J. L., & Griffiths, T. L. (2015). Random walks on semantic networks can resemble optimal foraging. Psychological Review, 122(3), 558–569.
- Hills, T. T., Jones, M. N., & Todd, P. M. (2012). Optimal foraging in semantic memory. Psychological Review, 119(2), 431–440.
- De Deyne, S., Navarro, D. J., Perfors, A., Brysbaert, M., & Storms, G. (2019). The “Small World of Words” English word association norms. Behavior Research Methods, 51, 987–1006.