The knowledge library
Every source behind this research is a page in a browsable, searchable library, with a graph view showing how the pieces connect. Like the Engine it is a working beta, and a public version is planned. The rest of this page explains how it was built.
A working implementation of an emerging pattern
This knowledge base did not start as an original idea about knowledge. It is a practical build of a pattern with a long lineage, taken further where the research needed it.
A personal, curated store of documents with associative trails between them. The connections mattered as much as the documents.
An LLM incrementally builds and maintains the wiki, so the connections stay current. You curate and question; it does the bookkeeping.
The same idea, plus a verification regime it lacked, turned from a personal tool into a public research artefact.
More than five hundred sources, filed by kind
The corpus is filed by what kind of thing each source is, not by topic. That sounds bureaucratic and turned out to matter: a district plan and a peer-reviewed paper carry very different weight, and the file type is the first signal of how far to trust a claim.
| Type | What it holds | How it is weighted |
|---|---|---|
lit | Academic and conference papers | Highest. Cited directly |
reg | District plans, standards, regulations | Authoritative for what is permitted |
ds / rd | Datasets and technical datasheets | Primary for figures |
cr | Compiled research, including AI synthesis | Never cited alone; must resolve to a primary |
url | Web captures | Supporting, dated at capture |
mm | Video and audio, with transcripts | Supporting |
ot | Everything else | Case by case |
14 weeks, 544 source pages
Every source page in the corpus, by the week it was added. The line is the running total. Both are read out of the project's version history rather than a tally kept by hand, so the shape is what actually happened, including the quiet weeks.
The corpus holds 544 source pages, of which 535 are published to the library above. The busiest week added 138, in the fortnight the method changed from retrieving sources one at a time to running retrieval and verification in parallel. Counting starts 18 May 2026.
From a file to a trusted claim
Following the pattern above, an LLM does the bookkeeping and a human does the curating. Every source becomes a typed page with structured metadata, linked to the concepts and claims it supports. Automated checks run over the whole corpus and refuse to pass if a page is missing required fields, or if a raw file exists with no page pointing at it.
Filed, typed, duplicate-checked by content hash.
A separate pass, so the writer is not also the checker.
Tied to the concepts and claims it supports.
Whole-corpus checks; no orphan files, no missing fields.
It was not designed. It was forced by an error.
Karpathy's pattern trusts the LLM to summarise faithfully. Ours learned not to. Partway through, AI assistants were caught attributing figures to the right organisation but the wrong document. Plausible, well-formatted, and wrong. The response was a verified-source-only rule: a figure may be used only if the underlying primary document has been retrieved and the exact number confirmed in its text. Not the summary. The document.
A knowledge centre of two kinds of page
It is a knowledge centre, not a document store. It is built from two kinds of page. Every ingested document gets its own page: a summary, its metadata, and the claims it supports, with the source file itself referenced rather than hosted. On top of those sit synthesis pages, which draw many sources together into a topic, renewable heating, food production, the self-sufficiency method, and connect the concepts that matter to the project.
Kerr (2024), community energy in Aotearoa
Micro-hydro
Sources for the method
Karpathy, A. (2026). LLM Wiki. gist.github.com/karpathy.
Bush, V. (1945). As We May Think. The Atlantic.