18mu-blog
A personal knowledge site with Markdown as its only content source: an Astro static build served by a single Cloudflare Worker that also hosts the semantic map snapshot.
On this page[4]
What it is
18mu-blog is the codebase behind this site: a personal knowledge site whose only human-maintained source is Markdown.
It solves a single problem — making “write something” the only thing a person has to do. Content is written as Markdown, and pages, indexes, tags, RSS, the sitemap, and deployment are all produced by the build, so no list is ever maintained by hand.
The site carries four content lines: finished Chinese writing, public bilingual project write-ups, de-identified For Agent records, and a resume driven by structured facts.
How it is structured
The site has three layers, and each one depends only on the layer beneath it.
Content layer. Markdown is the only human-maintained public source. Blog, projects, and For Agent records are three content collections whose field contract lives in content.config.ts and is checked by a validation script before every build. Resume experience takes a different path: its facts live in fact-base/ and are read only at build time, never published as a collection.
Delivery layer. Astro builds the collections into a static site in .deploy/, and a single Cloudflare Worker does two jobs — it hands those assets to readers, and it answers the public API. The assets are bound to that Worker through the ASSETS binding, so the whole site is one deployable unit.
Semantic layer. Only published blog articles enter semantic processing. The Worker calls an external embedding service, computes pairwise cosine similarity, clusters by threshold, and writes each article’s nearest neighbours into Cloudflare KV. The knowledge map and random walk islands read the resulting snapshot from a public endpoint.
Design decisions worth noting
The content contract comes before the content. Each collection’s fields are defined by a schema, and a relation check runs before every build: a related field must point at a record that exists and is published. A wrong slug fails the build instead of leaving a dead link on a page.
Derived data can be rebuilt at any time. The semantic snapshot is derived entirely from published blog Markdown, with no hand-maintained relationships. Clearing the KV entry and re-running the sync restores it.
Secrets never enter content. The embedding endpoint and key exist only in Worker configuration and Cloudflare Secrets. The browser reads a public snapshot without credentials, and the sync token is used only inside GitHub Actions.
The release path is Git. Production only accepts a committed and verified main; uncommitted work in the working tree is never published.
What problems it solves
- Writing content never requires maintaining an index, tag page, or sitemap — the build produces them.
- Projects, experience, and For Agent records relate to each other by slug, and a mistake is caught during the build.
- Articles are connected by semantic similarity rather than hand-applied tags, so topic clusters form on their own.
- The site has no login, comments, payments, or browser-side credentials, so a content update only ever changes Markdown.