Curation,
as a pipeline.
Six steps turn a noisy web into a small, right, machine-readable library. Every step leaves an artifact in the repo — nothing lives only in someone's head.
Curate
An editor picks high-signal sources: foundational papers, canonical repos, and posts from
first-party feeds. The seed list is code (ingest/seeds.py), so every
pick is reviewable and reversible. Curation is the product.
Fetch — metadata only
Fetchers pull only metadata from public, unauthenticated endpoints: the arXiv API,
the GitHub public API, and RSS feeds. No scraping behind ToS, no keys, no content copies.
Each fetch records fetched_at and the endpoint it came from.
Describe
One short description per entry — an abstract snippet or a curated summary, ≤800 characters. Short summaries, attributed sources, no rewrites: hallucination control is a content policy, not just a model problem.
Tag and graph
Entries get taxonomy tags (24 topics, lowercase-kebab) and links to use-case nodes
(graph/use_cases.json). The use-case graph is the spine: it routes
agents from a job to a reading list, in order of editorial trust.
Embed and expose
Titles, descriptions, and tags are embedded locally so agents can ask in natural language.
The machine surface is stable: search,
get_resource, list_topics —
plus llms.txt and sitemap.xml.
Refresh
fetched_at stamps every record; stale entries get flagged for
editorial review. The library is a living index, not a frozen markdown list — freshness is a
field, not a promise.
Index-only
Never copy or host third-party content. Link and describe. The record carries a license note; the URL is always canonical.
Sponsorship never outranks the graph
Sponsored entries are clearly marked metadata, ranked after editorial order. The graph is the product; the ad model is an experiment on top of it.
Agents see no ads
The machine surface is clean: three tools, structured records, no injected promotions. Disclosure fields are part of the record, not a wrapper.
No standing credentials
Ingest reads public, unauthenticated endpoints only. No API keys in files, no service accounts, nothing to rotate or leak.
Fail-closed on live writes
Nothing deploys publicly by default: no production pushes without an explicit owner gate.
Independent verification
Every artifact self-checks before it reports done: schemas validate, links resolve, pages lint. Trust, but verify — especially ourselves.
One schema, two readers
Sponsorship is data, not decoration: a disclosure block inside the record. Humans see a badge; agents see a field — and can filter it out entirely.
sponsor— who paid, in plain textdisclosure— required, machine-readable statementpriority— capped below editorial ranking
Revenue is an experiment, honestly framed: agent-discoverable ads are early and unproven. CPMs are not promised; the graph is not for sale.
// inside a record, never around it "sponsorship": { "sponsor": "Example Tooling Co.", "disclosure": "sponsored listing, paid placement", "priority": -1, // never above editorial order "agent_visible": false // ads are metadata, not messages }