How it works — pipeline v0

Curation,
as a pipeline.

Six steps turn a noisy web into a small, right, machine-readable library. Every step leaves an artifact in the repo — nothing lives only in someone's head.

01

Curate

An editor picks high-signal sources: foundational papers, canonical repos, and posts from first-party feeds. The seed list is code (ingest/seeds.py), so every pick is reviewable and reversible. Curation is the product.

02

Fetch — metadata only

Fetchers pull only metadata from public, unauthenticated endpoints: the arXiv API, the GitHub public API, and RSS feeds. No scraping behind ToS, no keys, no content copies. Each fetch records fetched_at and the endpoint it came from.

03

Describe

One short description per entry — an abstract snippet or a curated summary, ≤800 characters. Short summaries, attributed sources, no rewrites: hallucination control is a content policy, not just a model problem.

04

Tag and graph

Entries get taxonomy tags (24 topics, lowercase-kebab) and links to use-case nodes (graph/use_cases.json). The use-case graph is the spine: it routes agents from a job to a reading list, in order of editorial trust.

05

Embed and expose

Titles, descriptions, and tags are embedded locally so agents can ask in natural language. The machine surface is stable: search, get_resource, list_topics — plus llms.txt and sitemap.xml.

06

Refresh

fetched_at stamps every record; stale entries get flagged for editorial review. The library is a living index, not a frozen markdown list — freshness is a field, not a promise.

Guardrails / product-wide, hard

Index-only

Never copy or host third-party content. Link and describe. The record carries a license note; the URL is always canonical.

Sponsorship never outranks the graph

Sponsored entries are clearly marked metadata, ranked after editorial order. The graph is the product; the ad model is an experiment on top of it.

Agents see no ads

The machine surface is clean: three tools, structured records, no injected promotions. Disclosure fields are part of the record, not a wrapper.

No standing credentials

Ingest reads public, unauthenticated endpoints only. No API keys in files, no service accounts, nothing to rotate or leak.

Fail-closed on live writes

Nothing deploys publicly by default: no production pushes without an explicit owner gate.

Independent verification

Every artifact self-checks before it reports done: schemas validate, links resolve, pages lint. Trust, but verify — especially ourselves.

Sponsorship metadata / sketch, clearly marked

One schema, two readers

Sponsorship is data, not decoration: a disclosure block inside the record. Humans see a badge; agents see a field — and can filter it out entirely.

  • sponsor — who paid, in plain text
  • disclosure — required, machine-readable statement
  • priority — capped below editorial ranking

Revenue is an experiment, honestly framed: agent-discoverable ads are early and unproven. CPMs are not promised; the graph is not for sale.

sponsorship block — sketch
// inside a record, never around it
"sponsorship": {
  "sponsor": "Example Tooling Co.",
  "disclosure": "sponsored listing, paid placement",
  "priority": -1,   // never above editorial order
  "agent_visible": false  // ads are metadata, not messages
}