Start from the job.
Agents don't browse — they arrive with a job. Nine use-case nodes route the taxonomy into paths: each one lists the curated entries that answer it. Click a row to expand.
The canon: agent-computer interfaces, SWE-bench as the yardstick, and the frameworks that survived contact with real repos.
- arxiv:2310.06770 — SWE-bench: can LLMs resolve real GitHub issues?
- github:princeton-nlp/SWE-agent — SWE-agent — the agent-computer interface
- github:aider-AI/aider — aider — pair programming in the terminal
- github:AllHandsAI/OpenHands — OpenHands — open platform for software agents
- arxiv:2305.16291 — Voyager — lifelong-learning agent with a skill library
From Toolformer's learned API calls to the Model Context Protocol's standard interface — typed, callable surfaces.
Roles, conversations, and SOPs: when one agent is not enough.
Benchmarks, harnesses, tracing: measurement before vibes.
- arxiv:2310.06770 — SWE-bench — real-issue evaluation
- arxiv:2308.03688 — AgentBench — evaluating LLMs as agents
- github:openai/evals — openai/evals — canonical eval harness examples
- github:Arize-ai/phoenix — Phoenix — LLM observability and eval tracing
- github:langfuse/langfuse — Langfuse — traces and evals
Paged context, reflection, hippocampus-inspired retrieval — the whole memory stack.
RAG from first principles: the original paper, the lost-in-the-middle caveat, and the engines.
Chains, trees, and search over thoughts — deliberate problem solving.
Guaranteed structure: grammar constraints, validated schemas, compiled pipelines.
Local and high-throughput inference, and the library underneath.
Graph, not keywords. Use-case nodes live in
graph/use_cases.json and every entry links to the nodes it
supports (use_cases field, schema v1). Sponsorship never
outranks this ordering.
From use-case to citation in one call.
Each entry behind these rows ships as a full record — the machine view your agent cites. See one end-to-end.