// Generated from the feature catalog
Thousands of Indexed Documents
Answers Grounded in Real Documentation · Shipped · workstation, home
Man pages, Arch Wiki, FreeBSD Handbook, Homebrew, and TLDR pages — all indexed locally.
The RAG corpus is built from 20+ specialized scrapers that acquire system administration documentation from diverse sources. Every document is verified, deduplicated, and indexed on disk. The full source list and licensing details are maintained in a dedicated data sources document.
Scrapers acquire content from man pages (Linux and macOS), Arch Wiki, Stack Exchange (Stack Overflow, Server Fault, Unix & Linux, Ask Different), FreeBSD Handbook, Homebrew formulae, TLDR pages, and vendor documentation. The index builder processes, chunks, and embeds each document for retrieval.
Limits and invariants
Section titled “Limits and invariants”No scraping of paywalled or copyrighted content without explicit license verification.
Where this lives
Section titled “Where this lives”halbert_core/halbert_core/rag/scrapers/halbert_core/halbert_core/rag/index_builder.pydocumentation/RAG-DATA-SOURCES-2026-08-24.md