Skip to content

// Generated from the feature catalog

Thousands of Indexed Documents

Answers Grounded in Real Documentation · Shipped · workstation, home

Man pages, Arch Wiki, FreeBSD Handbook, Homebrew, and TLDR pages — all indexed locally.

The RAG corpus is built from 20+ specialized scrapers that acquire system administration documentation from diverse sources. Every document is verified, deduplicated, and indexed on disk. The full source list and licensing details are maintained in a dedicated data sources document.

Scrapers acquire content from man pages (Linux and macOS), Arch Wiki, Stack Exchange (Stack Overflow, Server Fault, Unix & Linux, Ask Different), FreeBSD Handbook, Homebrew formulae, TLDR pages, and vendor documentation. The index builder processes, chunks, and embeds each document for retrieval.

No scraping of paywalled or copyrighted content without explicit license verification.

  • halbert_core/halbert_core/rag/scrapers/
  • halbert_core/halbert_core/rag/index_builder.py
  • documentation/RAG-DATA-SOURCES-2026-08-24.md