CLAUDE.md had grown to 152 KB (~21k words, ~40k tokens) and is loaded into
every coding-agent session. The cost is not the cache read, it is attention:
the rules that are genuinely invariant were drowning in the mechanics of
subsystems that most tasks never touch.
The split criterion is blast radius, not importance. A rule a change anywhere
could violate stays in CLAUDE.md — the commit rule, the production/schema
constraint, domain neutrality, the event-bus rule, the crate boundaries, and
the module map. The mechanism of one subsystem moves to dev-docs/, opened on
entry to that subsystem via a routing table at the top of CLAUDE.md.
Nothing was rewritten: every section was moved verbatim by line range and
verified line-by-line against the original. The only edits are cross-reference
repairs ("see the DB section" -> a link), the promotion of headings in the
extracted files, and a condensed "Current state" whose full text now lives in
dev-docs/users-auth-and-boot.md.
CLAUDE.md: 152 KB -> 31 KB. Twelve subsystem files plus an index under
dev-docs/, which now carries the same standing rule as docs/ and CHANGELOG.md:
a change to a subsystem updates its dev-doc in the same change.
No CHANGELOG entry: this is documentation for coding agents with no observable
effect on the application.
7.3 KiB
Skald dev-docs — architectural reference for coding agents. Index: README.md · Entry point: ../CLAUDE.md
Read this when: you touch LLM clients, providers.yaml, retriability, request logging, token streaming or multimodal attachments.
The LLM stack
The client layer (crates/skald-core/src/llm/)
LLM client abstraction (OpenAI-compat, Anthropic, Ollama…). OpenAI-compatible provider types are runtime data, not code: providers/declared.rs loads providers.yaml at boot (see Config in ../CLAUDE.md); only non-OpenAI-compatible or bespoke providers (anthropic, ollama, openai, openrouter) stay native. Retriability (Model::is_retriable, agent-loop) keys on the real HTTP status carried by ModelError { status }, not a substring of the message — a model id/token count containing "404"/"401" cannot mis-classify; 401/403/404/422 don't retry, 400/429/5xx/network do. Request logging is the logging.rs::LoggingModel decorator, attached by the caller's ModelSelector (loop_adapters/selector.rs::SkaldSelector::with_log) — never by LlmManager, which builds one shared client per model and cannot know whose traffic it serves. The decorator's RequestLogTarget carries the owner: metadata → llm_requests in the registry (user_id, the column the UI filters on), payload bodies/headers → llm_request_payloads in that user's own encrypted DB, keyed by request_id; session + frame come from the request's own conversation/frame, so kernel rounds, sub-agent frames and compaction summaries are all attributed with no extra plumbing (ModelRequest::log is unused here)
Token streaming & reasoning display
The chat streams tokens live, as a parallel best-effort side-channel that never alters the turn's authoritative flow: the final Done (or Thinking) event still carries the complete content and the frontend treats it as truth.
- Client seam (
core-api::chatbot):ChatbotClient::chat_with_tools_raw_streaming(..., delta_tx: mpsc::Sender<StreamDelta>)— default impl ignores the channel and calls the bufferedchat_with_tools_raw, so providers without streaming (Ollama, LM Studio) are untouched.StreamDelta::{Text, Reasoning}splits visible answer from chain-of-thought. Senders usetry_send(deltas drop when the channel is full) — streaming must never backpressure the HTTP read. - SSE implementations (
crates/llm-client):OpenAiClient(stream:true+stream_options.include_usage,reasoning_content/reasoningdeltas, index-basedtool_callsaccumulation, usage from the final chunk) andAnthropicClient(stream:true;message_start/content_block_*/message_deltaevents;thinking_delta→ reasoning,input_json_delta→ tool input). Both reassemble the sameLlmTurn+LlmRawMetathe buffered path returns (the payload log stores a synthesized buffered-shaped body). Failure policy: if the stream dies before any delta the client retries buffered on the same model (providers rejectingstreamkeep working); a mid-stream failure propagates to the normal model-fallback logic. Framing is shared (llm_client::SseDecoder). Anthropic's buffered path now also parsesthinkingblocks intoreasoning_content(previously discarded). - Loop wiring:
call_llm_roundcreates the delta channel per attempt and a forwarder task maps deltas toServerEvent::TokenDelta { kind: content|reasoning, delta }on the turn's event channel (drained before the round's outcome events, so ordering holds); cancellation drops the in-flight future as before. A mid-stream fallback is handled client-side: the frontend clears its pending bubble onmodel_fallback. - Reasoning surfacing:
reasoning_contentridesDone/Thinkingevents (so buffered providers show it live too) and is projected asreasoningon assistant/thinking history items (build_items); persistence inchat_history.reasoning_contentand the echo back into context predate this feature. - Frontend (
chat-session.js+copilot-render.js, shared by desktop copilot and mobile chat-page):token_deltaaccumulates into a pending assistant bubble (in-place mutation + ~15 Hz flush, blinking caret);done/thinkingfinalize it in place,error/llm_failed/model_fallbackdrop it,tool_start/agent_donefinalize orphan bubbles (reasoning-only rounds, sub-agent final rounds that emit noDone). The reasoning block is a muted, collapsed-by-default native<details>(renderReasoning,.reasoning-blockincopilot-messages.css, i18n keychat.reasoning) — open state survives re-renders, and it renders identically from live events and from history.
Multimodal attachments
Uploads go through one centralized seam — ChatHub::save_upload (behind ChatHubApi::save_upload, backed by skald_core::uploads::save_to_home) — so every surface persists identically and no two callers can drift on placement (the class of bug where the agent was handed a path it couldn't reach). The seam writes into the caller's container home under uploads/{session_id}/ (agent path uploads/{session}/{name}, the UPLOADS_SUBDIR const in core-api/user_fs.rs), collision-dedupes the name, and prefers the sniffed magic-byte MIME over the client claim. The web handler (POST /api/{source}/uploads) buffers each field with a 256 MiB cap then calls the seam; the Telegram plugin downloads bytes then calls the same seam via handle.chat_hub().save_upload("telegram", …). Because the file lands in the home (bind-mounted at /root), it is reachable by the fs-tools, execute_cmd, and the file viewer (GET /api/file, per-user via resolve_view_path) — there is no /data static route anymore (removed: it was require_auth-only, not ownership-scoped, and also exposed internal server state under data/). Attachment metadata travels as structured JSON in chat_history.metadata — never as persisted text.
At context-build time (the crate's projection), attachments of the current turn (the user/agent rows following the last completed assistant reply, including across in-flight tool rounds) are partitioned by agent_loop::projection::media, with loop_adapters/media_source.rs deciding which files may be handed over (§6 containment): when the resolved model's LlmEntry.capabilities include the modality (vision → image_url parts, video → video_url parts), the file is inlined as a base64 data-URL content part — but only if it resolves (through the caller's UserFs, via resolve_host_path) under the home's uploads/ dir, its sniffed MIME is in the allowlist, and it fits the budgets (4 files / 10 MiB image / 32 MiB video / 48 MiB total per turn). Everything else — older turns, other kinds, any failed check — keeps the textual <system-extra> path block (built by core_api::message_meta::attachments_block / system_extra; the tag name is the single SYSTEM_EXTRA_TAG constant), so a non-vision model produces a byte-identical payload to before. OpenAiClient forwards parts verbatim; AnthropicClient translates image_url data URLs to image blocks (video unsupported; Anthropic models get vision by editing the model row's capabilities — no catalog refresh writes them). On LLM fallback mid-round, messages are rebuilt with the replacement model's capabilities.