Files
Daniele 902f47ecd8
Nightly Build / build (push) Successful in 10s
docs: split CLAUDE.md into an always-loaded core plus dev-docs/
CLAUDE.md had grown to 152 KB (~21k words, ~40k tokens) and is loaded into
every coding-agent session. The cost is not the cache read, it is attention:
the rules that are genuinely invariant were drowning in the mechanics of
subsystems that most tasks never touch.

The split criterion is blast radius, not importance. A rule a change anywhere
could violate stays in CLAUDE.md — the commit rule, the production/schema
constraint, domain neutrality, the event-bus rule, the crate boundaries, and
the module map. The mechanism of one subsystem moves to dev-docs/, opened on
entry to that subsystem via a routing table at the top of CLAUDE.md.

Nothing was rewritten: every section was moved verbatim by line range and
verified line-by-line against the original. The only edits are cross-reference
repairs ("see the DB section" -> a link), the promotion of headings in the
extracted files, and a condensed "Current state" whose full text now lives in
dev-docs/users-auth-and-boot.md.

CLAUDE.md: 152 KB -> 31 KB. Twelve subsystem files plus an index under
dev-docs/, which now carries the same standing rule as docs/ and CHANGELOG.md:
a change to a subsystem updates its dev-doc in the same change.

No CHANGELOG entry: this is documentation for coding agents with no observable
effect on the application.
2026-08-24 18:04:43 +01:00

16 KiB

Skald dev-docs — architectural reference for coding agents. Index: README.md · Entry point: ../CLAUDE.md

Read this when: you touch any table, accessor under crates/skald-core/src/db/, or the memory/report stores.


DB tables (sqlx SQLite)

database/system.db — the path is a constant (core::db::SYSTEM_DB_PATH), not configurable. init_system_pool creates the directory; SQLite only creates the file. Per-user files are database/{userid}.db, created by UserManager::register_user and encrypted with SQLCipher.

The schema is split into two buckets (§5.1), and the split is the point:

  • create_registry_tables — instance-wide, readable without any user key: users, roles, llm_providers, llm_models, transcribe_models, tts_models, image_generate_models, plugins, plugin_access + plugin_user_configs, approval_rules, tool_permission_groups, config, known_tools, llm_requests, mcp_catalog, mcp_global_servers + mcp_global_access, oauth_providers, role_capabilities, shared_folders + shared_folder_members, projects + project_members, supervision, system_agent_coverage, system_agent_user_settings. The MCP tables back the Connectors model (§7/§14/§15 — see its own section); oauth_providers (accessor db/oauth_providers.rs) holds one row per identity provider (Google…) — endpoints + client_id/client_secret + redirect_uri, admin-owned household secrets (§4/§15b), never a per-user token. The last two pairs are junction-backed membership: shared_folder_members (accessor db/shared_folders.rs) for the on-disk shared folders (§6), project_members (accessor db/project_members.rs) for projects (see projects-and-files.md) — both let a member be read-only (can_write) and both drive the container mount topology + the fs routing. Their FKs are registry→registry (same file), which is allowed — unlike an owner→registry key.
  • create_owner_tables — one owner's content, identical schema in every file that has it: chat_sessions, chat_sessions_stack, chat_history, chat_llm_tools, chat_summaries, session_scratchpad, session_mcp_grants, stack_mcp_grants, scheduled_jobs, job_runs, system_agent_runs, system_agent_state, mcp_user_servers, mcp_events, sources, secrets, user_config, llm_request_payloads, memory_docs (+ FTS5 memory_docs_fts), reports. user_config is the per-user twin of the registry config table and deliberately does not share its name: the two hold different namespaces (instance settings the admin owns vs. one member's own preferences, the notification home being the first), and a same-named table in both files would turn a wrong-pool call into a silent read of the other scope — instead of the "no such table: config" that revealed /sethome writing owner state through db::config against a {userid}.db, which also had the notification consumer dropping every batch it ever built. mcp_user_servers (a user's activated per-user connectors) carries catalog_name as a bare TEXT snapshot of mcp_catalog.name, never a FK — an owner→registry key would fail every INSERT; for an OAuth connector it also snapshots oauth_provider + deliver_json, and its api_key column holds the refresh token (in the SQLCipher-encrypted file, so no column crypto). Because memory_docs is an owner table, one definition backs private memory in each {userid}.db and shared memory in system.db (the household owner) — see the memory namespace note below. (projects/project_tickets were owner tables in the single-user past: projects are shareable now, so projects + project_members are registry tables and project_tickets is gone.)

The schema is no longer greenfield (see the production note in ../CLAUDE.md): a full recreate is not an option anymore. db::ensure_columnALTER TABLE … ADD COLUMN swallowing the "duplicate column" error, a no-op on a fresh DB where the CREATE TABLE already carries it — is therefore not a convenience for dev boxes anymore but the only change shape that is currently safe, and additive-with-a-default is the shape to design towards. Used for the OAuth columns on mcp_catalog / mcp_user_servers. Anything destructive waits for real versioning.

No foreign key in the owner bucket may point at a registry table. SQLite cannot enforce a key across files, not even through ATTACH, and sqlx turns on PRAGMA foreign_keys: the CREATE TABLE succeeds and every INSERT fails. db::tests::owner_tables_stand_alone_with_foreign_keys_on enforces this by running the owner schema against a database holding nothing else, then inserting a row into each table. One key crossed and was fixed: chat_history.model_db_id (dropped — write-only, and llm_requests.model_name already records the model).

Memory namespace (blueprint §5). memory_docs (accessor db/memory_docs.rsget/upsert/list/search(FTS)/delete) backs a virtual note store surfaced through the fs-tools, not the disk. Two sibling roots (not the blueprint's nested memory/{userid} + memory/shared): user-memory/… routes to the caller's own pool (ToolContext::pool), shared-memory/… to the system pool (a singleton captured in fs::register_all). tools/fs/classify_memory() decides on the raw first path component (a .. in the tail clamps inside the store, never escapes to disk); read_file/write_file/list_files/edit_file/insert_at_line/replace_lines/search_file override run_with to route memory paths (each extracting a pure transform shared with its on-disk execute) and leave every other path on disk. The HTTP surface routes them the same way: GET /api/file classifies before resolve_view_path and serves the note from memory_docs (caller's pool / system pool), so the file viewer opens user-memory/… and shared-memory/… like any file, and show_file_to_user accepts memory paths too (existence-checked on the right pool). Approval (seeded in seed_fs_path_rules): user-memory/* is @fs_any allow (private, frictionless); shared-memory/* is @fs_read allow + @fs_write require — reads free, writes need approval so the agent can't silently push one person's data into shared memory. grep_files stays disk-only (regex-across-tree ≠ FTS); ranked full-text recall over notes is a separate tool, memory_search (tools/fs/memory_search.rs), over the memory_docs FTS index — allowed by a path-less rule (it takes query, not path).

Supervision + coverage (registry). supervision(subject_user_id, supervisor_user_id) (accessor db/supervision.rs) is the §0.1 supervision edge — a generic directed edge between two users, deliberately attribute-free, whose domain reading ("a parent watches a child") lives only in seed data and UI copy. It answers two questions with one table: whom does a background agent look at (subjects()) and who may read what it produced (supervisors_of(), which is what reports.audience = 'supervisors' resolves against). Both FKs are registry→registry, so the cascade is real in both directions. system_agent_coverage(agent_id, subject_user_id, covered_through) (accessor db/system_agent_coverage.rs) is the per-subject watermark that makes "everything since last time" a window: it sits between system_agent_runs (a history for the human, skips idle passes) and system_agent_state (attempt marker, advances on every tick and before the work — which is precisely why it can never delimit the window the work is about), and differs from both by advancing only on a completed pass, so a crash re-covers rather than skips. Deriving it from the last report's period_end was the obvious alternative and is wrong for one ordinary reason: a supervisor deleting an old report would rewind the scheduler and regenerate the report they just discarded — a document is the user's to delete, scheduler state is not. Registry rather than owner because the pass runs in some supervisor's runtime and which one depends on who is logged in that night; the acting user's file would give one subject two unsynchronised clocks.

Reports (db/reports.rs, blueprint §13). The documents system agents write about a stretch of time — a daily review of a supervised account, a weekly "what you struggled to get done" digest. The second two-homes table, for the same reason as memory_docs and with the same mechanics: one owner schema, and the file a row lands in is its audience. A {userid}.db row is that user's own report, behind SQLCipher; a system.db row is an instance report, written about someone for the people who supervise them and therefore cleartext to whoever owns the box — deliberately, since they are the intended reader (§2). Which file a producer writes into falls out of its own AgentScope with no new concept (PerUserctx.pool, Instance → the registry pool it already holds), and the subject of an instance report cannot see it because their tools only ever reach their own pool — the invisibility is structural, so nothing anywhere filters by reader. subject_user_id/producer_user_id/run_id are bare snapshot columns, never FKs (owner→registry would fail every INSERT; for an instance row the system_agent_runs trace sits in the acting user's file). kind is producer-declared text, not an enum (§0.1). Rows are immutable but for mark_read, whose read_at IS NULL guard makes acknowledgement shared and first-reader-wins — two admins, one alert, dealt with once. Consequence worth internalising: since the admin cannot open the subject's encrypted sessions, there is no click-through to the evidence — whatever justifies a report must be narrated in its body, under the same rule the shared memory lint already follows (say which conversation and what kind of problem, without reproducing the sensitive line). Currently there is no producer, no API and no UI — the table, its accessor and its tests are the whole of it.

Memory injection into the prompt: AgentSystemContext::load_inject_memory (loop_adapters/system.rs) routes each meta.inject_memory entry — user-memory/… → owner pool, shared-memory/… → the shared (system.db) pool, both via memory_docs::get; anything else (data/…, $WD/…) is a disk read. The shared pool is threaded ChatSessionManagerUserLoopRuntimeAgentSystemContext. assistant and project-coordinator inject user-memory/index.md + shared-memory/index.md.

Prompt substitutions: an AGENT.md may carry <!-- KEY --> placeholders; agents::resolve_includes turns each into a __KEY__ sentinel, replaced at request time. Several are resolved by the system-context source itself (loop_adapters/system.rs) from the session owner (user_id) + registry (shared_pool) + their UserFs, so every source (WS, mobile, cron, sub-agents) gets them with no caller plumbing: __SKILLS_LIST__ (the generated skills index — see Filesystem & containers), __SANDBOX_COMMANDS__ (the sandbox command hint — see below), __SHARED_FOLDERS__ (the user's shared-folders table) and __USER_PROFILE__ (the owner's directory profile: Name, Date of birth with age computed at build time, Sex, Preferred language, admin Notes — unset values render as explicit unknown / not specified, the Notes line is omitted when empty). Any other key comes from the per-call SendMessageOptions::system_substitutions map.

system.db still gets both bucket functions — but no longer because the migration is unstarted. It gets the owner schema because it is the owner of shared memory (memory_docs) plus, for now, the globally-scoped secrets (SecretsStore is built on the system pool and shared by reference into every UserContext; the global runtime's config now lives in the registry table mcp_global_servers, and per-user connector config in each user's owner mcp_user_servers). The global runtime no longer writes mcp_events there: notification persistence is an explicit McpManager::new argument (EventLog::{Persist,Discard}), Discard for the ownerless global runtime and Persist for each per-user one, because an event belongs to whoever it happened to and its only reader (event triage) is per-user. Every other owner table is created there but never written to anymore — the global owner-bound managers that would write them (chat/jobs/etc.) are inert (see users-auth-and-boot.md). Fully dropping create_owner_tables from system.db is blocked on the §4 scope decision for secrets, not on call-site migration.

users (crates/skald-core/src/db/users.rs) holds the directory plus auth material. It lives in the system DB, which the box owner can read, so it must never store anything that derives a user's key. Credentials is an enum mirroring the table's CHECK: an encrypted user carries a wrapped DEK (whose AEAD tag is the password verifier — hence no password_hash); a cleartext user carries an ordinary verifier, or none. User is deliberately not Serialize and its Debug redacts key material — use User::summary() for anything leaving the process. role_id references roles(id) (the roles table is now seeded before users in create_registry_tables). A nullable locale column (additive via ensure_column) holds the per-user UI language override; role-driven conventions live in the free-form roles.attrs JSON — never new columns per attribute — parsed at a single point by the typed db::roles::RoleAttrs (ui_mode, permission_groups, chat_agent, auto_grant — the last one being why that struct's Default is hand-written, see default-access.md): ui_mode (see frontend.md) plus the role's security-group set (roles.permission_group = the default group, attrs.permission_groups = additional allowed groups; Role::effective_groups() = the union, roles::role_allows_group() gates it with admin short-circuiting to all). See the security-group picker in the frontend section. The role's default entry (chat) agent is attrs.chat_agent — the neutral chat-type agent members of the role land on (§0.1: data, not an enum). Resolved by roles::default_chat_agent_for_user(registry_pool, user_id) — the single seam behind both the per-user ChatHub's default_agent (snapshotted at login in UserContextFactory::build, like fs/MCP access, so every session-creation path — explicit provision_session, lazy WS get_or_create_session, notify — honors it) and provisioning_for_source's non-project branch. Falls back to agents::DEFAULT_CHAT_AGENT ("assistant", the renamed former main) when unset. Seeded: admin/memberassistant, childrenkid (Companion). A per-user override is future work, layering on top in the same resolver. The stack root frame is created with the session's own agent_id (not a literal) — config.agent_id (from the frame) drives which prompt runs, so a wrong id there silently runs the wrong agent. The admin-managed directory profile lives in three more additive columns — birthdate (ISO YYYY-MM-DD), sex (free text), notes (admin-authored) — edited only from the Users admin page (set_directory_fields; validation — real non-future date, length caps — lives in the users_mgmt API, not the db layer) and rendered into agent prompts by the __USER_PROFILE__ substitution (see above). They are directory metadata written by the admin about the user, so the registry is their honest home under the §2 threat model.