Files
Skald-Circle/CLAUDE.md
T
dguiducci b198ac923b
Nightly Build / build (push) Successful in 6m57s
connectors: merge the catalog page into a row-list Connectors page
The standalone Catalog page added nothing the Connectors page could not
do: drop it (component, route, sidebar entry) and move its affordances
onto the Connectors page — the Add-connector dropdown (marketplace /
manual form, now at #connectors/new) and per-row removal for the admin.

Replace the card grid with a sharper row list (4px radius) built for
scanning status and acting; the marketplace's back link now returns to
#connectors. i18n keys renamed catalog.* -> connectors.add/new/*.
2026-07-27 21:01:32 +01:00

86 KiB
Raw Blame History

Skald (project-family) — codebase guide

Rust async web app (Tokio + Axum). Runs as a local chat server with LLM tool-calling and a sub-agent system.

Never git commit unless explicitly asked. Staging, building, running and testing are fine on your own initiative; creating a commit is not. Do the work, leave it in the working tree, and let the user commit — or ask them to — even when a commit looks like the obvious next step.

Commit messages must be in English.

What this repository is

A dedicated fork of Skald, turning a single-user personal agent into a multi-user assistant for a small trusted group — positioned at families, but see the neutrality rule below.

The design lives in blueprint/project-family.md. Read it before any architectural work; its sections are referenced by number (§0.1 neutrality, §5.1 database layout, §11 UserManager, §12 auth schema, §16 LLM privacy tiers, §17 sequencing). The blueprint/ directory is gitignored and not under version control — treat it as the source of truth, and never assume a section says what you remember.

Load-bearing decisions from that document:

  • Not upstreamable. Nothing here needs to preserve Skald's schema or be portable back to it.
  • Greenfield. No users in production ⇒ no migrations, no backwards compatibility. Tables get restructured, renamed and moved freely; the schema collapses into a single clean baseline v1.
  • Dual memory: a private per-user pool plus a shared pool. A user's private space is encrypted so that nobody else — the admin included — can read it through normal use of the system. Never claim "mathematically impossible": the honest promise is transparency plus verifiability (§3).
  • Threat model (§2): the adversary is the tempted admin, who owns the box but does not recompile the binary or dump RAM. Do not design against a forensic attacker.
  • Roles are data, not enums (§0.1): a roles table binds permission-group, run-context and data-handling attributes. "Children" is a seeded preset row, never a hardcoded type.

Event-driven coupling — think in events, not calls

Three global broadcast buses — never add a fourth without checking these first:

Bus Cap Events File
ChatEventBus 256 user message, assistant response, compaction done core-api/src/bus.rs
SystemEventBus 64 provider (un)registered, config key updated, job completed, session cancelled, user created/deleted/active-changed/mounts-changed, global connectors changed, connector reinstalled core-api/src/system_bus.rs
GlobalEvent (per-user) 512 all ServerEvent variants → WS clients + inbox lifecycle core-api/src/events.rs

Plus internal mpsc queues: per-source SourceInbox (message serialization) and a central notify queue (background agents → user).

The user-lifecycle reconciler is the worked example of the rule. Creating a user, deleting one, deactivating one, or changing a shared-folder/project membership all need Docker work (provision, tear down, stop, recreate with new bind mounts); enabling or reinstalling a connector needs live runtimes re-snapshotted. None of the endpoints that make those changes touches ContainerManager or the refresh helpers: each announces SystemEvent::User{Created,Deleted,ActiveChanged,MountsChanged} / McpGlobalServersChanged / ConnectorReinstalled after its DB write, and one subscriber — skald::wiring::spawn_user_lifecycle, spawned post-construction because it reacts through Skald's own accessors, holding only a Weak — does the reacting, sequentially and best-effort. Being off the response path matters for ConnectorReinstalled in particular: it re-copies files and restarts servers inside every live user's container, seconds of work the admin's install no longer waits on. The payoff is that a future endpoint granting membership cannot forget to remount, because remounting was never its job. Reactions never block the HTTP response, and a failure settles at the user's next login or at boot reconciliation.

Where the bus stops: reconciliation rides it, authorization does not. SystemEventBus is a lossy 64-slot broadcast whose contract is "best-effort, settles at the next login" — right for a stale mount, wrong for a revocation, where "settles later" is the failure. So deactivating or deleting a user splits in two: Skald::revoke_user_runtime runs synchronously in the handler, before it responds (revoke every session → evict + cancel the UserContextUserManager::lock, in that order, so nothing is left querying a pool we then close and the DEK leaves RAM per §9), while only the container half — stop or remove — rides the bus. Before this, active = 0 blocked the next login but left live sessions working: login checks the flag, require_auth only maps token → id. Same split for security groups (see the picker section) and for connectors, where the test is worth internalising because the call is literally the same function: Skald::refresh_global_mcp_access is announced (McpGlobalServersChanged) when a global connector is enabled or deleted — the first only makes something appear, the second is already enforced by stop_server — but called directly from global_set_access and user_connectors_set, where set_access/set_for_user replace a grant set and the refresh is what actually revokes. Both sync call-sites carry a DELIBERATELY SYNCHRONOUS comment, because they look identical to the announced ones. Never put an access revocation on a bus.

Before you add a direct function call or a new import between two components, stop and ask: is one component producing data another needs? If yes, add a variant to an existing bus and spawn a subscriber. Don't call some_manager.log_thing(...) from the producer — emit a ThingHappened event on SystemEventBus and let the manager subscribe.

A new mpsc::channel or broadcast::channel is a code-review flag. Nine times out of ten you want one of the three buses above. If you truly need a new one, be ready to explain why none of the existing three fits.

The core is domain-neutral — this is a hard rule

"Family" is positioning, not architecture. Schema, engine, API, identifiers and comments must never contain family, household, parent, child or minor. A pivot to teams, small orgs or care settings must not require renaming anything.

Domain concept Technical primitive
the group implicit — it is the instance. No group entity. Future multi-group ⇒ tenant / workspace, never family
shared memory memory/shared
parent / admin role admin
child / minor a data-driven role defined by the admin
"the parent reads the child's data" a generic supervision edge between users

Domain words are allowed only in seed data, preset labels, UI copy and positioning.

Current state

UserManager (§11) is now consumed. Login exists (crates/skald-core/src/auth/mod.rs: SessionStorelogin/user_of/logout plus revoke_user, the admin-side "drop every session of this user" used by Skald::revoke_user_runtime; the deny-by-default middleware is src/frontend/api/guard.rs, whose require_auth maps token → id and does not re-read the row, which is exactly why revocation must be pushed rather than polled; first admin created by skald-setup), and the per-user owner-bound runtime is UserContext (crates/skald-core/src/skald/user_context.rs) — resolved by Skald::user_context / the frontend's require_context, keyed off UserManager::pool_of, and carrying its own CancellationToken (a child of the instance one) so a single user's cron/hub/MCP loops can be stopped without touching anyone else's. The frontend owner call-sites (WS, sessions, inbox, approval-pending, projects, uploads, run-context, cron) route through the per-user pool; dev/stats read llm_requests — a registry table — from system.db, which is correct. The "owner-without-a-user" question resolved to there isn't one: every owner content belongs to a logged-in user (the admin included). The global owner-bound bundles (Conversation/Tasks: the "ownerless" ChatSessionManager, ChatHub, cron TaskManager) are still constructed but inert — their loops never spawn and nothing consumes their accessors; removing them is pending follow-on work (kept for now because RunContextManager shares the Conversation bundle and is used, being registry-backed). See blueprint §19.

Direction of travel, decided but not yet executed: strip the power-user surface (self-rewriting, arbitrary shell, dev-agent suite, ticket system) and move to a binary-first layout — the app is built once and run from a compiled binary, not executed from its own source tree.

Workspace layout

The application core is the skald-core crate; the binaries are shells around it.

Crate Role
crates/skald-core/ Storage, identity, crypto, LLM stack, tools, MCP, sessions. Knows nothing about what runs it: no HTTP server and no concrete plugin cratePluginManager only ever sees Arc<dyn Plugin> from core-api
skald (root, src/) The server shell: main.rs, the Axum frontend/, config.rs. Constructs the plugin list and hands it to Skald::new. Runs headless as a background daemon under the run.sh supervisor
crates/skald-setup/ Guided first-run setup — a terminal shell over skald-core. Creates the first admin and seeds the instance through the shared seam skald_core::setup::initialize_instance (apply the chosen seed profile → register_user(admin) → set default locale) — the same function the web setup calls, so the two shells can't drift. Asks profile, interface language, whether to encrypt — default yes — and password. A separate binary so the server never links TTY-prompt deps, and so a future GUI installer is a third shell over the same seam. run.sh runs it before the server loop; it prompts only when users is empty and stdin is a terminal, otherwise a no-op. --check reports readiness by exit code (0 done, 1 needed)
crates/core-api/ The contracts both sides share: Plugin, Tool, event buses, provider types

Two rules keep the boundary real, and both are enforced by the compiler:

  • The core never names a plugin. A plugin contributes tools through Plugin::tools(self: Arc<Self>) — the sibling of http_router() — so nothing in the core has to downcast to a concrete type. Naming one would drag every plugin in the tree into the core, including a C build via plugin-transcribe-whisper-local.
  • The core never learns about the process shell. There is no in-core restart hook — the former restart tool and its tools::restart::set_restart_handler seam were removed. The only coupling to the supervisor is now the run.sh exit-code protocol (exit 255 ⇒ re-exec the same binary by path), a seam no code currently triggers (kept for a future admin-driven restart). The live expression of this principle is skald_core::boot, which emits startup lines each shell renders (src/boot_format.rs here).

Plugin visibility & per-user config. The admin surface is split in two: #plugin-catalog (plugin-catalog.js) is a status board — one card per plugin with an enable toggle + health dot + a Configure button — and #plugin-detail?id=<id> (plugin-detail.js) holds the instance-config form + per-user access checklist for one plugin (the plugin counterpart of connector-detail.js). The user-facing half is #plugins (plugins-page.js): granted plugins + their per-user config forms. Enable/disable + instance config + access grants are gated by the plugin.manage capability (admin-only by construction). Visibility is opt-in: a row in plugin_access(plugin_id, user_id) grants a user sight of an enabled plugin (plugin_id is bare TEXT, never a FK — a plugins row exists only after the first toggle). A plugin with a non-empty Plugin::user_config_schema() exposes per-user settings, stored in plugin_user_configs (admin-readable system.db — never secrets) and applied through the Plugin::update_user_config hook, whose default just stores the blob via the PluginUserConfigApi on PluginContext.user_config. Telegram is the reference impl: the user pastes the bot's pairing code in their Plugins page, the override turns it into a chat_id → user_id binding (same write path as the telegram_pairing tool) and stores a {linked, chat_id} status blob for the UI. Endpoints: admin GET/PUT /api/plugins[/{id}] + GET/PUT /api/plugins/{id}/access; user GET /api/plugins/mine + PUT /api/plugins/{id}/my-config.

Plugin HTTP routes & web pages. Every plugin's http_router() mounts at boot under /api/plugin/<id>/enabled or not: two shared gates wrap each router (require_auth, then guard::plugin_enabled_gate, which re-checks the DB flag per request and answers 404 while disabled), so enable/disable serves/stops routes immediately with no restart, and plugin responses carry Cache-Control: no-cache. The router contract: cheap and safe to build pre-start, handlers tolerant of the not-running state (resolve runtime state per request through a shared cell, as mobile-connector does). A plugin may also contribute frontend pages via Plugin::web_pages() (PluginPage { page_id, title, icon, entry, admin_only, priority }): GET /api/plugins/pages returns the caller's visible pages (admin: all; others: non-admin_only pages of granted, enabled plugins) with entry_url resolved, and the sidebar renders them as menu entries routed #plugin/<plugin_id>/<page_id>. A single <plugin-page-host> (web/components/plugin-page-host.js) dynamic-imports the fragment ES module the plugin serves from its own router, registers its default-exported HTMLElement class, and mounts it with the plugin-id attribute — the fragment talks to its backend only through /api/plugin/<id>/… and runs with full session privileges (plugins are trusted: they ship in the binary). The frontend knows nothing about plugin page contents or behavior.

skald_core::boot emits curated startup lines on the boot tracing target; each shell decides how to render them (src/boot_format.rs here). The core says what happened, never how it looks.

Key modules

Path Role
src/main.rs Thin entry point: tracing → Skald::newWebFrontend::start → shutdown. Builds a tokio runtime and blocks on async_main, which runs the backend until a SIGINT/SIGTERM. Exposes run_backend() / shutdown_backend()
crates/skald-core/src/skald/ Skald — headless application core. mod.rs (struct + staged new() / shutdown()), runtime.rs (cross-cutting Runtime context), bundles.rs (8 domain bundles + build()), wiring.rs (wire() + spawn_background()), supervisor.rs (TaskSupervisor), accessors.rs (per-manager accessor facade — the API surface the frontend uses)
crates/agent-loop/ The LLM loop itself, as a standalone crate: kernel (round loop, fallback, tool fan-out), LoopManager, HistoryStore, projection (history→wire), DelegateTool (sub-agents), recovery.rs (restart), compaction.rs, plus the shipped model clients (models/). Knows nothing about Skald — see the loop section below
crates/skald-core/src/loop_adapters/ Skald's side of that crate's traits: history store, model selector, approval gate, tool set + bridges, agent catalog, event translator, projection knobs, async executor. This is where "how Skald does it" lives
crates/skald-core/src/session/handler/ What is left of the session layer: mod.rs (ChatSessionHandler + handle_message), kernel_turn.rs (the three loop entry points), config.rs, interface_tools.rs, media.rs
crates/skald-core/src/session/manager.rs Creates/retrieves ChatSessionHandler per session
crates/skald-core/src/chat_hub/ ChatHub: broadcast events to all connected WS clients
crates/skald-core/src/chat_event_bus.rs Global async bus for cross-session events
crates/skald-core/src/agents.rs Discovers agents from agents/*/, loads meta + system prompt
crates/skald-core/src/tools/ Built-in tools: exec (runs inside the caller's per-user Docker container via docker exec, as the non-root host uid — sudo for system installs — with a robust /stop that reaps the command's process-group; see container/; the only live path is run_with (needs ToolContext) — the context-free Tool::execute/execute_async now error (HOST_PATH_ERROR) instead of the old host sh -c, so nothing can run a command outside the sandbox), list_agents, fs/* (route user-memory//shared-memory/ to memory_docs, and every other physical path through ctx.fs to the caller's per-user host workspace — see DB tables + container), notify, ast_outline, image_generate, MCP tools, plugin tools, cron tools
crates/skald-core/src/container/ ContainerManager (§6): per-user Docker containers (the execution sandbox). Docker is a hard requirementcheck_docker() fails Skald::new (→ shell exits) if the daemon is unreachable. Builds our own skald-runtime image (python+node+sudo; tag is versioned skald-runtime:v2 so a Dockerfile change forces a rebuild) once from the embedded Dockerfile, then reconcile_all() at boot ensures one running container skald-{userid} per active user. Each container runs as the host uid:gid (--user, §6 UID coherence) with --init (tini reaps zombies); ensure() self-heals a container whose --user is stale (e.g. an old root one) by recreating it, and injects a passwd/shadow entry post-create so sudo (NOPASSWD, in the image) resolves the arbitrary uid. build_user_fs() assembles a user's UserFs (home {WD}/homes/{userid}/root, plus each shared/{name} they belong to). Shells the docker CLI (no client crate)
crates/skald-core/src/tool_catalog.rs ToolCatalog: unified tool listing façade (wraps ToolRegistry + McpManager)
crates/skald-core/src/events.rs ServerEvent enum streamed over WebSocket to the frontend
crates/skald-core/src/db/ sqlx SQLite — see below
crates/skald-core/src/users/ UserManager (§11): user directory CRUD on system.db, credential check, and the map userid → SqlitePool of unlocked databases. The pool is the unlock token — its connect options carry the DEK as SQLCipher's raw key, so an open pool means the key is in RAM (§9) and dropping it re-locks. Knows nothing about cookies: whatever maps an HTTP session to a user id sits above it
crates/skald-core/src/crypto/ Envelope encryption (§4/§5.1). A random 256-bit DEK encrypts {userid}.db; users.database_password holds it sealed with AES-256-GCM under Argon2id(password, salt). The AEAD tag is the password verifier — one derivation both authenticates and yields the key, and no second hash sits in the admin-readable DB. Cleartext users store the Argon2id output directly, compared constant-time. Argon2 runs in spawn_blocking behind a 2-permit semaphore (256 MiB per derivation)
src/config.rs Loads config.yml; LLM clients, strength, data root. All relative paths (db, logs, data, …) resolve against the launch cwd
crates/skald-core/src/mcp/ MCP runtimes + the McpProvider seam (§7): the shared host global runtime and the per-user container runtimes, unioned per session as UserMcpView. See the MCP connectors section
crates/skald-core/src/plugin/ Plugin system: discovery, enable/disable, tool registration, per-user access grants + per-user config
crates/skald-core/src/cron/ Scheduled job runner
crates/skald-core/src/tic/ TicManager: one tick of the TIC system agent for one user (run_for). No timer of its own — the instance-wide scheduler is skald::wiring::spawn_system_agents. See the system-agents section
crates/skald-core/src/compactor.rs Context compaction policy — when to compact and with which model; the mechanics are agent_loop::compaction. Model for the summary call: the instance-wide Settings pick (compaction_model, a PropertyType::LlmModel config property declared by compactor::config_set) wins; else AUTO by compaction.strength (config.yml); a missing configured model degrades to the same AUTO path
crates/skald-core/src/approval/ Approval rules engine
crates/skald-core/src/clarification/ ClarificationManager: background-session question/answer
crates/skald-core/src/elicitation/ ElicitationManager + bridge: MCP server-initiated input (elicitation/create), surfaced in the Inbox; secrets never logged/persisted
crates/skald-core/src/inbox.rs Inbox: unified façade for pending approvals + clarifications + elicitations (wraps ApprovalManager, ClarificationManager, ElicitationManager). The managers already emit the *Requested/*Resolved lifecycle events on the per-user bus; ws.rs forwards them to every connected client of that user regardless of source, so the web UI updates live (see sidebar.js row)
crates/skald-core/src/llm/ LLM client abstraction (OpenAI-compat, Anthropic, Ollama…). OpenAI-compatible provider types are runtime data, not code: providers/declared.rs loads providers.yaml at boot (see Config); only non-OpenAI-compatible or bespoke providers (anthropic, ollama, openai, openrouter) stay native. Retriability (Model::is_retriable, agent-loop) keys on the real HTTP status carried by ModelError { status }, not a substring of the message — a model id/token count containing "404"/"401" cannot mis-classify; 401/403/404/422 don't retry, 400/429/5xx/network do. Request logging is the logging.rs::LoggingModel decorator, attached by the caller's ModelSelector (loop_adapters/selector.rs::SkaldSelector::with_log) — never by LlmManager, which builds one shared client per model and cannot know whose traffic it serves. The decorator's RequestLogTarget carries the owner: metadata → llm_requests in the registry (user_id, the column the UI filters on), payload bodies/headers → llm_request_payloads in that user's own encrypted DB, keyed by request_id; session + frame come from the request's own conversation/frame, so kernel rounds, sub-agent frames and compaction summaries are all attributed with no extra plumbing (ModelRequest::log is unused here)
crates/skald-core/src/transcribe/ Transcription providers
crates/skald-core/src/image_generate/ Image generation providers
crates/skald-core/src/memory/ Agent memory tools
src/frontend/mod.rs WebFrontend: wires router_factory, starts plugins, runs Axum
src/frontend/server.rs Axum router, static file serving
src/frontend/api/ HTTP + WebSocket handlers — State<Arc<Skald>>
web/components/ Lit web components (see below)

DB tables (sqlx SQLite)

database/system.db — the path is a constant (core::db::SYSTEM_DB_PATH), not configurable. init_system_pool creates the directory; SQLite only creates the file. Per-user files are database/{userid}.db, created by UserManager::register_user and encrypted with SQLCipher.

The schema is split into two buckets (§5.1), and the split is the point:

  • create_registry_tables — instance-wide, readable without any user key: users, roles, llm_providers, llm_models, transcribe_models, tts_models, image_generate_models, plugins, plugin_access + plugin_user_configs, approval_rules, tool_permission_groups, config, known_tools, llm_requests, mcp_catalog, mcp_global_servers + mcp_global_access, oauth_providers, role_capabilities, shared_folders + shared_folder_members, projects + project_members. The MCP tables back the Connectors model (§7/§14/§15 — see its own section); oauth_providers (accessor db/oauth_providers.rs) holds one row per identity provider (Google…) — endpoints + client_id/client_secret + redirect_uri, admin-owned household secrets (§4/§15b), never a per-user token. The last two pairs are junction-backed membership: shared_folder_members (accessor db/shared_folders.rs) for the on-disk shared folders (§6), project_members (accessor db/project_members.rs) for projects (see the Projects section) — both let a member be read-only (can_write) and both drive the container mount topology + the fs routing. Their FKs are registry→registry (same file), which is allowed — unlike an owner→registry key.
  • create_owner_tables — one owner's content, identical schema in every file that has it: chat_sessions, chat_sessions_stack, chat_history, chat_llm_tools, chat_summaries, session_scratchpad, session_mcp_grants, stack_mcp_grants, scheduled_jobs, job_runs, system_agent_runs, mcp_user_servers, mcp_events, sources, secrets, llm_request_payloads, memory_docs (+ FTS5 memory_docs_fts). mcp_user_servers (a user's activated per-user connectors) carries catalog_name as a bare TEXT snapshot of mcp_catalog.name, never a FK — an owner→registry key would fail every INSERT; for an OAuth connector it also snapshots oauth_provider + deliver_json, and its api_key column holds the refresh token (in the SQLCipher-encrypted file, so no column crypto). Because memory_docs is an owner table, one definition backs private memory in each {userid}.db and shared memory in system.db (the household owner) — see the memory namespace note below. (projects/project_tickets were owner tables in the single-user past: projects are shareable now, so projects + project_members are registry tables and project_tickets is gone.)

Schema is greenfield (no migrations, §0), but a purely additive column lands on an existing DB in place: db::ensure_column runs ALTER TABLE … ADD COLUMN and swallows the "duplicate column" error, a no-op on a fresh DB where the CREATE TABLE already has the column. Used for the OAuth columns on mcp_catalog / mcp_user_servers so a dev box need not be wiped for an additive change (a full recreate is still valid).

No foreign key in the owner bucket may point at a registry table. SQLite cannot enforce a key across files, not even through ATTACH, and sqlx turns on PRAGMA foreign_keys: the CREATE TABLE succeeds and every INSERT fails. db::tests::owner_tables_stand_alone_with_foreign_keys_on enforces this by running the owner schema against a database holding nothing else, then inserting a row into each table. One key crossed and was fixed: chat_history.model_db_id (dropped — write-only, and llm_requests.model_name already records the model).

Memory namespace (blueprint §5). memory_docs (accessor db/memory_docs.rsget/upsert/list/search(FTS)/delete) backs a virtual note store surfaced through the fs-tools, not the disk. Two sibling roots (not the blueprint's nested memory/{userid} + memory/shared): user-memory/… routes to the caller's own pool (ToolContext::pool), shared-memory/… to the system pool (a singleton captured in fs::register_all). tools/fs/classify_memory() decides on the raw first path component (a .. in the tail clamps inside the store, never escapes to disk); read_file/write_file/list_files/edit_file/insert_at_line/replace_lines/search_file override run_with to route memory paths (each extracting a pure transform shared with its on-disk execute) and leave every other path on disk. The HTTP surface routes them the same way: GET /api/file classifies before resolve_view_path and serves the note from memory_docs (caller's pool / system pool), so the file viewer opens user-memory/… and shared-memory/… like any file, and show_file_to_user accepts memory paths too (existence-checked on the right pool). Approval (seeded in seed_fs_path_rules): user-memory/* is @fs_any allow (private, frictionless); shared-memory/* is @fs_read allow + @fs_write require — reads free, writes need approval so the agent can't silently push one person's data into shared memory. grep_files stays disk-only (regex-across-tree ≠ FTS); ranked full-text recall over notes is a separate tool, memory_search (tools/fs/memory_search.rs), over the memory_docs FTS index — allowed by a path-less rule (it takes query, not path).

Memory injection into the prompt: AgentSystemContext::load_inject_memory (loop_adapters/system.rs) routes each meta.inject_memory entry — user-memory/… → owner pool, shared-memory/… → the shared (system.db) pool, both via memory_docs::get; anything else (data/…, $WD/…) is a disk read. The shared pool is threaded ChatSessionManagerUserLoopRuntimeAgentSystemContext. assistant and project-coordinator inject user-memory/index.md + shared-memory/index.md.

Prompt substitutions: an AGENT.md may carry <!-- KEY --> placeholders; agents::resolve_includes turns each into a __KEY__ sentinel, replaced at request time. Two are resolved by the system-context source itself (loop_adapters/system.rs) from the session owner (user_id) + registry (shared_pool), so every source (WS, mobile, cron, sub-agents) gets them with no caller plumbing: __SHARED_FOLDERS__ (the user's shared-folders table) and __USER_PROFILE__ (the owner's directory profile: Name, Date of birth with age computed at build time, Sex, Preferred language, admin Notes — unset values render as explicit unknown / not specified, the Notes line is omitted when empty). Any other key comes from the per-call SendMessageOptions::system_substitutions map.

system.db still gets both bucket functions — but no longer because the migration is unstarted. It gets the owner schema because it is the owner of shared memory (memory_docs) plus, for now, the globally-scoped secrets (SecretsStore is built on the system pool and shared by reference into every UserContext; the global runtime's config now lives in the registry table mcp_global_servers, and per-user connector config in each user's owner mcp_user_servers). The global runtime no longer writes mcp_events there: notification persistence is an explicit McpManager::new argument (EventLog::{Persist,Discard}), Discard for the ownerless global runtime and Persist for each per-user one, because an event belongs to whoever it happened to and its only reader (TIC) is per-user. Every other owner table is created there but never written to anymore — the global owner-bound managers that would write them (chat/jobs/etc.) are inert (see "Current state"). Fully dropping create_owner_tables from system.db is blocked on the §4 scope decision for secrets, not on call-site migration.

users (crates/skald-core/src/db/users.rs) holds the directory plus auth material. It lives in the system DB, which the box owner can read, so it must never store anything that derives a user's key. Credentials is an enum mirroring the table's CHECK: an encrypted user carries a wrapped DEK (whose AEAD tag is the password verifier — hence no password_hash); a cleartext user carries an ordinary verifier, or none. User is deliberately not Serialize and its Debug redacts key material — use User::summary() for anything leaving the process. role_id references roles(id) (the roles table is now seeded before users in create_registry_tables). A nullable locale column (additive via ensure_column) holds the per-user UI language override; role-driven conventions live in the free-form roles.attrs JSON — never new columns per attribute — parsed at a single point by the typed db::roles::RoleAttrs (ui_mode, permission_groups, chat_agent): ui_mode (see the frontend section) plus the role's security-group set (roles.permission_group = the default group, attrs.permission_groups = additional allowed groups; Role::effective_groups() = the union, roles::role_allows_group() gates it with admin short-circuiting to all). See the security-group picker in the frontend section. The role's default entry (chat) agent is attrs.chat_agent — the neutral chat-type agent members of the role land on (§0.1: data, not an enum). Resolved by roles::default_chat_agent_for_user(registry_pool, user_id) — the single seam behind both the per-user ChatHub's default_agent (snapshotted at login in UserContextFactory::build, like fs/MCP access, so every session-creation path — explicit provision_session, lazy WS get_or_create_session, notify — honors it) and provisioning_for_source's non-project branch. Falls back to agents::DEFAULT_CHAT_AGENT ("assistant", the renamed former main) when unset. Seeded: admin/memberassistant, childrenkid (Companion). A per-user override is future work, layering on top in the same resolver. The stack root frame is created with the session's own agent_id (not a literal) — config.agent_id (from the frame) drives which prompt runs, so a wrong id there silently runs the wrong agent. The admin-managed directory profile lives in three more additive columns — birthdate (ISO YYYY-MM-DD), sex (free text), notes (admin-authored) — edited only from the Users admin page (set_directory_fields; validation — real non-future date, length caps — lives in the users_mgmt API, not the db layer) and rendered into agent prompts by the __USER_PROFILE__ substitution (see above). They are directory metadata written by the admin about the user, so the registry is their honest home under the §2 threat model.

Filesystem & containers (blueprint §6)

Each user has one permanent Docker container (skald-{userid}, our own skald-runtime image with python+node), created on user creation and started at boot (ContainerManager, crates/skald-core/src/container/). Docker is required: a missing daemon fails Skald::new and the process exits. The container runs as the host uid:gid (not root) so files created in-container and by the host-side fs-tools share ownership on the bind mounts (matters on native Linux; masked on macOS Docker Desktop). Because that user isn't root, the image ships passwordless sudo (a passwd/shadow entry is injected at create) so an agent can still sudo apt-get install …; --init runs tini as pid 1 to reap zombies.

The agent sees one namespace, routed on the first path component. The choke point is UserFs (core-api/src/user_fs.rs, a pure value type carried in ToolContext.fs), plus resolve_host_path() in tools/fs/mod.rs:

Agent path Backing Routed by
user-memory/… SQLite ctx.pool ({userid}.db) classify_memorymemory_docs
shared-memory/… SQLite system.db classify_memorymemory_docs
shared/{X}/… host {WD}/shared/{X} (if a member) UserFs::host_base_and_tail
projects/{O}/{S}/… host {WD}/projects/{owner_userid}/{S} (if a member) UserFs::host_base_and_tail
~/…, relative host {WD}/homes/{userid} UserFs::host_base_and_tail

Two views, one storage: the fs-tools run host-side in the Skald process on {WD}/homes/{userid} + {WD}/shared/{X}; execute_cmd runs inside the container (docker exec -w <container-path> skald-{userid} sh -c …, via ExecuteCmd::run_with) on the same paths bind-mounted (homes/{userid}/root, shared/{X}/root/shared/{X}, read-only when can_write=0). A file written in the container appears to the host fs-tools and vice versa.

Containment (resolve_host_path): every physical fs-tool op canonicalizes the resolved path (following symlinks) and prefix-checks it against its mount base, fail-closed. Since the same tree is writable from inside the container, a symlink planted there that points outside the home/shared root is caught here — the host-side tool never escapes the user's workspace. grep_files stays disk-only (regex ≠ FTS; memory → memory_search) but resolves its root the same way. execute_cmd's workdir is an agent path mapped to its container path via UserFs::to_container.

The threading: UserContext.fs (built by container::build_user_fs at login, snapshotting shared memberships) → ChatSessionManagerChatSessionHandler.fsToolContext.fs. Admin CRUD is wired (src/frontend/api/shared_folders.rsGET/POST /api/shared-folders, PATCH/DELETE /api/shared-folders/{id}, POST/DELETE .../members[/{user_id}]; UI shared-folders.js): a create/describe/delete + per-member can_write surface, and each mutation emits SystemEvent::UserMountsChanged, on which the lifecycle reconciler runs Skald::refresh_user_mounts — rebuilding the affected user's fs + container mounts in place, so a membership change lands without a re-login (blueprint §6's "admin CRUD" + "membership refresh without re-login" TODOs, now closed; it still settles at next login/boot if the live remount fails). execute_cmd /stop is robust: the command runs under setsid -w in its own process-group (leader pid recorded in a container pidfile), and a KillReaper drop-guard reaps that group on /stop or timeout via a detached docker exec that walks /proc and kills members by positive pid (the container's dash mishandles kill -<pgid>); the pidfile is passed positionally ($1), and the container's --init (tini) reaps the killed processes so no zombies accumulate. Per-user MCP connectors now run inside this container (§7) — the container infra enabled it; see the MCP connectors section.

Projects

A project is a shareable, self-service workspace: a folder at {WD}/projects/{owner_userid}/{slug} plus membership in the registry. projects (accessor db/projects.rs — slug is immutable, UNIQUE(owner_user_id, slug)) + project_members (junction with can_write; the owner is always a write-member, so a private project = one member). Sharing is not admin-gated: the owner and any write-member can add/remove/re-grant members and edit metadata; only the owner can delete. Each membership mutation emits SystemEvent::UserMountsChanged for the affected user; the lifecycle reconciler remounts their container in place (Skald::refresh_user_mounts), so the folder is browsable at once (the explorer reads host-side) and reachable from execute_cmd a moment later. The mount appears in the agent namespace as projects/{owner_username}/{slug} (host keys on the stable userid, agent path on the username) — read-only members get a read-only bind mount in the container.

API (src/frontend/api/projects.rs): GET/POST /api/projects, GET/PUT/DELETE /api/projects/{id}, POST /api/projects/{id}/members, DELETE .../members/{user_id}, POST /api/projects/{id}/session. ProjectDetail carries root_path — the agent path of the folder, computed server-side (owner username ≠ owner_name, which may be a display name) — the explorer's root. A project-{id} chat source provisions the project-coordinator agent with a project RunContext (provisioning_for_sourceskald_core::projects::build_project_run_context: project_root + a system block with name/description/folder/members); every member keeps their own private project-{id} session — only the folder is shared.

UI (web/components/projects/): index.js (<projects-page> host — hash-routed: #projects, #projects/{id}, #projects/{id}/sharing, back/forward-aware), project-list.js (card grid + create/edit/delete modal), project-board.js (<project-board-section> — the detail page: header with Open chat, then a Files / Sharing tab bar using the .project-tab-bar styles in css/projects/board.css), project-files.js (<project-files-panel> — the explorer). The mobile app has its own read-only shared/projects-page.js (list → open project chat).

The explorer (project-files.js): one directory at a time via GET /api/files/dir?path=… (new endpoint in src/frontend/api/files.rs: immediate children with name/path/is_dir/size/created_at/modified_at, dirs-first; same resolve_view_path scoping as /api/file). Breadcrumb rooted at the project (/ = root_path); file click → window.openFile (existing viewer); folder click → navigate. Live: it subscribes the open directory on the existing /api/file/watch socket (web/lib/file-watcher.js singleton — notify NonRecursive on a dir reports its direct children) and reloads debounced 300 ms, so files created by other members or by the agent in-container appear without a refresh. Write actions (new folder, upload incl. drag&drop, rename, delete) are shown only to can_write members and ride the existing /api/file endpoints — POST gained dir:true (mkdir), DELETE handles directories (remove_dir_all), and binary upload is the new POST /api/file/upload?path=… (raw body, 256 MiB DefaultBodyLimit). Server-side write gate: all /api/file write handlers now call UserFs::can_write_to(agent_path) (core-api) — home → true, shared//projects/ → the membership's can_write, docs/ → false — closing the host-side bypass of the read-only bind mount (the container mount only gates in-container writes).

MCP connectors (blueprint §7/§14/§15)

MCP servers are surfaced to users as "Connectors" (UI naming; mcp/schema stays neutral, §0.1). The old single owner table mcp_servers, the agent-facing register_mcp/delete_mcp tools, and the mcp kinds of list_items/toggle_item are gone. Connectors are now admin-curated and user-activated through the Connectors UI/API — never written by the agent, which closes the §14 RCE vector (prompt-injection → agent writes+registers a local script → arbitrary code on the box).

Two runtimes, one view (§7). A session's MCP tools are the union of:

  • Global runtime — shared, stateless connectors (web-search, Tavily…) that run on the host, connected at boot from mcp_global_servers by McpManager::initialize. Filtered per user by mcp_global_access.
  • Per-user runtime — the connectors a user has activated, run inside their container, started at first login from that user's owner mcp_user_servers and living until restart (§9; the docker exec -i children die via kill_on_drop when the UserContext drops).

McpProvider (mcp/provider.rs) is the trait the session code talks to, so all_tool_defs / render_mcp_list / ActivateTools never learn which runtime owns a server. McpManager implements it directly (used for the inert ownerless bundle, §19); UserMcpView implements it as global user, where accessible_global is a snapshot of mcp_global_access captured when the UserContext is built (like fs membership). Both runtimes share McpManager::connect_all(specs, boot); McpServerSpec + global_row_spec/user_row_spec turn a DB row into a connectable spec (a per-user local_script spec targets the user's container).

Authorization is a capability on the role, not if role==admin (§0.1/§14 — db/role_capabilities.rs): mcp.register_remote + mcp.register_local_from_catalog are self-service (seeded on every new role by roles::create via seed_defaults); mcp.register_local_script + mcp.manage_catalog are admin-only. admin holds every capability by construction (short-circuit in has()). API handlers gate through require_cap.

Tables (see DB section) — registry: mcp_catalog (admin-vetted templates; holds only the schema of what an activation must supply, never live creds — plus, for OAuth, oauth_provider + oauth_scopes_json + deliver_json), mcp_global_servers + mcp_global_access, oauth_providers (per-provider client creds), role_capabilities. Owner: mcp_user_servers (per-user activations; api_key encrypted at rest — the refresh token for an OAuth one — catalog_name/oauth_provider/deliver_json bare TEXT snapshots).

Endpoints (src/frontend/api/mcp.rs, mounted in api/mod.rs) — admin: /mcp/catalog (GET/POST/DELETE), /mcp/global (list/enable/delete + /{id}/access GET/PUT), /mcp/providers (GET/POST + DELETE /{name} — OAuth provider creds, secret never returned to the browser). User: /mcp/available, /mcp/activate, /mcp/activated (+ DELETE /{id} to deactivate), /mcp/oauth/start + /mcp/oauth/complete (the §15 OAuth login), /mcp/login/status + /mcp/login/reset (the §15 QR/device login — see below). connectors.js (<connectors-page>) is the single Connectors surface — a row list, one row per connector (there is no separate catalog page): the user view (activate/deactivate + granted globals) always, plus the admin affordances when role_id === 'admin' — the Add connector dropdown (from the Marketplace, or manually via the #connectors/new sub-page), per-row removal from the catalog, and the Sign-in providers modal. The Marketplace stays its own page (marketplace.js), reached from that dropdown and linking back to #connectors. connector-detail.js (<connector-detail-page>) is a connector's own page and hosts both the OAuth login panel and the QR login panel.

Dependency reconciler (mcp::install::ensure_installed). Copying a local-script connector's files into a container never installed its deps. ensure_installed closes that: a content-hash reconciler keyed on the connector's source files (not a version string) that, when the hash changed, re-copies the files and installs deps inside the container — npm ci --omit=dev (node, from package.json) and/or pip install --target .pydeps (python, from requirements.txt, put on the server's PYTHONPATH by user_row_spec). Runs at activation and on every per-user startup path (UserContext build, remount) via mcp::prepare_local_connector, so a fresh container installs from scratch, an updated connector re-installs, and an unchanged one is a hash-match no-op. Deps are therefore never vendored — connectors ship package.json/requirements.txt, not node_modules/. Authoring contract for connectors lives in scripts/CONNECTOR_MANIFEST_GUIDE.md.

Connector versioning. mcp_catalog carries version (INTEGER — the update-comparison key), version_string (semver, display) and version_release_date (ISO, display), snapshotted from the feed on install. The marketplace list computes update_available = feed version > installed version (strict) and surfaces it as an "Update" button (marketplace.js). The integer is the UI signal; the actual re-install trigger is the reconciler's content-hash.

OAuth per-user connectors (blueprint §15 — copy-paste flow)

OAuth2 authorization-code + PKCE is wired for per-user connectors (Gmail is the first). The consent is a human copy-paste, not a headless action: no callback route into the (NAT'd, hostname-less) box, and no client secret on the public feed.

  • Providers, not per-connector URLs. The client is per-provider (one Google app covers Gmail/Calendar/Drive): oauth_providers holds auth_url/token_url/client_id/client_secret/redirect_uri/extra_params, admin-entered via the Sign-in-providers modal (Google preset fills all but the two secrets; redirect_uri = the static oauth/show.html page, extra_params = access_type=offline+prompt=consent so Google returns a refresh token). The manifest only names auth.provider + auth.scopes + auth.deliver — never URLs or secrets (feed is remote data, §14).
  • Flow (mcp/oauth.rs): activate on an OAuth catalog entry persists a pending mcp_user_servers row (files installed, command wired, no token) and returns needs_oauth — it does not start the server. /mcp/oauth/start builds the consent URL (PKCE S256 + opaque state) and stashes the verifier in a RAM-only, TTL'd flow store keyed by state; the user approves in a browser, the provider lands the code on oauth/show.html, they paste it back. /mcp/oauth/complete exchanges code+verifier for a refresh token (client_secret sent server-side), stores it in the row's api_key, flips to ready, and starts the server. PKCE makes an intercepted code worthless; a restart drops in-flight flows (mirrors the RAM-only session model).
  • Credential delivery = env, nothing on disk. The manifest's deliver ({as,format,env}, parsed as mcp::DeliverSpec) says how the token reaches the server. user_row_spec_resolved assembles the credential (google_authorized_user JSON = client creds from the provider + refresh token) and injects it as an env var (GMAIL_CREDS_JSON) on the docker exec — never a file, coherent with §2 (the tempted admin doesn't read /proc). The server reads it via Credentials.from_authorized_user_info. Ran both at OAuth-complete and at login-time per-user startup.
  • Google needs a Web-application client: a Desktop client rejects an https:// redirect (loopback only), so the oauth/show.html redirect must be registered on a Web app OAuth client, and exact-match under Authorized redirect URIs — redirect_uri_mismatch otherwise.

QR / interactive device login (blueprint §15 — polling flow)

For a per-user connector whose credential is produced by pairing (auth.type: "qr"; WhatsApp is the first, on Baileys — the slim skald-runtime image has no Chromium, so a browser-based client is out), there is no code to paste and the server must run to produce the QR. The seam is a generic tool contract, reusable for future device kinds (SSH…):

  • login_status tool contract. A connector needing an interactive login exposes one tool, login_status, returning JSON {state, qr?, message} (state: connecting|need_scan|ready|logged_out; qr is a data-URL PNG only while need_scan). Skald calls it directly, never the agent.
  • Flow. activate on a qr entry inserts a pending mcp_user_servers row and starts the server (unlike OAuth, which defers), returning needs_login/login_kind:"qr". /mcp/login/status ensures the server is running (restarts a pending one), calls login_status, and returns its state; on ready it flips the row's auth_state so all_startable picks it up next login. /mcp/login/reset calls the connector's logout tool to re-arm (link a different device). The connector-detail.js QR panel polls login/status and renders the QR.
  • Credential = on-disk session, not a token. The connector persists its session inside its own dir (e.g. ./auth/), under the bind-mounted home so it survives a container recreate — the honest §4 gap (admin-root-readable), not memory_docs.
  • Node 18 gotcha: the container ships Node 18; Baileys uses the Web Crypto global, so the server must globalThis.crypto ??= require('crypto').webcrypto or it dies pre-QR with "crypto is not defined".

Deferred: SSH and other §15 device kinds (would reuse the login_status contract), deliver.as=file, and non-Google OAuth providers are unimplemented paths that error clearly rather than half-work. No boot seed of catalog presets; the admin populates the catalog from the Marketplace.

System agents (TIC)

A system agent runs on a user's behalf without being asked. TIC — the background event processor — is the only one, and the surface is written so a second needs no new machinery.

It is per-user, and every part of the design falls out of that. The events it reads (mcp_events) are in the caller's own encrypted database, pushed there by connectors running in the caller's container; the notification it emits goes to the caller's hub; the trace it leaves (system_agent_runs) is in that same file. TicManager (crates/skald-core/src/tic/) therefore owns no timer and no user list: it exposes run_for(user_id, pool, sessions, hub), one tick for one user, over deps unpacked from that user's UserContext. Building it against the ownerless Conversation bundle was exactly what made the pre-multi-user version inert — it wrote sessions into system.db, notified a hub with no subscribers, and resolved tool paths against a container that does not exist.

One scheduler, sequential. skald::wiring::spawn_system_agents is the instance-wide loop — spawned post-construction with a Weak<Skald>, like spawn_user_lifecycle and for the same reason (it resolves per-user runtimes through Skald::user_context). Each pass walks the directory and runs the agent for one user at a time: a pass is N container round-trips and N LLM calls, and nobody is waiting on a background tick, so concurrency would only spike the box every interval. A ConfigKeyUpdated on the interval key cuts the current wait short; the enabled flag is re-read per pass.

A locked user is skipped, and that is the normal case, not an error. The pool is the unlock token (§9): a user who has not logged in since the last restart has no readable events, no session store — and no place to record the skip, since the only file that could hold it is the one we cannot open. Hence system_agent_runs has no skipped status: the skip is an INFO log line and nothing else. Their events keep accumulating and the first pass after they log in picks them up.

The run log is theirs, not the admin's (db/system_agent_runs.rs, owner table, no user_id column — the file is the owner). A run summarises what landed in someone's inbox, so GET /api/system-agents/runs is scoped through require_context with no admin override: everyone, admin included, sees their own runs. stats is a JSON blob of the agent's own counters (never event contents). The write is split start/finish (unlike job_runs, written once at the end) so a crash leaves a visible running row, swept to failed by the next start for that agent — safe precisely because the scheduler is sequential and single-instance. An idle tick writes nothing: a row only exists when there were events, or the log becomes a heartbeat instead of a history.

The configured security group is not applied verbatim. tic.security_group is an instance-wide admin setting; handing it to a restricted member's run would give their background agent a tool set their role never granted. It goes through run_context::reconcile_group_for_user — the same seam a persisted group takes — degrading to the role default when the role disallows it. With nothing configured the run still starts from role_default_run_context, never None, because None means the catch-all group, which is wider.

UI: #system-agents (web/components/system-agents.js, sidebar group extensions, visible to everyone — there is nothing to gate when the data is the caller's own). It replaced the old #tic "TIC Sessions" debug page, which listed chat_sessions WHERE source='tic' and so inferred runs from leftover ephemeral sessions rather than recording them.

Multimodal attachments

Uploads go through one centralized seamChatHub::save_upload (behind ChatHubApi::save_upload, backed by skald_core::uploads::save_to_home) — so every surface persists identically and no two callers can drift on placement (the class of bug where the agent was handed a path it couldn't reach). The seam writes into the caller's container home under uploads/{session_id}/ (agent path uploads/{session}/{name}, the UPLOADS_SUBDIR const in core-api/user_fs.rs), collision-dedupes the name, and prefers the sniffed magic-byte MIME over the client claim. The web handler (POST /api/{source}/uploads) buffers each field with a 256 MiB cap then calls the seam; the Telegram plugin downloads bytes then calls the same seam via handle.chat_hub().save_upload("telegram", …). Because the file lands in the home (bind-mounted at /root), it is reachable by the fs-tools, execute_cmd, and the file viewer (GET /api/file, per-user via resolve_view_path) — there is no /data static route anymore (removed: it was require_auth-only, not ownership-scoped, and also exposed internal server state under data/). Attachment metadata travels as structured JSON in chat_history.metadata — never as persisted text.

At context-build time (the crate's projection), attachments of the current turn (the user/agent rows following the last completed assistant reply, including across in-flight tool rounds) are partitioned by agent_loop::projection::media, with loop_adapters/media_source.rs deciding which files may be handed over (§6 containment): when the resolved model's LlmEntry.capabilities include the modality (visionimage_url parts, videovideo_url parts), the file is inlined as a base64 data-URL content part — but only if it resolves (through the caller's UserFs, via resolve_host_path) under the home's uploads/ dir, its sniffed MIME is in the allowlist, and it fits the budgets (4 files / 10 MiB image / 32 MiB video / 48 MiB total per turn). Everything else — older turns, other kinds, any failed check — keeps the textual <system-extra> path block (built by core_api::message_meta::attachments_block / system_extra; the tag name is the single SYSTEM_EXTRA_TAG constant), so a non-vision model produces a byte-identical payload to before. OpenAiClient forwards parts verbatim; AnthropicClient translates image_url data URLs to image blocks (video unsupported; Anthropic models get vision by editing the model row's capabilities — no catalog refresh writes them). On LLM fallback mid-round, messages are rebuilt with the replacement model's capabilities.

Token streaming & reasoning display

The chat streams tokens live, as a parallel best-effort side-channel that never alters the turn's authoritative flow: the final Done (or Thinking) event still carries the complete content and the frontend treats it as truth.

  • Client seam (core-api::chatbot): ChatbotClient::chat_with_tools_raw_streaming(..., delta_tx: mpsc::Sender<StreamDelta>) — default impl ignores the channel and calls the buffered chat_with_tools_raw, so providers without streaming (Ollama, LM Studio) are untouched. StreamDelta::{Text, Reasoning} splits visible answer from chain-of-thought. Senders use try_send (deltas drop when the channel is full) — streaming must never backpressure the HTTP read.
  • SSE implementations (crates/llm-client): OpenAiClient (stream:true + stream_options.include_usage, reasoning_content/reasoning deltas, index-based tool_calls accumulation, usage from the final chunk) and AnthropicClient (stream:true; message_start/content_block_*/message_delta events; thinking_delta → reasoning, input_json_delta → tool input). Both reassemble the same LlmTurn + LlmRawMeta the buffered path returns (the payload log stores a synthesized buffered-shaped body). Failure policy: if the stream dies before any delta the client retries buffered on the same model (providers rejecting stream keep working); a mid-stream failure propagates to the normal model-fallback logic. Framing is shared (llm_client::SseDecoder). Anthropic's buffered path now also parses thinking blocks into reasoning_content (previously discarded).
  • Loop wiring: call_llm_round creates the delta channel per attempt and a forwarder task maps deltas to ServerEvent::TokenDelta { kind: content|reasoning, delta } on the turn's event channel (drained before the round's outcome events, so ordering holds); cancellation drops the in-flight future as before. A mid-stream fallback is handled client-side: the frontend clears its pending bubble on model_fallback.
  • Reasoning surfacing: reasoning_content rides Done/Thinking events (so buffered providers show it live too) and is projected as reasoning on assistant/thinking history items (build_items); persistence in chat_history.reasoning_content and the echo back into context predate this feature.
  • Frontend (chat-session.js + copilot-render.js, shared by desktop copilot and mobile chat-page): token_delta accumulates into a pending assistant bubble (in-place mutation + ~15 Hz flush, blinking caret); done/thinking finalize it in place, error/llm_failed/model_fallback drop it, tool_start/agent_done finalize orphan bubbles (reasoning-only rounds, sub-agent final rounds that emit no Done). The reasoning block is a muted, collapsed-by-default native <details> (renderReasoning, .reasoning-block in copilot-messages.css, i18n key chat.reasoning) — open state survives re-renders, and it renders identically from live events and from history.

The LLM loop (agent-loop)

The loop is a standalone crate (crates/agent-loop/) that knows nothing about Skald: it owns control flow (rounds, model fallback, tool fan-out, recording), the projection of history into wire messages, sub-agent delegation, restart recovery and compaction. Skald supplies content through the traits in crates/skald-core/src/loop_adapters/. Nothing in session/handler/ shapes a Value anymore — there is exactly one projection in the workspace.

One LoopManager per user (UserLoopRuntime, loop_adapters/runtime.rs, blueprint D12), built by ChatSessionManager: it owns the event bus, the live-loop registry (which conversations are running, /stop, recovery, shutdown), the store, the approval gate, the hooks, the agent catalog and the delegate tool. A turn contributes only what is its own — the agent's prompt, its tool set, its model pin — via turn_params.

Per-turn state rides the Extensions type-map (loop_adapters/scope.rs::TurnScope): the gate and the catalog live as long as the user, so they cannot capture a session id or a permission group — they read the turn's scope from the call's extensions. A call with no scope is denied, never run with permissive defaults.

Three entry points, all in session/handler/kernel_turn.rs:

entry when what it does
run_kernel_turn a user message repairs a dangling call from a crashed turn, then manager.start_turn
recover_turn WS connect, async result delivery, background wake-up Recovery::run — no new message, continue what was interrupted
resolve_pending_call an approval answered after a restart run the call with the gate skipped, then continue

The event translator (loop_adapters/translate.rs) is the ONE bus subscriber turning LoopEvents into the session's ServerEvents; byte-parity with the pre-kernel event sequence is its contract.

Sub-agents

  • A sub-agent is a tool, not an interception: DelegateTool (registered under the legacy names execute_task / execute_subtask, D11, each keeping its exact legacy schema) opens a child frame and runs a normal loop in it. The parent simply awaits a slow tool call. Max depth MAX_AGENT_DEPTH = 5.
  • Parallel batches are the kernel's generic fan-out: a round whose calls are all concurrency_safe (a sync delegate is) runs concurrently, bounded by max_parallel_calls. The ordering invariant is unchanged — ids allocated in call order (phase 1) → concurrent execution (phase 2) → recording in call order (phase 3) — so the model reconstructs results by id. Any mixed batch stays sequential. Siblings share the session scratchpad; concurrent writes to the same key are last-writer-wins by design.
  • mode: "async" submits a durable scheduled_jobs row through loop_adapters/async_task.rs::CronExecutor and returns a receipt immediately; when the job finishes, DurableSink writes the result into the parent conversation (synthetic assistant + a completed task_completed call) and resumes it. mode: "cron" is scheduling, not delegation, and stays on the cron interface tool.
  • A child's model is never inherited from the parent: passing a concrete name would bypass AUTO selection, so sub-agents auto-select unless explicitly overridden (args.clientmeta.json client → AUTO by strength).
  • list_agents returns task agents only (never chat/system ones like the entry agent).

Restart recovery (agent_loop::recovery)

A crash loses RAM (the approval oneshot, the cancellation token), never truth: every state transition is a store write. So recovery does not have a mode of its own — it makes the history well-formed and then runs a normal loop on it:

  1. Reap an interrupted parallel batch (≥2 active frames at one depth is impossible for a linear stack): fail their spawning calls, close the frames. Deliberately lossy.
  2. Resolve the deepest frame's non-terminal calls. A Running one is re-gated and re-executed unless the tool says otherwiseexecute_cmd declares RestartHint::MarkInterrupted (D7), because a command may already have had its effect. An AwaitingHuman one is re-asked (the card reappears).
  3. Un-wedge: a child that finished but whose result never reached its parent propagates without calling the model again.
  4. Cascade to the root, resolving each parent call with its child's result — every frame running as its own agent, from the catalog, never the root's (B3).

Cancelled and Rejected are terminal and are never re-executed. Anti-double-driving goes through the manager's registry (a recovery claims the conversation like a live turn), not a host-side flag.

Cancellation (stop)

  • The turn's CancellationToken is minted by LoopManager::start_turn and cloned by value down the whole call tree; a delegate passes ctx.cancel.child_token(). It is never re-read from a field mid-turn, which is what makes /stop sticky across sub-agent recursion.
  • ChatSessionHandler::cancel()manager.cancel(&conversation). The token is checked at each round boundary and before each tool call, wrapped around the in-flight LLM call (tokio::select!, aborting the request), and around execute_cmd (dropping the future → kill_on_drop). Parent and child share the tree, so a cancelled child stops the parent by construction.

Compaction

agent_loop::compaction owns the mechanics: split point (never between an assistant turn and its tool results), transcript, prompt (SUMMARY_PREFIX / preamble / template live there now), the single no-tools model call, the saved summary row. skald-core/src/compactor.rs owns the policy: the token threshold, the ephemeral guard, which model summarises (compaction_model from Settings, else AUTO by compaction.strength), and publishing CompactionEvent on the chat bus. The DTL re-anchor is the on_compacted hook (loop_adapters/hooks.rs::DtlReanchorHook). The next turn needs nothing: the assembler reads the latest summary from the store.

Approval gate

The rule engine ApprovalManager::check returns Allow/Deny/Require per tool call (default rules seeded on first boot; the catch-all * require @999999 gates anything not explicitly allowed — e.g. execute_cmd, execute_task, writes outside whitelisted paths). It is wired to the loop as loop_adapters/gate.rs::ApprovalGate (agent_loop::gate::Gate). A Require registers a oneshot in the in-memory pending map keyed by request_id and emits an approval event over WS.

Resolution is source-agnostic: the WS + Inbox paths resolve by request_id; the inline chat card resolves by the durable tool_call_id via POST /api/tools/:tool_call_id/resolve (resolve_tool in src/frontend/api/sessions.rs), which derives the owning session from the tool call's own stack row — never a hardcoded source. Live pending cards fire the oneshot. Post-restart there is one path for every tool, LoopManager::resolve_pending: the call runs with the gate skipped (the human just decided) but with the session's real ToolContext — owner pool, per-user container — so a resolved write_file/execute_cmd acts on the user's workspace, never the server cwd/host (this was a §6 escape); then the conversation continues, including a sub-agent dispatch, which simply opens its child frame like any other call. The endpoint returns as soon as the work is scheduled and the result streams over the bus.

The diff preview in a PendingWrite event (loop_adapters/preview.rs::read_current_content, driven by the SkaldWritePreviewHook) routes exactly like the fs-tools: user-memory//shared-memory/memory_docs on the right pool, every other agent path → the caller's host workspace via resolve_host_path(&self.fs, …). It must never use the cwd-relative fs::resolve — that showed a bogus "new file" on overwrites (or the diff of a same-named cwd file), so the user would approve the wrong diff.

Tool visibility in the Security-groups UI (GET /api/approval/tools): tools injected outside the ToolRegistry (interface/plugin/provider tools) would otherwise be un-configurable. ToolCatalog::list_all() covers registry tools + a static synthetic_tools() list of core interface tools; everything else is captured by crates/skald-core/src/tool_discovery.rs (ToolDiscovery), which taps the tool set the loop offers each round (SkaldToolSet::defs) and upserts every offered tool into the known_tools table (in-memory seen-set guard → background DB write). list_tools merges known_tools (deduped, category: "dynamic") so any tool offered at least once becomes gate-able. Drift-proof by construction; core never hardcodes plugin tool names.

Restart

There is no in-app restart anymore. The agent-callable restart tool and its set_restart_handler seam were removed (blast radius = the whole box: it dropped every user's session and in-RAM DEK from one user's chat — a power-user leftover, out of place in the multi-user model). Nothing in the process now calls libc::_exit(-1).

The supervisor protocol survives but is currently unreachable in-app: run.sh still re-executes the binary by path when it exits 255, but no code produces that exit code. Restarting is therefore a manual/admin operation.

To pick up config.yml / providers.yaml / database changes (read only at startup), or to load new code (./build.sh installs the new binary via atomic rename): stop the server and let run.sh loop, or re-run ./run.sh. A future admin-only restart action (endpoint/button gated by an admin capability) would re-use the 255 ⇒ re-exec seam — it is intentionally kept for that.

run.bat is still stale (cargo run) and must be fixed.

Build & run

./build.sh      # release build → bin/skald and bin/skald-setup (atomic install)
./build.sh -d   # debug profile; extra args are forwarded to the server build
./run.sh        # first-run setup, then the supervisor loop — never compiles

build.sh builds and installs both binaries; any forwarded args go to the server only.

run.sh resolves the server binary as $SKALD_BINbin/skaldtarget/release/skald, and warns when sources are newer than it. Before the loop it runs skald-setup (found next to the server, or $SKALD_SETUP_BIN); a non-zero exit there — a failed or cancelled wizard — stops run.sh before the server starts. Server exit 0 stops the loop, 255 re-executes, anything else propagates.

In a debug build, Argon2id at 256 MiB is unoptimised and takes far longer than the ~1s of a release build — skald-setup -d will feel stuck at the password step. Use the release binary for anything interactive.

Tracing filter: RUST_LOG=skald=debug,info

Adding an agent

Create agents/<id>/meta.json and agents/<id>/AGENT.md. The agent is discovered at runtime (no restart needed for prompt edits). Optionally set "client": "<name>" in meta.json to pin a specific LLM.

Documentation

docs/ is not developer documentation — it's written for the in-app LLM, not for a human reading the repo, and is mounted read-only into every user's container at ~/docs/ (see the Filesystem & containers section: docs_host on UserFs, DOCS_DIR in container/mod.rs). It explains the software's UX (plugins, and eventually agents/connectors/memory/roles/…) in plain terms, in English, so the assistant can help a non-technical user configure things instead of guessing. docs/index.md is the entry point (general index of feature pages); docs/plugins/<plugin id>.md covers each built-in plugin. The three type: chat agents (assistant, kid, project-coordinator) are told in their AGENT.md to read docs/index.md when a user asks how the software works. Standing rule: every change that impacts the UX must update docs/ in the same change — a new/renamed feature page plus the docs/index.md index entry. It goes stale like any other doc, except users actually see this one.

Config

Copy default.config.yamlconfig.yml. Never commit config.yml (contains API keys).

providers.yaml (repo root, cwd-relative like config.yml) declares the OpenAI-compatible LLM provider types — endpoints, UI metadata, per-model JSON field mapping, id-glob enrichment rules, reasoning knobs. Loaded at boot by llm::providers::declared; edit + restart the process, no rebuild. An invalid entry is logged and skipped, never fatal; an id colliding with a native provider is skipped. Adding a new OpenAI-compatible provider is a YAML edit, not a Rust file. The shipped file is validated by a unit test (declared::tests::shipped_providers_yaml_is_valid).

Python environment

All Python scripts (MCP servers, setup scripts) use a local virtualenv at .venv/ in the project root.

run.sh creates it automatically on first launch (using uv if available, otherwise python3 -m venv) and installs requirements.txt. It then prepends .venv/bin to PATH before starting the app, so every child process — MCP server launches, execute_cmd shell calls — resolves python3 to the venv automatically. No manual activation needed. Python is optional: if neither uv nor python3 is found, the app starts normally and only Python-based MCP servers will be unavailable.

To add a Python dependency: add it to requirements.txt. It will be installed on the next ./run.sh invocation if .venv does not yet exist — or run uv pip install -r requirements.txt manually.

Frontend components (web/components/)

All extend LightElement from web/lib/base.js (Lit). ChatSession (web/lib/chat-session.js) is the shared base for WS-connected chat UIs.

The chat is the home page. <app-copilot> is a single persistent element with two layout modes driven by the route (llm-page-change): mode="full" on the home route (it fills the workspace — the conversation IS the landing page, with a welcome hero + prompt suggestions as its empty state) and mode="dock" on every other route (the classic resizable side panel). Same element ⇒ WS, tabs, scroll and drafts survive navigation; you watch files/projects update live while the conversation keeps going. Collapse only applies to the dock. The old dashboard content (hero, LLM stats charts, pending inbox, quick guide) lives on as the separate #dashboard page; the debug toggle moved to the Settings page.

Theme (web/css/variables.css): warm "paper" palette (terracotta accent, light by default, warm-charcoal dark), generous radius (--radius-sm/md/lg), 16px-base chat type, WCAG-fixed contrasts, global :focus-visible ring and prefers-reduced-motion support. Everything consumes CSS variables — never hardcode a hex in a component stylesheet.

i18n (web/lib/i18n.js + web/i18n/{en,it,fr}.js): t(key) helper, I18nMixin re-renders on locale-changed. Resolution order: user preference (users.locale, editable on the profile page) → instance default (registry config key ui_locale, editable by the admin in Settings — declared in skald_core::i18n::config_set) → English. Server-side, never re-implement that chain: skald_core::i18n::resolve_locale(pool, user_locale) is the one function (with default_locale(pool) and language_name(locale) for prompt rendering); they read through db::config because the bus only matters for writes and callers like the system-context source hold pools, not the manager. Pre-auth screens use the localStorage cache. Default locale is English. First-run setup asks the language in both shells — the console wizard writes ui_locale via skald_core::i18n::set_default_locale (no system bus exists there), the web setup page sends locale to POST /api/setup/user, which writes it through GlobalConfigManager::set. Supported locales are centralized in skald_core::i18n::SUPPORTED_LOCALES and enforced server-side on every write. Translated so far: chrome (sidebar/topbar), chat + approval cards, login/setup, profile, inbox; deep admin pages are still English (fallback is automatic per-key). Copy is the only place domain words may appear (§0.1).

Plugin & backend i18n — two seams, both keyed the same way. A plugin page fragment (served from its own router) localizes client-side: it ships a web/i18n.js module (export default { en, it, fr }, keys namespaced plugin.<id>.<key>) and calls addStrings(dicts) (in web/lib/i18n.js) once at module load to merge into the host's shared DICTS, then uses the same t()/I18nMixin as the app (the fragment imports them from the absolute /lib/i18n.js — the same module instance the host uses, so t() and locale-changed are shared; no endpoint, no per-locale fetch — all locales ride in the fragment, so a language switch is instant). Mobile-connector is the reference: common.js registers the dict + re-exports t, and MobileBase extends I18nMixin(LitElement). Backend-generated strings (a plugin's HTTP error/response text, notifications) go through core_api::i18n: a plugin declares Plugin::i18n() -> Vec<LocaleBundle> (mobile-connector loads them from embedded i18n/{en,it,fr}.json via include_str!), the PluginManager merges every plugin's bundles once at boot into an I18nCatalog (skald_core::i18n) and injects it as PluginContext.i18n: Arc<dyn I18nApi>. At request time the handler resolves the caller (Caller.user_id from the auth layer) and calls i18n.for_user(user_id, key, args).await — which reads users.locale, runs it through the same resolve_locale chain, and renders locale → en → key with {name} placeholders. The frontend surfaces these already-translated: jf() throws the server's response text verbatim. Front and back keep separate tables (UI labels ≠ error strings; overlap is minimal) but share the plugin.<id>. namespace convention. The mechanism is general (any plugin, and eventually the core, registers the same way); only mobile-connector uses it so far.

Role-driven interface (§0.1 — data, not enums): roles.attrs JSON may carry "ui_mode": "simple". /api/auth/me resolves it via RoleAttrs (admin is always full) and the sidebar renders chat + inbox only for simple-mode members; the role editor exposes it as an "Interface" select. Hiding links is never access control — routes stay capability-gated server-side. MeResponse also carries locale, default_locale and encrypted.

Security-group picker (per-session, runtime, role-gated). A security-group is a permission bundle only — a tool_permission_groups id, driving tool visibility/approval — not a "mode" (no system-context injection; the RunContext.system_prompt substrate exists but is unused by the picker). The role carries the user's allowed set (default permission_group + attrs.permission_groups, §0.1); a new non-project session inherits the role's default group (sessions.rs::createrole_default_run_context). The chat surface switches it at runtime like the model pill: copilot.js renders a shield pill (hidden when ≤1 group) fed by GET /api/my/security-groups (the caller's role set, joined with group names; admin → all); selecting one sends the WS control message {type:"select_security_group", group} (chat-session.js::_selectGroup, twin of select_client). The server (ws.rs::handle_select_security_group_msg) validates against the role, persists it on chat_sessions.run_context, updates the live handler, and broadcasts ServerEvent::SecurityGroupSelected so every open tab re-syncs (the initial state is sent on WS connect). Enforcement is server-side via the shared run_context::validate_run_context_for_role (used by both the WS path and the REST set_session_run_context): a non-admin may only pick a group in its role's effective set (else 403), and every other RunContext field (system_prompt, allow_fs_writes/allow_fs_reads, working_directory) is discarded — closing an fs-escalation hole; admin passes through unchanged.

Selection is gated once; the persisted group is re-checked on every load. validate_run_context_for_role runs at selection time, and the result is persisted on chat_sessions.run_context — so on its own it let a group survive the role that granted it, indefinitely and across restarts (revoke ops from a role, and every session that had already picked it kept running on it). The fix is a second, narrower seam: run_context::reconcile_group_for_user, run by ChatSessionManager::get_or_create_handler on every handler build, which treats the stored group as advisory and degrades it when the owner's current role no longer allows it. Three properties are load-bearing: (a) it degrades to the role's default group (role_default_group, the same seam sessions.rs uses for a new session, so start-group and fallback-group cannot drift) — never to None, because a missing group means the catch-all default, whose rules are the fallback tier under every other group, so clearing widens; (b) it touches only security_group, unlike the selection path, so a project session's server-built project_root/system_prompt survive a permissions edit; (c) on uncertainty (unknown user, unreadable role, DB error) it leaves the stored group alone — guessing could only widen. The liveness half is Skald::revalidate_security_groups_for_{user,role}, called synchronously from the roles API (update) and the users API (role reassignment), which reconciles already-open handlers, persists, and emits SecurityGroupSelected so the pill re-syncs. Same rule as revocation: authorization is pushed, never left to the bus.

The role editor (roles-page.js) sets the default group + an allowed-groups checklist (→ attrs.permission_groups) + a default-assistant select (→ attrs.chat_agent) fed by GET /api/agents filtered to type:chat minus project-coordinator (source-driven); the same exclusion is enforced server-side in the roles API (validate_chat_agent).

File Element Notes
copilot.js <app-copilot> The chat surface (_wsSource='web'): full/dock roving layout, welcome hero empty state, privacy chip, composer with model pill, slash-command autocomplete
shared/chat-page.js <chat-page> Mobile chat (_wsSource='mobile')
copilot-render.js (helpers) renderMsg, renderTool, renderDiff, etc. — shared by copilot and chat-page
sidebar.js <app-sidebar> Nav sidebar; role-driven (ui_mode); inbox badge is live — the chat WS forwards the inbox lifecycle events (approval_requested/resolved, clarification_*, elicitation_*) regardless of source, chat-session.js re-dispatches them as the inbox-changed window event, and the sidebar (+ agent-inbox.js) refreshes on it; a 60 s poll remains as fallback
topbar.js <app-topbar> Top nav bar; per-user avatar color hashed from the username
dashboard-page.js <dashboard-page> #dashboard — status hero, LLM stats charts, pending inbox, quick guide
shared/file-viewer-base.js FileViewerBase (base) Shared file-viewer engine (fetch, kind detection, markdown/PDF/SVG/LaTeX, watcher, _renderBody); driven by _show/_hide. Extended by desktop + mobile
file-viewer-page.js <file-viewer-page> Desktop file viewer: FileViewerBase + hash routing via window.openFile(path)#file_viewer?path=...
shared/file-viewer-mobile.js <mobile-file-viewer-page> Mobile file viewer: FileViewerBase + prop-driven (visible/path), full-screen with back button
agents.js <agents-page> Agent discovery and config
agent-inbox.js <agent-inbox-page> Pending approvals + clarifications from background sessions
approval-rules.js <approval-rules-page> Approval rule management
cron-jobs.js <cron-jobs-page> Scheduled job management
connectors.js <connectors-page> MCP Connectors row list (one row per connector): user activate/deactivate + granted globals; admin also gets the Add connector dropdown (Marketplace / manual form at #connectors/new), per-row removal from the catalog, and the Sign-in providers modal (§7/§14/§15)
plugins-page.js <plugins-page> #plugins — user half: granted plugins + schema-driven per-user config form
plugin-catalog.js <plugin-catalog> #plugin-catalog — admin status board: one card per plugin (enable toggle + health dot + Configure → #plugin-detail)
plugin-detail.js <plugin-detail> #plugin-detail?id=<id> — one plugin's admin page: instance-config form (config_schema) + per-user access checklist (plugin twin of connector-detail.js)
plugin-page-host.js <plugin-page-host> Host for plugin-contributed pages (#plugin/<plugin_id>/<page_id>): dynamic-imports the fragment module, registers its element, mounts it with plugin-id
system-agents.js <system-agents-page> #system-agents — the caller's own run history for the background system agents (TIC): agent, start, status, duration, counters; row → the run's session
shared-folders.js <shared-folders-page> #shared-folders — admin-only CRUD for on-disk shared folders (§6): create/describe/delete + per-member read-only/read-write grants; description feeds the assistant's __SHARED_FOLDERS__ context
projects/ <projects-page> #projects — host + list + board; the board is tabbed (Files explorer with live watcher + write actions, Sharing members), deep-linked #projects/{id}[/sharing]. See the Projects section
connector-detail.js <connector-detail-page> A connector's own page (#connector?name=X): env/secret form + Test, the OAuth login panel (sign in → paste code → complete, §15), global enable + per-user access grants
shared/connector-common.js (helpers) Shared Connectors vocabulary: statusOf (incl. needs_login for a pending OAuth row), STATUS_LABEL, schema normalization, jf fetch
llm-providers.js <llm-providers-page> LLM provider management
models-hub.js <models-hub-page> Models hub landing (LLM / Transcription / Image)
models-llm.js <models-llm-section> LLM model CRUD + drag-and-drop priority
models-transcribe.js <models-transcribe-section> Transcription model CRUD
models-image.js <models-image-section> Image generation model CRUD
mobile-app.js <mobile-app> Mobile app shell
shared/settings-page.js <settings-page> Mobile settings: per-user avatar, locale picker (I18nMixin), profile/preferences