3f74dc26f2856f5e6dcab7956863a4fa9f0fe424
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
efb5b1dc33 |
feat: let an agent ask what its connectors are, instead of guessing
Nightly Build / build (push) Successful in 7m44s
An agent that wanted to know which MCP servers it had called `list_mcp_servers` — a tool that has never existed anywhere in this repo — and got "unknown tool". It was not a random hallucination: the prompt block says "the system prompt shows available servers", and `render_mcp_list` returned an empty string when nothing was connected. The model read a promise, found no table, and invented the discovery tool the text implied. The `mcp` kinds of `list_items`/`toggle_item` had been removed to close the §14 RCE vector, which was right for the write half and left no read half at all. So `list_items` gains `type: "mcp"` and returns the whole picture in one call, split into four buckets that each answer a different question: what is already loaded (call its tools directly), what is ready for `activate_tools`, what is installed but unusable and why, and what the user could still activate. Conflating the first two is what produced the original failure, so they stay apart. Every entry carries a derived note and a next step; when the step is a human one, it says so and names the UI page, because there is no tool for it. Read-only, and structurally so: `toggle_item` deliberately gains nothing, and the new `McpDirectory` trait exposes exactly one method. Enabling a connector from a tool is the thing §14 removed, and a wider seam here is how it would come back. Deny-by-default survives the report — an ungranted connector is not named at all, since a listing of what to ask for is itself a leak — except for a catalogue manager, who cannot administer what they cannot see. Three sources answer three questions and none is redundant: the registry says what exists and who may have it, the owner database says what was activated, and the live runtimes say what is connected right now — a row can read `ready` while its process is dead. The live half reaches the tool through the turn's extension map, alongside the pool and the fs view; with no live view the durable picture still renders, so freshness is an improvement and never a precondition. The static `__MCP_LIST__` table stays as it was, because it is frozen per conversation for prompt-cache stability. Its empty case now says so out loud and points at the tool. |
||
|
|
f900d803f2 |
fix: one tool-set recipe per session, so a tool cannot vanish between rounds
Nightly Build / build (push) Successful in 7m36s
Two "unknown tool (not in this turn's tool set)" failures, one disease: the turn's tool set was rebuilt from a different recipe depending on which entry point happened to drive it. A sub-agent got `ask_user_clarification`, `execute_subtask` and `activate_tools` and nothing else — while `agents/common/tools.md` and every reporting agent's prompt tell it to register its output with `update_scratchpad`. The child could see the scratchpad injected into its context but had no way to write to it. It now gets the scratchpad and todos tools, on the parent's `scratchpad_sid`: one blackboard per session, as the surrounding code already declared. `show_file_to_user` was injected per message by the WS handler, while `resume_session` and `resolve_pending_call` rebuilt the list with `execute_task` alone. So approving a card, or reconnecting mid-turn, continued the *same* conversation with the tool silently gone. There is now a single recipe, `ChatHub::session_interface_tools`, used by all three paths and fed by a builder the shell installs once through `Skald::set_interface_tools_builder`: the core keeps owning the tool, the shell keeps owning the policy of who gets it — Telegram still does not, since it cannot act on OpenFile. |
||
|
|
6cb4ea0ce8 |
feat: let the fs-tools reach the whole container, and stop rebuilding the system prefix every round
Nightly Build / build (push) Successful in 7m33s
Two changes to what a turn costs and what it can see.
## The system prefix is frozen per conversation
`AgentSystemContext::system_context` is called once per round and reassembled
`base` from disk and SQLite each time, so an agent writing `user-memory/index.md`
in round 3 turned round 4 — seconds later, with the provider cache certainly
warm — into a full miss. `base` is the head of every provider's cache key, so it
is the most expensive string in the request to touch.
`PrefixCache` builds it once per (conversation, agent) and holds it on
`UserLoopRuntime`. The refresh rule is the only free one: rebuild once the
conversation has been idle longer than a provider's cache could survive
(20 min). The clock is idle time of the conversation, not time since a file
changed, and reading restarts it — every get is a request about to go out.
Writes are deliberately not reacted to. The agent's own edits are already in the
context, two messages downstream. A write from elsewhere is invisible until the
TTL: that is precisely where an immediate rebuild costs the most, and the
cheaper freshness path already exists — a `read_file` result appends, and
appending invalidates nothing. The injection header now says so.
Also removes 4 DB queries and 2 file reads per round.
## The security boundary is the container, not the mounted subtree
`read_file /tmp/cv.txt` answered "path escapes your workspace" and the agent
re-read the file with `cat`. It was right to refuse — /tmp exists only inside
the container — but the refusal protected nothing: `execute_cmd` already runs
there with passwordless sudo. The mount is the fast path, not the perimeter.
`resolve_target` now routes an absolute path through `container_to_agent`
first. Landing on a mount takes the host path, which also fixes a real bug:
`/root/x` IS `~/x`, yet every tool rejected it, because `PathBuf::join` with an
absolute tail discards the base and the result then failed the prefix check
(`/root/shared/{X}/…` too). Landing nowhere means container-only, served by the
new `container::exec_fs` over `docker exec`, with paths passed positionally so
a path containing `$(…)` stays data. Membership still holds: `/root/shared/{X}`
for a non-member fails exactly as `shared/{X}` does.
Single-file tools get this without a second implementation: `fs::Shuttle` pulls
the file out, runs the unchanged tool on the copy, and pushes it back if the
bytes changed. `list_files` lists in place, `read_file` reads container paths as
text (a shuttled copy cannot back a MediaRef), and `grep_files` refuses them
with a pointer to `rg` rather than approximating its own semantics. The viewer
follows the same routing, so the user can open what the agent read.
Host containment is untouched and still guards every mounted path — it is the
defence against a symlink planted in the container pointing at the host's /etc,
and the container branch never touches the host filesystem at all.
Verified end-to-end against a live skald-runtime:v3 container: write/read/edit
on /tmp round-trip, /etc/os-release reads, binary and shell-metacharacter paths
survive, and /root/notes.md lands in the host home.
|
||
|
|
da0830aefa |
feat: an hour-precision clock that says so, and a cron tool that names the real timezone
Nightly Build / build (push) Successful in 7m30s
The datetime block claimed second precision it never had. It is built once per request and a turn can run for minutes, so `17:54:31` is a lie by the time the model reads it — and the model, having no way to know, wrote cron expressions from it. Rounding was already there but configurable (`round_minutes`, shipped at 60) and justified by the prompt cache. That justification was false: the block is the LAST system message, after the whole conversation, so the cached prefix is identical from turn to turn whatever the timestamp says. Rounding buys nothing for caching today. So the knob goes and the granularity becomes part of the contract: always truncated to the hour, stated in words, with a pointer to `date` for the cases that need the minute. `DatetimeConfig` keeps only `enabled`. - truncation happens in the DISPLAYED zone, not on the UTC epoch: +05:30 zones would otherwise render 20:30 — an hour off and not on an hour boundary, which reads as precise again. - the weekday is spelled out. "next Tuesday" is a far more common ask than the minute, and weekday-from-date is exactly the arithmetic models get wrong. Also fixes a real bug found on the way: `execute_task` told the model, twice, that cron expressions are evaluated in Europe/London — hardcoded, while TaskManager uses the configured timezone. On a non-UK box every scheduled job was written against the wrong clock. The description now names the zone the scheduler actually uses (`TaskManager::timezone_name`), and the assistant's AGENT.md stops repeating the literal. CLAUDE.md: record that the instance is in production. The greenfield licence has expired — schema changes need a versioning mechanism, and per-user SQLCipher files mean it cannot be a boot-time sweep. |
||
|
|
baf68878e4 |
fix: stop shrinking conversations behind the user's back — both automatic context guards ship off
Nightly Build / build (push) Successful in 7m35s
The shipped default combined a sliding history window with no compaction, which is the worse of the two available trades in both directions it is measured on. `max_history_messages: 30` is a sliding tail window (`projection::window` — `drain(..len - max)`). Past 30 messages it drops from the head on *every* turn, so the prompt prefix changes on every single request and every provider that caches one (Anthropic breakpoints, OpenAI automatic prefix caching) misses every time. It also drops those messages with no summary standing in for them: silent amnesia, not just a cold cache. Compaction rewrites the prefix once per compaction and leaves a summary behind — yet it was the half that was commented out, while the window's own doc-comment already said the two were exclusive. Both are now `Option` and both ship unset, so nothing shrinks a conversation unless a human types `/compact`. Which surfaced the real bug: `/compact` did not work either. The compactor was `Option<Arc<ContextCompactor>>` keyed on the config section existing, so commenting out `compaction:` disabled the manual command too — `force_compact` returned `Ok(false)` and the chat answered "compaction disabled". Manual compaction is a command a user types; it cannot depend on an admin having filled in a token threshold. The compactor is now built unconditionally and `threshold_tokens: Option<u32>` arms only the automatic pass; `try_compact` early-returns without it, `force_compact` deliberately never consults it. The projection accordingly yields to the *automatic* pass rather than to the compactor's existence (`LoopConfig.auto_compaction_enabled`), so a configured message cap is not silently voided by `/compact` merely being available. `CompactionConfig::Default` is hand-written for the same reason `RoleAttrs`'s is: a derived one gives `keep_recent: 0`, which would compact away every recent message on any box omitting the section — now the default. Also fixes two documentation bugs in the same file: `event_triage` was documented nested under `llm:`, where it parses fine and is then silently ignored (it is a top-level field), and `datetime` was documented twice with conflicting examples. A new test asserts the shipped default actually deserializes and that both guards are off — a field the default omits must be genuinely optional, or a brand-new install fails to boot. Automatic compaction returns later, triggered off the resolved model's own context window instead of a hand-tuned token count that cannot know which model is answering. |
||
|
|
4f10528368 |
feat: conversation review — a nightly report on a supervised person's conversations
Nightly Build / build (push) Successful in 7m40s
The first AgentScope::PerSubject system agent, and the reason that scope exists. Once a night, for each person with a supervision edge, it reads every message that person and the assistant exchanged since the previous review — across all their conversations — and writes one report for the people who supervise them. Schema (all registry except reports): - supervision(subject_user_id, supervisor_user_id): the generic §0.1 edge, answering both 'whom does a background agent look at' and 'who may read what it produced', with real FKs so deleting a user cascades both ways - system_agent_coverage(agent_id, subject_user_id, covered_through): the per-subject watermark that makes 'everything since last time' a window — neither system_agent_runs (history for humans) nor system_agent_state (advances before the work), and advanced only on a completed pass so a crash re-covers instead of skipping - reports (owner schema, the second two-homes table after memory_docs): instance rows land in system.db, deliberately cleartext to the box owner, who is the intended reader (§2); the subject cannot see them structurally The pass reads the subject's database inside a supervisor's runtime, so the ephemeral session and run row land in the watcher's file; iteration is over subjects, so two parents watching one child get one review; and the subject need not be logged in when their space is unencrypted — via the new UserManager::open_unencrypted, which refuses an encrypted user outright (no key to be had) and never registers the pool as unlocked. The agent declares the new AgentMeta flag allow_tools: false, so its turn gets an empty tool registry — nothing for a prompt injection in the transcript to call — and produces its report as its final assistant message, read back from chat_history and parsed (NOTHING_TO_REPORT sentinel, no row on quiet days). chat_history::conversation_window is the transcript query; its four filters (non-ephemeral, depth 0, non-synthetic, non-empty) each guard a specific way the review would otherwise be wrong, and tool calls are absent by construction. Cadence is Run at (hour) rather than Interval — 4am local by default — with due-ness answered inside has_work against the coverage watermark, so a machine off for three days covers the whole stretch in one pass. Reports announce ReportCreated on the system bus (no subscriber yet). run_ephemeral_turn gains a per-pass system_substitutions map, which the review uses to hand the model the subject's profile under __SUBJECT_PROFILE__ — the system-context substitutions describe the session owner, the wrong person here. docs/system-agents.md gains the conversation review section; CLAUDE.md documents the scope, the tables and the tool-less design. |
||
|
|
6f35c53d93 |
activate_tools: diagnose a group instead of pretending it activated
A group name that resolved to no running MCP server was granted anyway, persisted in `activated_tools`, and reported as a success with "registered but not yet running — tools will appear after reconnect". Every part of that was false: nothing registers connectors anymore (the agent-facing `register_mcp` went away with §14), no reconnect will ever produce the tools, and the model — believing it had succeeded — called `mcp__x__…` a round later and failed there instead of here. The junk grant row stayed in the session forever, resolving to zero tool defs on every projection. `SkaldToolActivator` now resolves first and acts only on what resolved, walking the connector states in order: the built-in `config` group, then the servers running in this user's view, then `mcp_user_servers` (owner pool), then `mcp_global_servers` + `mcp_global_access`, then `mcp_catalog` + `mcp_catalog_access` (registry pool), then unknown. Only `activated` touches the in-memory grant set and the DB; everything else leaves no trace at all. The result is a JSON object keyed by group name — `status` (activated / needs_login / not_activated / not_authorized / unavailable / unknown), `tool_prefix`, `tool_count`, `description`, `message` — with the same shape whether the call succeeded or not, so the model parses one thing rather than prose. `message` is written to be relayed to a non-technical user and says who can fix it: the user in Connectors, or the admin. When no group at all activates, the tool fails with that same JSON, so the model reports the diagnosis instead of proceeding. A diagnosis query that errors is logged and falls through to the next candidate — a broken lookup must never become a false claim about a connector. The activator needs the registry pool, the user id and the config defs; both construction sites are updated, so a sub-agent gets the same diagnosis as the root agent. Removes `tools/activate_tools.rs` and its interface-tool registration: it was unreachable (`SkaldToolSet::find` prefers natives, and `NATIVE_NAMES` explicitly drops the legacy interface tool of that name) but carried the same wrong string, waiting to be fixed twice. |
||
|
|
73c720e9ef |
llm: restore request logging lost in the agent-loop migration
Nightly Build / build (push) Successful in 6m50s
The LLM-requests page has been empty since
|
||
|
|
a3e1b0add0 |
memory: reshape the two stores into a maintained wiki
Nightly Build / build (push) Successful in 6m50s
Memory was a scrapbook: notes accumulated, nothing kept them consistent, and shared memory had no rule saying what belonged in it. This adopts the LLM-wiki pattern — the assistant maintains an evolving artifact rather than re-deriving knowledge each session. The schema (agents/common/memory-wiki.md, included by the three type:chat agents) adds an append-only log.md beside each index.md, names Ingest / Recall / Lint as habits, and states the rule for shared memory: write it there only if you would say it out loud with every member in the room. One person's health, results or another member's opinion of them stays private — that is what shared folders, not shared memory, are for. Tampering is the reason the rules are shaped this way. Shared facts carry provenance and are superseded rather than erased, and a member who contradicts a fact they did not write gets a logged CLAIM instead of an overwrite: only the originator or an admin can turn it into a change. The approval gate cannot enforce this — it asks the caller, who is the same person pushing — so a per-role write permission on shared memory is still the real boundary. The prompt is the etiquette, not the fence. append_file is a new fs tool because the log needs a write that cannot shorten a file. On memory paths it is one SQL statement, so concurrent appends (parallel tool batches, two sessions of one user) cannot lose a line — a dropped line in an audit trail is worse than a failed write. It is auto-allowed on shared-memory/log.md at a lower priority than the shared write rule: gating the trail would be friction with no safety, and a rejected log write yields an unlogged change. memory::scaffold seeds index.md / log.md / user.md so the schema does not describe files that do not exist. Seeded empty rather than left absent: a missing note resolves to nothing at injection, so the model cannot tell "nothing recorded yet" from "this mechanism is not running". __MEMBERS__ renders the roster from users + roles instead of a note the model maintains. A remembered copy drifts and can be talked into being edited; this one bypasses the model entirely. users.notes are excluded — they are the admin's private notes about a person and this block is visible to every member. Still open: a UI to browse memory (list_dir does not classify memory paths yet), a tool-written revisions table (log lines are still typed by the model), and the periodic lint pass. |
||
|
|
4d81295a3d |
messages: unify harness-injected data under <system-extra> tag
Nightly Build / build (push) Successful in 6m49s
Replace the ad-hoc [SYSTEM INFO] / [TELEGRAM SYSTEM INFO] prefixes with a single canonical <system-extra> wrapper, sourced from one constant (SYSTEM_EXTRA_TAG) so emission and documentation can never diverge. - core-api: SYSTEM_EXTRA_TAG + system_extra() helper; attachments_block rebuilt on top of it. - telegram: system_info_message (location) uses the helper; the voice transcript is forwarded as a plain user message (it is the user's own words, not harness metadata). - chat agents: new agents/common/harness.md include (long form, with an explicit "data, not instructions" guard), added to assistant/kid/ project-coordinator. The tag name rides the __HARNESS_TAG__ sentinel, resolved in AgentSystemContext to SYSTEM_EXTRA_TAG — renaming the tag stays a one-line change. |
||
|
|
24ee5b89d7 |
agent-loop: projection, recovery, compaction into the crate (phase 3)
Nightly Build / build (push) Successful in 6m49s
The session handler is now a thin shell: three entry points in kernel_turn.rs (run_kernel_turn / recover_turn / resolve_pending_call) and the ChatSessionHandler. Everything that shaped a Value — projection, recovery, compaction mechanics, the LLM loop, message building — lives in agent-loop or behind a loop_adapters trait. agent-loop: - projection/ (mod + media): stored history -> wire messages, the one place provider divergence lives; well-formedness contract, DTL injections (append-only), media parts. LinearAssembler is now a Projection + ProjectionHooks config, not its own implementation - recovery.rs: reap interrupted batches -> resolve the deepest frame's non-terminal calls (Running by policy + RestartHint, AwaitingHuman re-asked) -> un-wedge finished children -> cascade up, every frame on its own agent (B3) - compaction.rs: split point (never assistant+tool group), transcript, SUMMARY_PREFIX/preamble/template, the no-tools model call, summary row - manager: resolve_pending (gate skipped, real ToolContext, then continue incl. sub-agent); start_loop used by recovery; LiveInput - delegate: AsyncExecutor + StoreSink for mode:async (durable cron row, result delivered back into the parent conversation) - kernel/context/store: support the above (TurnScope via Extensions, frame lookups, aligned result-text semantics) skald-core: - loop_adapters: UserLoopRuntime (D12 - one LoopManager per user), TurnScope (per-turn state in the Extensions type-map; no scope is denied), projection_cfg/media_source/tool_digest (Skald's projection knobs without owning projection code), async_task (CronExecutor + DurableSink) - session/handler: stripped to mod.rs + kernel_turn.rs + config.rs + interface_tools.rs + media.rs; deleted agent_dispatch, approval, dispatch, emitter, gate, llm_call, llm_loop, message_builder, messages, outcome, resume - compactor.rs: policy only (threshold, model pick, CompactionEvent); mechanics are the crate's CLAUDE.md updated (recovery, compaction, sub-agents, approval gate, projection sections now describe the crate-owned flow). |
||
|
|
3fca7867fa | agent-loop: drop leftover empty wiring module (phase 2) | ||
|
|
0297fe71bd |
agent-loop: root turn driven by the library kernel (phase 2)
ChatSessionHandler now runs the root turn on the agent-loop kernel instead of run_agent_turn; sub-agents follow on the same kernel via DelegateTool. The old loop stays for resume/recovery until phase 3. agent-loop: - DelegateTool + AgentCatalog/AgentProfile (full toolset override, per-child selector/assembler, frame-scoped get), StaticCatalog, FilteredToolSet; sync flow with sticky child_token; batch via the generic fan-out - manager: start_loop skips the registry (children are not double-driving; they ride the parent's token tree); LoopParams gains selector/token overrides - store: get_frame, get_call, set_call_extras; HistoryStore result text aligned to raw-stored semantics (projection formats) - events: ApprovalRequired.request_id, AgentSpawned/Finished parent info; AskUserTool with_name + suggested_answers alias + Question.frame skald-core (loop_adapters + handler): - SkaldAssembler (byte-parity port of MessageBuilder's projection: scratchpad/summary/window, DTL Kimi/Anthropic injection, media, user coalescing, reasoning echo) + AgentSystemContext (prompt layers, substitutions, MCP list, shared folders, user profile) - SkaldAgentCatalog (build_sub_agent_config port), SkaldHumanChannel, scratchpad/todos tools, execute_task sync/async alias, LegacyInterfaceTool, PendingLiveInput - ApprovalGate: PendingWrite diffs via LoopEvent::Host (memory/disk routed like the fs-tools); SkaldWritePreviewHook for executed-write diffs; EventTranslator LoopEvent→ServerEvent (display meta, preview, FileChanged, AgentStart/Done, root-only Done/Truncated/Cancelled) - handle_message: builds TurnParams and drives the kernel; resume of pending tools runs first (results belong to the previous turn); ChatEvent publication stays handler-side; /stop cancels the live loop - ToolRegistry.get_tool/all_tools; def builders made pub(crate) Full workspace suite green (179 skald-core, 34 agent-loop, adapters incl.); two pre-existing doc-test failures fixed along the way. |
||
|
|
d50abbb0fa |
agent-loop: Skald adapters behind the crate traits (phase 1)
New skald-core::loop_adapters module — implements the agent-loop trait
surface over existing infrastructure, unused by the current loop (wired
in phase 2):
- SqliteHistory: HistoryStore over chat_sessions_stack/chat_history/
chat_llm_tools/chat_summaries, no schema change; CallState maps 1:1 on
the existing status strings; wire call ids synthesized as tc_{id}
- SkaldSelector: ModelSelector over LlmManager with the agent's strength
captured per-turn (D14); DtlMode → ToolRendering mapping (D15)
- SkaldActivationSource + SkaldToolActivator: DTL catalog + persistence
(activated_tools, anchored at the triggering message) behind the
crate's protocol traits; unifies the grants/persistence split
- ApprovalGate: port of run_approval_gate (pre-approved, engine, fs
fast-path, auto-deny, AwaitingHuman + block on human); a closed human
channel maps to the new GateDecision::Suspend in agent-loop
- SkaldToolSet + CoreToolBridge/McpToolBridge: core-api and MCP tools
run inside the crate's kernel (execution bridged, execute_cmd keeps
its teardown; D7 MarkInterrupted for shell)
- agent-loop: re-export async_trait at root; EventSink::new made public
17 adapter tests green (temp-DB integration); full workspace suite green
(pre-existing honcho-client doc-test failure untouched: missing dev-deps).
|