system agents: generalise the scheduler and add the two memory lints
Nightly Build / build (push) Successful in 7m14s
Nightly Build / build (push) Successful in 7m14s
Memory is kept as a maintained wiki, and a wiki nobody prunes rots. This adds
the scheduled maintenance pass, and generalises the machinery TIC had grown so
that a background agent is a trait impl rather than a loop of its own.
Two lint agents, not one. The private pass runs per user over `user-memory/`
and reports to them; the shared pass runs once over `shared-memory/`, where the
interesting defect is different — a note failing the table rule, i.e. private
business written where every member can read it. It names the note and the
category without repeating the content, since restating it spreads the very
thing being flagged. Both share `agents/common/memory-lint.md`.
Both are read-only, and that is enforced twice: the prompt says report-never-
repair, and `shared-memory/*` writes are already `@fs_write require`, so an
agent that tried to fix something would raise an approval card from an
unattended pass, which is auto-denied. Read-only is the only design that works
here, not merely the safe one.
One scheduler for cadences three orders of magnitude apart. TIC runs every few
minutes, a lint weekly — the case that tempts a second loop. It stays one
because the wake-up decides nothing: `base_tick` picks only how often to look,
and whether an agent runs for a user is `is_due` against persisted state.
Due-ness moves out of the run log into a new owner table, `system_agent_state`.
The two answer different questions: the run log skips idle ticks so it stays a
history rather than a heartbeat, while scheduling needs every attempt. Reading
due-ness off the log would re-run an idle agent on every tick and never bring a
weekly one due once its last productive run aged out. Persisting it is also
what makes a long interval survive a restart — an in-memory deadline is fine at
TIC's scale, but a weekly agent on a box rebooted every few days would have it
re-armed before it ever fired.
The shared store belongs to nobody, so `AgentScope::Instance` runs that pass as
the first unlocked admin. An ownerless run would write its trace into system.db,
which the runs endpoint shows to nobody by design, and its notify() would have
no recipient; attributing it to a user keeps the whole per-user surface working
unchanged.
Settings move to where the run log is. `ConfigSet` gains `owner`, so placement
is data on the set rather than a page that knows set names; the System agents
page grows one tab per agent holding its description, its settings (admin only)
and its runs — "why did this do nothing last night?" is half a schedule
question and half a log question. The form is shared with the Config page, and
writes still go through PUT /api/config/{key}.
Fixes an authorization gap found on the way: neither /api/config handler took
the caller into account, so any authenticated session could read and write
instance-wide config. The sidebar hiding the page is presentation, not access
control. Both are now admin-gated.
This commit is contained in:
@@ -30,6 +30,7 @@ pub mod scratchpad;
|
||||
pub mod shared_folders;
|
||||
pub mod sources;
|
||||
pub mod system_agent_runs;
|
||||
pub mod system_agent_state;
|
||||
pub mod tool_permission_groups;
|
||||
pub mod users;
|
||||
|
||||
@@ -935,6 +936,33 @@ pub async fn create_owner_tables(pool: &SqlitePool) -> Result<()> {
|
||||
.execute(pool)
|
||||
.await?;
|
||||
|
||||
// When each system agent last *attempted* a pass for this user — the
|
||||
// scheduler's state, deliberately kept apart from `system_agent_runs`.
|
||||
//
|
||||
// The two answer different questions and conflating them breaks both. The run
|
||||
// log is a history for the human: an idle tick writes nothing there, or it
|
||||
// degenerates into a heartbeat. Scheduling needs the opposite — every attempt,
|
||||
// productive or not — because "is this agent due?" is `now - last_attempt >=
|
||||
// interval`. Reading due-ness off the run log would re-run an idle agent on
|
||||
// every pass, and a weekly agent would never come due at all once its last
|
||||
// productive run aged out.
|
||||
//
|
||||
// Persisting it is what makes a long interval survive a restart. An in-memory
|
||||
// deadline is fine at TIC's scale — a few minutes, re-armed on boot — but a
|
||||
// weekly agent on a machine rebooted every few days would have its deadline
|
||||
// reset before it ever fired, and would simply never run.
|
||||
//
|
||||
// Owner table for the same reason as the run log: when an agent last ran for
|
||||
// someone is that person's activity, not the registry's.
|
||||
sqlx::query(
|
||||
"CREATE TABLE IF NOT EXISTS system_agent_state (
|
||||
agent_id TEXT PRIMARY KEY,
|
||||
last_attempt_at TEXT NOT NULL
|
||||
)",
|
||||
)
|
||||
.execute(pool)
|
||||
.await?;
|
||||
|
||||
sqlx::query(
|
||||
"CREATE TABLE IF NOT EXISTS mcp_events (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
|
||||
Reference in New Issue
Block a user