Introduction
Glomeris is an evidence-first, policy-constrained developer storage autopilot
for macOS. When a machine enters disk pressure, Glomeris discovers meaningful
reclaimable storage (Cargo target dirs, node_modules, Homebrew’s cache,
Xcode DerivedData, Docker’s reported build cache), explains why a resource is
or is not safe to remove, and executes only policy-approved cleanup actions —
re-measuring actual freed bytes rather than trusting an estimate.
The canonical safety invariant
AI can recommend. Policy decides. Executor verifies. Filesystem reality wins.
Concretely, as implemented today:
- An optional BYOK LLM planner (
glomeris::actions::llm) may rank and explain candidates, but it can only select from a closed, typed set of pre-registered action IDs and real resource IDs the crate already found — it can never compose a raw command or an arbitrary path (see BYOK LLM Planner). - A single deterministic module (
glomeris::policy::engine::classify) decidesAUTO_SAFE/ASK/PROTECTEDfor every resource. Nothing upstream of it, including any LLM output, can bypass that decision (see Safety Model). - The executor (
glomeris::executor::execute) never trusts a previously computed decision at face value: it re-collects evidence and re-runsclassifyimmediately before mutating anything, and aborts rather than acts if the fresh read disagrees with what was approved. - Every execution re-measures actual reclaimed bytes after the fact rather than assuming the estimate was correct.
What Glomeris is not
- Not cross-platform. The MVP targets macOS only. There is no Windows or
Linux support, and the
platform::macosmodule is compiled out entirely on other operating systems. - Not a cloud service. There is no backend, no telemetry, and no account
system. Everything runs locally: a CLI binary, an optional per-user
launchdagent, and an optional menu-bar app. The one thing that can leave your machine is a BYOK LLM request you configure yourself, to an endpoint you name — bounded and previewable (see BYOK LLM Planner). - Not a generic system optimizer. Glomeris only understands a fixed, named set of developer-tool-owned resource kinds (Cargo, npm/pnpm/yarn, Homebrew, Xcode, Docker). It does not attempt to clean arbitrary “junk” files or ordinary user documents.
- Not an automatic deleter of ordinary user documents. Every resource
kind Glomeris knows about is a build/tool cache with a defined owning
tool; an unrecognized resource kind (
ResourceKind::Unknown) is classifiedPROTECTEDunconditionally rather than falling through to any default treated as safe. - Not a GUI with authority of its own. There is a SwiftUI menu-bar app
(see Menu Bar App) — an earlier version of this page said
there was not. It is a thin client: it shells out to the same
glomerisbinary and renders what comes back. It classifies nothing and decides nothing; a CI guard (scripts/check-no-policy-label-branching.sh) fails the build if Swift code branches on a policy label, so the one way the GUI could quietly grow its own safety opinion is checked rather than trusted.
Two meanings of “autopilot”
Both are used in this book, so they are worth separating once:
- The product is a storage autopilot in the sense above — it discovers, explains and verifies on its own rather than asking you to audit paths by hand. That is what the first paragraph means.
glomeris autopilotis one specific command: an unattended run inside a policy envelope you set on the command line — byte, action, time and resource-kind limits,AUTO_SAFEonly unless you pre-authorise otherwise. See Autopilot.
glomeris autopilot run is the only thing that deletes unattended.
detect, explain, clean, execute, free and emergency all need you to
invoke them. The optional launchd agent does run unattended, but it only
polls disk pressure and notifies — it executes no action and never touches the
filesystem it is watching (see Daemon Lifecycle).
Project status
This is an experimental MVP, not a stable release — see Known Limitations for a specific, current accounting of what is and is not implemented.
The command set described here is held to the code mechanically:
tests/help_golden.rs pins every help surface byte for byte, and
scripts/check-docs-cover-cli-commands.sh fails CI if a command ships without
a section in CLI Reference. Prose can still go stale, which
is what Known Limitations is for.
Installation
Prebuilt release archives (recommended)
Every tagged release is built and published automatically by
cargo-dist
(.github/workflows/release.yml), which attaches the binary tarballs
directly to each GitHub
Release with checksums.
Download the one matching your Mac’s architecture, extract it, and copy the
glomeris binary onto your PATH (e.g. /usr/local/bin).
Two macOS targets are built for every release: aarch64-apple-darwin
(Apple Silicon) and x86_64-apple-darwin (Intel, cross-compiled) — see
dist-workspace.toml. Pick by architecture; nothing selects it for you.
Homebrew tap
brew tap Chisanan232/tap
brew trust --tap chisanan232/tap # Homebrew 7 and later only; see below
brew install glomeris
brew upgrade glomeris picks up later releases the same way, and Homebrew
selects the right architecture automatically.
The tap is Chisanan232/homebrew-tap,
and dist-workspace.toml points cargo-dist at it so each release can publish
a regenerated formula there (HORO-1069, HORO-1320).
Three things about this path are worth knowing before you use it.
Homebrew 7 will not load a third-party tap until you trust it. On Homebrew
7.0 and later, brew install glomeris straight after brew tap fails with
Refusing to load formula chisanan232/tap/glomeris from untrusted tap. That is
Homebrew protecting you from arbitrary Ruby in a tap you have not vouched for,
not a broken formula. brew trust --tap chisanan232/tap records the decision in
~/.homebrew/trust.json (or under $XDG_CONFIG_HOME/homebrew/) and only needs
doing once. Older Homebrew versions have no brew trust and do not need this
step.
A manually copied binary already on your PATH blocks the symlink. If you
previously followed the prebuilt-archive instructions above and copied
glomeris into /opt/homebrew/bin or /usr/local/bin, Homebrew installs into
its Cellar but refuses to link over your file, and warns that the Homebrew copy
is shadowed. Remove your manual copy, or let Homebrew take ownership:
brew link --overwrite glomeris --dry-run # lists exactly what it would remove
brew link --overwrite glomeris
The formula currently tracks v0.2.0 and is updated by hand. The release
workflow’s publish-homebrew-formula job needs a credential for the tap
repository that does not exist yet, so no release has regenerated the formula
automatically. That credential is the remaining half of HORO-1320. Until it is
configured, the tap can lag the newest tagged release — check the
releases page if you need
the very latest, or use a prebuilt archive above.
Build from source
Glomeris requires Rust/Cargo (2021 edition). There is no other runtime dependency.
git clone https://github.com/Chisanan232/glomeris.git
cd glomeris
cargo build --release
The resulting binary is at target/release/glomeris. Copy it onto your
PATH (e.g. /usr/local/bin) if you want to run it as glomeris directly.
Installing the menu-bar app
See Menu Bar App for what GlomerisMenuBar.app actually
shows and lets you do once it’s installed.
The app runs the glomeris CLI for everything it displays. A release build
ships its own copy inside the bundle, so it works with no separate CLI
install; otherwise it looks along PATH and then in /opt/homebrew/bin and
/usr/local/bin, which covers both Homebrew prefixes and the manual-copy
instructions above. Menu Bar
App documents the exact
order and why PATH alone is not enough for a GUI app.
From the next tagged release onward, GlomerisMenuBar.app.zip is published
as an extra asset on that GitHub
Release alongside the CLI
tarball, by a separate macos-app-release.yml workflow that fires once the
release is published. That workflow landed after the current latest release
was tagged, so releases published so far carry only the CLI tarballs — to
get the app before the next tag, build
macos/GlomerisMenuBar/GlomerisMenuBar.xcodeproj yourself. It’s built and
ad-hoc signed automatically
(codesign --force --deep --sign -, identity -) — this seals the bundle
well enough to run, but it is not signed with an Apple Developer ID and
not notarized. Developer ID notarization is tracked as a future
follow-up, not done today.
That matters the moment you download the zip through a browser: macOS tags
the extracted app with a “quarantine” flag, and Gatekeeper will refuse to
just open it — double-clicking shows “Apple could not verify that
GlomerisMenuBar is free of malware” with no direct way to proceed. This
is expected for an ad-hoc-signed app and does not mean the download is
broken or unsafe; it’s the same warning any non-notarized app gets. Follow
these one-time steps to open it:
- Unzip
GlomerisMenuBar.app.zipand moveGlomerisMenuBar.appwherever you want to keep it (e.g./Applications). - Double-click it. You’ll see the “could not verify” warning — click Done (or Cancel) to dismiss it. This first attempt is expected to fail; it’s what unlocks the next step.
- Open System Settings → Privacy & Security, scroll down to the Security section, and you’ll see a line like “GlomerisMenuBar” was blocked to protect your Mac with an Open Anyway button next to it. Click it, then authenticate with your password/Touch ID when prompted.
- Double-click
GlomerisMenuBar.appagain. A dialog reappears asking if you’re sure — click Open. macOS remembers this decision, so every later launch works with a plain double-click, no repeat of these steps.
This System Settings path is Apple’s own documented method for opening software from an unidentified developer and is the one to use on current macOS (Sequoia and later, including Tahoe) — recent macOS versions have locked down the older shortcut of control-clicking the app and choosing Open, so that shortcut may not offer an Open option at all anymore. If it does work for you, it’s faster: control-click the app, choose Open from the menu, then click Open in the dialog that appears — but treat the System Settings steps above as the reliable path.
Terminal fallback (last resort). If neither of the above surfaces an Open Anyway/Open option — for example the security exception was already dismissed, or you’re scripting the install — remove the quarantine flag directly:
xattr -d com.apple.quarantine /path/to/GlomerisMenuBar.app
Only run this against an app you actually intend to trust: it deletes the flag that tells Gatekeeper to check the app at all, so use it deliberately, not as a routine habit.
Platform support
Glomeris targets macOS only. daemon, emergency, and free are gated on
#[cfg(target_os = "macos")] and print an error and exit non-zero on any
other OS; scan and detect do not have this restriction, since they only
touch std::fs/std::process and don’t depend on the macOS-only
platform::macos module.
Verifying the build
cargo test
cargo clippy --all-targets -- -D warnings
Both are part of this repository’s CI (.github/workflows/ci.yml, macOS
runner) and should pass on a clean checkout.
Menu Bar App
GlomerisMenuBar.app is a native SwiftUI menu-bar app (Epic HORO-1043,
“Glomeris MVP 2.0”) that gives the glomeris CLI a graphical front end. See
Installation for how to
download and open it.
Thin client, not a second decision-maker
This is the project’s core invariant for the app, stated verbatim in the
Swift source (macos/GlomerisMenuBar/Sources/GlomerisMenuBarApp.swift):
GlomerisMenuBar is a THIN CLIENT ONLY, over the existing Rust
glomerisCLI’s--jsonoutput. This Swift codebase renders and lets a human approve/decline what the Rust CLI already decided.NO policy classification (
AUTO_SAFE/ASK/PROTECTED), NO evidence correlation, NO action planning, and NO filesystem execution logic may EVER live here.
In practice this means every screen below is a formatting step over one of
glomeris’s own --json reports (status, daemon status, detect,
explain, history, actions history, execute, autopilot show|enable|revoke), spawned as a subprocess. The app never re-derives “is
this safe to clean” from
policy_label or reasons — it reads the already-computed executable,
offered_actions, and refusal_reason fields the CLI provides for exactly
this purpose (see CLI Reference). If you trust the CLI’s
policy decisions on the command line, you’re trusting the exact same
decisions in the menu bar — the app has no separate opinion.
How the popover reads
Every section is a titled card in a scrolling column, ordered by the questions you open the panel with: what the disk is doing now, what could be reclaimed, what has already happened. Three presentation rules hold across all of them:
- Plain language first, the CLI’s own token second.
PRESSUREDreads as “Running low”,aborted_by_revalidationas “Stopped safely” — and the raw token stays visible beside or beneath it, because this is an evidence-first product and you have to be able to match what the panel says against--jsonand against the docs. - Three separate axes, never merged. Storage impact (how much space),
safety (what the policy allows) and evidence quality (how sure the CLI
is) are shown as distinct badges. A large
AUTO_SAFEcandidate is an opportunity, not a hazard; a smallPROTECTEDone is still protected. Size is therefore drawn in a neutral tone at every magnitude. - Colour is never the only signal. Every badge carries a symbol and a
word as well as a tint, so the state survives greyscale, a
colour-vision deficiency and a tinted wallpaper. A refusal
(
PROTECTED) is deliberately not painted like a failure: it is the policy working, and there is nothing for you to fix.
Disk space, and Background monitor
Two cards, polled on appear and every 10 seconds while the popover is open. The two facts are deliberately never collapsed into a single “healthy” indicator:
- Current disk pressure (
glomeris status --json) — used percent, free space, and pressure state (one ofHEALTHY,WARN,PRESSURED,CRITICAL,EMERGENCY). - Daemon health (
glomeris daemon status --json) — whether thelaunchdagent is loaded, and how long ago the poll loop last recorded a heartbeat ("Last heartbeat: 42s ago", or"no heartbeat recorded").
A launchd-loaded-but-wedged daemon and an actually-polling one stay
visibly distinguishable — the app never merges loaded and
heartbeat_age_secs into one boolean.
Reclaimable space
Shows the most recent glomeris detect --json scan (“Last scanned:
<time>”) plus a Refresh button. Nothing is scanned automatically:
there is no appear-triggered scan, no timer, and no background polling
loop for candidates — detect runs only when you explicitly tap Refresh,
streaming live per-detector progress via --progress-json while it works
(button label switches to “Scanning…”).
Five states that are easy to conflate are kept distinct, because each is a different claim about your disk:
| State | What it says |
|---|---|
| “No scan yet” | Nothing has been looked at. Not a clean bill of health. |
| “Nothing worth reclaiming” | Scanned, and there is genuinely nothing — good news. |
| “Nothing found where Glomeris could look” | Scanned, found nothing, but part of the search never answered — so whatever is there is unknown rather than absent. Names which checks did not finish. |
| “No candidates match this filter” | Things were found; you are just not looking at them. |
| A scan failure | Says what failed. An empty list is never shown in its place. |
The third state also has a non-empty counterpart (HORO-1484): when a detector
fails but others still found candidates, the rows are shown as normal with
“This list may be incomplete” added below them. The rows are real and stay
on screen; what they may not do is look like the whole account. A detector
whose tool is simply not installed is not a failure and produces neither
message — docker being absent is normal, brew being asked and failing is
not.
Each row shows the candidate’s kind, what cleaning it would free, and its safety verdict in words. Tapping a row opens its detail view — there is no inline “Clean” button in this list.
Rows arrive in the CLI’s own order — biggest reclaimable size first — and
are shown exactly as detect returned them. The app does no ranking of its
own: deciding what matters most is a judgment, and it belongs next to the
evidence in Rust rather than in a thin client that would then drift from it.
The View options menu (the funnel next to Refresh) offers two things:
- Order — “Biggest first”, which is the CLI’s order untouched, or “Path (A–Z)”, an alphabetical index for finding a resource whose path you already know. There is deliberately no third “largest first” option: that is “Biggest first”.
- Show — All, Safe to reclaim, Asks first, Protected, or Not enough evidence. This hides rows and does nothing else; it cannot enable, authorise or perform anything, and the list has no action affordance for it to unlock. Protected items are a first-class filter value rather than something hidden by default: “what on this machine is off-limits, and why” is a reasonable question, and quietly omitting them would teach you that protection means invisibility.
When a filter is active the header reads “N of M items” so a shortened list can never be mistaken for a smaller problem.
Size is not safety
A third badge appears on the rows worth pausing on — “Biggest wins” or
“Worth a look” — using the impact_tier the CLI already computed. It is an
emphasis hint about magnitude only, carries no safety colour, and appears on
no more rows than deserve it (ordinary and unmeasured candidates get no badge
rather than a badge saying “normal”, which would be noise).
Storage impact, safety, and evidence confidence are three separate axes and the popover never collapses them. A large candidate may be protected; a small one may be perfectly safe to reclaim. Nothing is distinguished by colour alone — every badge pairs its colour with a distinct symbol and words — and VoiceOver reads each row as size, then safety, then emphasis, so the badge is never the only route to the fact.
Candidate detail
Opened by tapping a candidate row; this is the sole place a “Clean” button
exists in the whole app. It calls glomeris explain --json for the exact
resource, which supplies fingerprint_token — the same token
execute --observed-fingerprint later pins consent to. This is also why
there’s no per-row Clean button in the candidates list above: cleaning a
resource always goes through one explain call that captures the
fingerprint the confirmation flow needs.
The sheet is ordered by the questions you have when you open it: may I
clean this, is it worth cleaning, on what evidence — and last, collapsed
behind a disclosure, the literal values explain --json returned, which
stay selectable so they can be quoted in a bug report.
- The Clean button’s enabled state reads
executablefrom theexplainreport directly — never a re-derived guess frompolicy_label. - If the resolved action
requires_confirmation(anASKorUNKNOWN_INCOMPLETEcandidate), tapping Clean shows a confirmation alert before doing anything. - Confirming runs
glomeris execute --action-id <id> --resource-id <id> --confirm-ask --observed-fingerprint <token> --progress-json, using the exact fingerprint token captured from theexplaincall above — never a fingerprint freshly re-observed at click time, which would defeat the whole point of fingerprint-pinning. - The outcome (
succeeded,failed,aborted_by_revalidation, or a specific refusal reason likeprotected/ask_consent_mismatch) is rendered from the realExecuteReport/ExecuteRefusalReportJSON the CLI printed — one specific message per outcome, not a generic success/failure toast.
AI Plan
The one card that talks to the internet, and only ever when you press its button. There is no timer behind it, nothing runs when the popover opens, and nothing retries — asking a provider costs money, so asking has to be something you did. Until you press it the card says exactly that: “Nothing has been sent anywhere.”
Pressing Ask AI for a plan runs glomeris llm-plan --json --progress-json with your project roots, streams the same per-detector
progress the Reclaimable space card shows, and offers a Stop button that
terminates the CLI child rather than just abandoning the result. Above the
button, permanently: “The model recommends. Glomeris decides what may run.”
and a note that asking costs money and that “Settings shows exactly what would
be sent, without sending it” — the privacy preview described under
Preferences — AI Provider below.
Each suggestion is one row, and every row is split down the middle by who said what:
- The machine’s half — safety class, storage impact, the size tier, and the evidence/confidence pair — is rendered as badges from the same vocabulary the candidates list uses, and comes from the real policy engine. It reads identically whether or not a provider was ever contacted. Below them, in plain words, is what Glomeris is willing to do: “Glomeris is willing to run this”, “Glomeris will ask you to confirm this before anything runs”, or the refusal, verbatim.
- The model’s half — its rationale — is an attributed, italic quotation under a “The model says” label. It is deliberately not a badge. A badge in this app means a verdict was reached; putting a model’s sentence in one would dress an opinion as a finding. A screen reader hears the machine’s verdict first, then “The model’s reason, which is advice and not a verdict: …”.
The rows appear in the order the provider returned them, and the card says so
(“The model’s order. Glomeris’s own ranking is the list above.”). No rank
number is drawn on a row, and the model’s own priority field is carried in
--json but never shown: the list position already makes that claim once, and
two competing numberings would read as though one of them were authoritative. Nothing in the app re-sorts the list either — Glomeris’s
ranking is the Reclaimable space card, which is why the AI Plan card sits
below it.
A recommendation cannot make anything runnable. A model that confidently
describes your SSH private key as a stale build directory gets its sentence
quoted, next to a PROTECTED badge and a refusal, with no action offered — the
row is built from the same executable/offered_actions/refusal_reason
fields the candidates list reads, which the CLI computed before the provider
was contacted. Suggestions naming a resource Glomeris never found, or an action
it does not have, are dropped by the CLI and the count of them is printed under
the list rather than quietly swallowed.
Tapping a row opens the same candidate detail sheet as the candidates list,
which issues its own explain call and owns the only Clean button in the app.
A plan item carries no fingerprint token, so there is no shortcut past that —
cleaning something a model suggested goes through exactly the path, and the
same confirmation, as cleaning something you found yourself.
Six situations, six distinct messages: never asked; asking; no provider
configured; the provider answered with nothing; the provider call failed
(quoted); and output this app could not read, which means the app and the
glomeris on PATH are different versions. The first three are not failures
and are not coloured like failures. The no provider configured message says
why a shell user can still land there — a Finder-launched menu-bar app inherits
no shell environment, so GLOMERIS_LLM_* variables exported in a terminal are
invisible to it — and is the one state with a Set up an AI provider… button
beside it, which opens Settings. The button is beside the message rather than
inside it on purpose: the copy layer is a pure token-to-words mapping, and a
message that could carry a control would be a message that could act.
Disk space history, and What Glomeris has done
Two independent, already-computed lists in two cards, each polled the same way as the status section:
- Disk space history — the last N pressure transitions from
glomeris history --json(e.g.WARN -> CRITICAL, with the used-percent/free-space reading at the time). The two ends of each transition are named exactly as the status card names them, because they are the same enum. - What Glomeris has done — the last N real-execution records from
glomeris actions history --json: which action ran against which resource, its policy label, outcome, and which ofexecute/free/emergencyproduced it. “Triggered by” is spelled out in words, because whether something was cleaned that you did not ask for is the question this card exists to answer. TheAUTO_SAFE-vs-refused/aborted visual distinction readsoutcome/abort_reasondirectly, never thepolicy_labeltext — andaborted_by_revalidationreads as “Stopped safely” rather than as a failure, because the resource changed between checking and acting and so nothing was touched.
An empty list in either card says only that it is empty. Neither claims an all-clear it cannot support: an empty pressure history could mean the disk has been steady, or it could mean nothing has been watching it, and the report does not distinguish those.
Neither list accepts a --project-root flag — both read global
daemon-state files (history.tsv/actions.jsonl), not project-scoped
detection state.
Preferences — Project Roots
The popover’s Project roots… footer button — or Cmd+, — opens Settings,
whose Projects tab is a simple list with add/remove controls for the
project-roots preference, backed by a small local store. (The footer exists
because the menu-bar item opens a window rather than a menu, so the popover is
the app’s only surface; Quit is there for the same reason.)
These are the same paths you’d otherwise pass repeatedly as --project-root <path> on the command line — the CLI’s cargo/node detectors only look
under directories they’re told about. This view is pure presentation: it
edits the stored list and does not itself call detect/explain/
execute; the roots are appended as --project-root arguments the next
time another card (Disk space, Reclaimable space, …) spawns the CLI.
Preferences — AI Provider
Settings’ second tab (HORO-1309) is where the BYOK provider is configured, for the common case of someone who runs the app from Finder and has never exported anything in a shell. Four cards, in the order the questions arise:
AI Provider — the API root address and the model name. Under each field, a
row saying where that field’s effective value came from: Configured here,
From the environment, or Not set. Per field, what you type here wins over
what the app inherited; a field left empty falls back to the inherited
GLOMERIS_LLM_* variable, including falling back to it being absent. Clearing
a field therefore returns you to the environment rather than switching a working
setup off. The rule is that the settings screen never lies: if the field shows
an address, that is the address used. Neither field has a Save button, because
neither has unsaved state — they write through as you type. Nothing here
validates the URL or knows which models exist; the CLI does both, and
Test connection is what reports it.
API Key — a SecureField, a Save button, and a Remove key button
marked as the destructive action it is. The key goes into the login keychain,
and from there into the environment of the glomeris child process at the
moment one is spawned. It is never a command-line argument (visible to ps, and
refused by the CLI by flag name), never this app’s preferences, never a file the
app writes, never a log line, and never displayed: there is no reveal control,
nothing reads it back, and the typed text is cleared as soon as it is handed to
the keychain — whether the write succeeded or not, so a failure cannot leave it
sitting in a field behind an error message.
Test Connection — one glomeris llm-check --json, described honestly as
“one tiny request — two words”, disabled until all three of the address, the
model and the key are available. The result is one of five outcomes in plain
words, each reading differently because each has its fix in a different place:
a rejected credential, an unreachable host, a reply the app could not use, a
base URL that cannot work, and success — which reports the path it posted to,
the model, and what came back. See
BYOK LLM Planner
for the outcome table. A failure is carried verbatim; nothing is paraphrased
into something more reassuring than what happened.
What Gets Sent — one glomeris llm-plan --print-payload --json, with the
same project roots a real plan would use, so it previews your payload rather
than a generic example. It splits the result into Leaves this Mac (the two
prompts, with a character count) and Stays on this Mac (the wire-alias
table, which is where the absolute paths are). The second group is blue, not
green: green in this app means AUTO_SAFE or succeeded, and data being
withheld is a deliberate hold, not a success.
Opening that preview sends nothing and cannot — --print-payload returns before
a provider is constructed — and the preview is the one command on this screen
that runs without the credential in its environment. Neither button does
anything until pressed: there is no timer, nothing runs when the window opens,
and nothing retries.
Preferences — Autopilot
Settings’ third tab (HORO-1310). Autopilot is the one feature
that grants standing permission to delete without being asked again, and the
person granting it is the least likely to be reading --help — so a grant that
could only be made and read on the command line would be a grant most users
would never read. This tab is where AC 6 of that ticket (“see exactly what
Autopilot is authorized to do before enabling it”) and AC 7 (“revoke
immediately”) are met for someone who never opens a terminal.
It is still a thin client. Every choice it offers arrives as data from
glomeris autopilot show --json: the checkboxes are allowlistable_kinds, the
steppers’ upper bounds are ceilings, the picker’s options are
pressure_states, the one question that may be answered in advance is
preauthorizable_reasons, and the “never available” lists are
never_allowlistable_kinds, never_preauthorizable_reasons and
never_executable_labels verbatim. If the CLI cannot be reached the tab draws
no form at all and says so, rather than falling back on its own idea of what
Autopilot allows — a settings screen offering a limit it invented would be a
policy decision made in Swift.
Eight cards, in the order the questions arrive:
- Autopilot — on or off, then what is in force right now in the CLI’s own
figures (including
max_bytes_human, so the number you read is formatted by the thing that enforces it), the envelope’s path, a Revoke now button and a Refresh. An enabled grant is drawn in the caution tone, never the critical one: a standing authorization you chose is not a fault. - What It May Reclaim — one checkbox per allowlistable kind, in the CLI’s own order, each with its plain-language name and what it is. Below them, the kinds that are never available whatever is granted here.
- How Much, Per Run — three steppers (actions, whole gigabytes, seconds), described as three independent budgets a run stops at whichever it reaches first, with the hard ceilings quoted underneath.
- When It May Act — the disk-pressure floor, including “Whenever there is something to reclaim” for no floor at all.
- Answering In Advance — the narrow
ASKpre-authorization, onekind:reasonpair at a time. Until a kind is ticked there is nothing to answer, and the card says that instead of offering a consent that would apply to nothing. - Grant This / Change This Authorization — the write. Disabled until at least one kind is ticked, and it says so.
- What the AI Decides — the
ai_authoritysentences, quoted rather than paraphrased, and the labels Glomeris will never delete whatever is authorized here and whatever a model recommends. - See What It Would Do — a copyable
glomeris autopilot run --dry-run.
Five properties of this screen are worth stating outright, because each is a way it could have been quietly wrong:
The form is pre-filled from what is in force, because enable replaces the
whole envelope. glomeris autopilot enable does not merge into the previous
grant (see Autopilot for why), so a form that
started from defaults would let you narrow one field and silently reset the
other five. Pressing Enable without touching anything re-grants exactly what
was already there. For the same reason the byte budget rounds up to the
next whole gigabyte: a 1.5 GiB grant shown as “1 GB” would mean that merely
opening this window and saving shrank a budget nobody touched.
The floors here are this screen’s, and they only narrow. The steppers stop
at 1 action, 1 GB and 5 seconds. The CLI accepts --max-actions 0 — an enabled
grant that can do nothing — which is coherent as an API and pointless as a
setting, so this window does not offer it. Nothing here can ask for more than
the CLI’s ceilings, which is a property of the values the report carries rather
than of a number typed in Swift.
Withdrawing a kind withdraws its advance consent with it. Unticking a kind
prunes any kind:reason pre-authorization that named it, so consent cannot
outlive the kind it applied to, or quietly come back when the kind is re-ticked.
There is no way to start a run from this window. Not a Run button, not a
dry-run button, no timer, nothing on appear: the only thing that runs
automatically is the read. autopilot run is also the one autopilot verb with
no --json, so the thin client has nothing to parse if someone adds a button
later without thinking about it. A run deletes, and the argument for keeping
deletion on a command you typed is the same one that keeps glomeris emergency
out of the app.
A revocation that failed is never reported quietly. The two write
directions fail in opposite ways — a failed enable leaves you with less
authority than you asked for, a failed revoke leaves standing deletion
authority in force — so they have separate wording, and every failed revoke
says that the authorization still is in force and how to withdraw it with
glomeris autopilot revoke. Exit 0 with output this app cannot read is treated
as a successful write in both directions, because the CLI prints its report
after the envelope has been saved.
The grant does not expire on its own. It is a file, it survives quitting and restarting, and it stays in force until it is revoked here or on the command line — which the card that grants it says in those words, because an authorization the user believes is temporary would be the worst kind of quiet.
How the app finds the glomeris CLI
Every screen spawns the CLI, so the app has to decide which binary that is. It checks these locations in order, and uses the first one that exists and is executable:
- Inside the app bundle, at
Contents/MacOS/glomeris. Release builds ship the CLI here, and it wins because it is version-matched to the app and covered by the bundle’s signature. Debug builds contain no embedded CLI, so development is unaffected. - Each absolute directory in
PATH, in order. Relative entries — including the empty entry that shells read as “the current directory” — are ignored, so nothing can substitute a binary by writing a file namedglomerisnext to the running process. /opt/homebrew/bin, then/usr/local/bin— the Homebrew prefix on Apple Silicon and on Intel respectively, and where Installation tells you to put a binary from a release tarball or a source build.
Step 3 is not redundant with step 2: a GUI app does not inherit your
shell’s PATH. An app launched from Finder gets
PATH=/usr/bin:/bin:/usr/sbin:/sbin, which contains neither Homebrew
prefix — so PATH alone would not find a brew installed CLI, even though
running glomeris in a terminal works fine.
Resolution happens on every invocation, not once at launch, so installing the CLI while the app is already running takes effect at the next poll without a restart.
If no binary is found, each section says so and lists the locations it searched, rather than failing silently or naming a path it only assumed.
Seeing which binary is in use
The Command-line tool card names the binary the app resolved: its full path, which of the three rules above chose it, and a fingerprint — the first twelve characters of the SHA-256 of its contents, with the whole hash in the tooltip.
The fingerprint is there because the version number cannot do this job. Two
glomeris binaries on one machine both reported version 0.2.0 while
disagreeing about whether an action with no scoped path may be offered at
all: the version is stamped from the crate version, so it does not change
between merges and is not a build identity. The hash is.
A release build records the hash of the CLI it ships, so the card can say whether the binary in use is that one. The four things it can say:
| The card says | What it means |
|---|---|
| Matches this app | The resolved binary is the one this app ships. |
| Not the build this app ships | Both hashes are known and differ. The app still works, and still drives that binary — what it does may differ from what the app describes. |
| Nothing to compare against | This app records no expected hash. Normal for a locally built app; see below. |
| Cannot be checked | The binary can be run but its contents could not be read, so no comparison was possible. |
A locally built app embeds no CLI. The embed step lives in the release
workflow, not in the Xcode project, so an app built with xcodebuild — or
from Xcode — contains no Contents/MacOS/glomeris and always falls through
to PATH or the Homebrew prefixes. Whatever is installed on the machine is
what your build drives, and that may be older or newer than the tree you
built from. A local build also records no expected hash, which is why the
card reads “Nothing to compare against” rather than reporting a mismatch.
If you are testing a change to the CLI from a local app build, the path on
the card is the one to check: cargo build alone does not put a binary
anywhere the app looks.
Every screen maps back to a CLI command
| Card | CLI command(s) |
|---|---|
| Disk space | glomeris status --json |
| Background monitor | glomeris daemon status --json |
| Reclaimable space + Refresh | glomeris detect --json --progress-json |
| AI Plan — Ask AI for a plan | glomeris llm-plan --json --progress-json |
| Settings → AI Provider — Test connection | glomeris llm-check --json |
| Settings → AI Provider — Show what would be sent | glomeris llm-plan --print-payload --json (no provider contacted) |
| Settings → Autopilot — on appear, Refresh | glomeris autopilot show --json |
| Settings → Autopilot — Enable / Save changes | glomeris autopilot enable --kinds <tag,...> --max-actions <N> --max-bytes <N> --max-duration <secs> --min-pressure <state|none> [--preauthorize-ask <kind>:<reason>]... --json |
| Settings → Autopilot — Revoke now | glomeris autopilot revoke --json |
| Candidate detail | glomeris explain <resource_id> --json --progress-json |
| Clean (with confirmation) | glomeris execute --action-id <id> --resource-id <id> [--confirm-ask --observed-fingerprint <token>] --json --progress-json |
| Disk space history | glomeris history --json |
| What Glomeris has done | glomeris actions history --json |
See CLI Reference for the full flag/exit-code/JSON-shape reference behind every one of these.
Quick Start
A short tour of the commands worth running first. The complete list, with
every flag and exit code, is in CLI Reference — and in the
binary itself, which is the better habit: glomeris --help, then
glomeris help <command> for any one of them.
If you would rather click than type, the menu-bar app covers detect, explain, clean and AI Plan over the same binary — see Menu Bar App.
Find out where to start
glomeris
Prints the version and one line pointing at glomeris --help. Nothing else;
a bare invocation reads no disks and changes nothing.
Discover what tool caches exist on this machine
glomeris detect
Runs every built-in detector (Xcode DerivedData, Homebrew cache, Cargo target
dirs, node_modules, Docker build cache) once and prints one line per
detector: found (<N> evidence), tool_absent, or failed: <reason>.
tool_absent is a normal, expected state — it means that tool isn’t
installed or has no cache yet, not an error.
Scan a directory for the largest entries
glomeris scan [path] [top_k]
Both arguments are positional, not flags, and both are optional: path
defaults to ., top_k defaults to 20. This walks the tree without
buffering it fully in memory and prints the top-K largest entries by
logical size. It is a generic size scan, unrelated to the detectors above —
see Architecture for how the two differ.
Try the bounded recovery loop (macOS only)
Set a goal for how full the disk should end up, and let the loop work toward it:
glomeris free --goal-used-percent 60
Or state the same thing as a free-space floor, which is what the loop itself works in:
glomeris free --target 10GB
# or
glomeris free --target 15%
Exactly one of the two is required, and they are different axes:
--goal-used-percent is target disk used, --target is a free-space
floor. --goal-used-percent 60 and --target 40% ask for the same end state.
--target accepts either an absolute size (B/KB/MB/GB/TB,
binary/1024-based) or a percentage of total capacity (0–100, suffixed
%). A --goal-used-percent that is not an improvement on your current usage
is refused before anything is deleted, rather than run and reported as a
success.
This one deletes. An earlier version of this page said no real deletion
could complete through this path; that stopped being true in HORO-994. Read
Safety Model first, and note that free declines every
ASK candidate rather than prompting (see
Known Limitations) — so it reclaims only what policy
classified AUTO_SAFE on its own.
To see what reaching the goal would take without touching anything:
glomeris free --goal-used-percent 60 --dry-run
That prints current usage, the free bytes still needed, and the estimated
reclaimable opportunity split by what policy would actually permit. It takes
no execution lock and mutates nothing. glomeris clean --dry-run remains the
way to see a plan for one pass rather than a goal.
Run unattended, inside limits you grant
glomeris autopilot # what am I allowing today?
glomeris autopilot enable --kinds node_modules \
--max-actions 1 --max-bytes 1073741824
glomeris autopilot run --dry-run # the real bounded plan
The one command that acts without you watching, so the grant is written to a
file you can read rather than inferred. It grants nothing until you enable
it, --kinds is required, and --max-bytes is a plain byte count. AUTO_SAFE
only unless you pre-authorise one specific ASK reason on one specific kind;
PROTECTED refuses unconditionally and no flag here changes that. See
Autopilot.
Try emergency mode (macOS only)
glomeris emergency
Takes no arguments by design, and acts machine-wide on everything it finds
AUTO_SAFE. See Emergency Mode for exactly what it does
and does not do before running it on a machine you care about.
CLI Reference
The built-in help is canonical. glomeris <command> --help renders from
src/cli/help.rs, which is the one place a command’s usage, flags, safety
semantics, examples and exit codes are written down, and whose output is held
to golden snapshots in tests/fixtures/help/ (HORO-1311). If this page and
--help ever disagree, --help is right and this page is stale.
What this page adds that --help deliberately does not: the JSON payload
shapes, the report field semantics, and the cross-references into the rest of
this book. It is the reference you read at a desk; --help is the one you read
mid-task.
glomeris
No arguments: prints glomeris <version> (from CARGO_PKG_VERSION) and a
pointer to glomeris --help. --version/-V prints the same version line on
its own.
Help surfaces
Everything below renders from the single COMMANDS table in
src/cli/help.rs. That table lived in src/main.rs until HORO-1311, where no
test could read it — so tests/shared_command_table.rs kept a hand-written
mirror of it, which drifted exactly as the pre-HORO-1050 duplication had
(HORO-1034: the unrecognized-command usage omitted llm-plan; by HORO-1311 the
test mirror omitted llm-check). There is now one table, read directly by both
the binary and its tests.
| Invocation | What it prints |
|---|---|
glomeris --help / -h / help | Every command, grouped by what it does to your machine, with a one-line summary each. |
glomeris <command> --help / -h | That command’s usage, safety statement, its subcommands (each with its own safety label), flags, examples, exit codes and related commands. |
glomeris help <command> | Identical to glomeris <command> --help. |
glomeris help exit-codes | The exit-status reference (see Exit codes). |
glomeris help <unknown-topic> | The list of topics that exist, to stderr, exit 2. |
glomeris <unknown-command> | A one-line error and a pointer to --help, to stderr, exit 2. |
glomeris <command> <bad-flag> | That command’s usage only, to stderr, exit 2. |
The groups in top-level help are INSPECT, PLAN, ACT, OBSERVE,
SERVICE and CONFIGURE, ordered so that the read-only commands come before
anything that can delete, and so that a first-time reader meets the product
before its preferences. A group heading describes consequence, not category: INSPECT says
“nothing is changed”, and ACT says its commands delete data and that every
deletion is policy-gated.
Note that a command’s group and its safety line describe only whether
Glomeris itself writes to the filesystem when you run it. They are not policy
classifications: AUTO_SAFE, ASK and PROTECTED classify resources, are
decided by the policy engine, and never appear in a help safety label. See
Safety Model.
Safety is declared per verb, not only per command
Four labels exist, in ascending order of consequence:
| Label | Means |
|---|---|
Read-only — changes nothing. | Nothing is written. |
Advisory — proposes, never executes. | Produces a plan; executes none of it. |
Writes only Glomeris's own state — never your files. | Writes the launch agent plist, the monitor’s history and heartbeat, the Autopilot envelope, or your stored preferences. Nothing you own. |
Can delete data — every deletion is policy-gated. | Deletes. |
For a command that takes subcommands, the consequence is a property of the
verb, not of the command name: daemon install writes a launch agent while
daemon status reads, and autopilot run deletes while autopilot show
prints a config file. Until HORO-1485 one label was declared per command, so
glomeris daemon --help printed “Read-only — changes nothing.” above
install, uninstall and run, and glomeris autopilot --help printed “Can
delete data” above show. No wording could have fixed either: the weakest
label is a false reassurance and the strongest a false warning, which is why
relabelling daemon as destructive was not an acceptable fix.
Each verb now carries its own label, printed beside it in the Subcommands
block. The command’s own safety line is the strongest claim reachable through
it, and when its verbs disagree it says so explicitly rather than picking one
of them:
Safety depends on the subcommand; each is labelled below. The strongest of
them: Writes only Glomeris's own state — never your files.
Two tests keep this honest, and they are deliberately not tests about strings.
src/cli/help.rs‘s command_safety_covers_every_reachable_subcommand asserts
the command’s label equals the strongest of its verbs’ — equality in both
directions, because under-stating and over-stating are both false — and
a_single_label_banner_is_true_of_every_verb_under_it asserts the same of the
rendered banner, since a correct table printed through a banner that ignored it
would still print a false claim. Neither could have caught the original defect
on its own: with no subcommands declared, both pass vacuously.
tests/subcommand_safety_is_honest.rs closes that gap from outside the table —
it reads src/main.rs’s dispatch arms, so a verb the binary accepts cannot go
undeclared, and it runs every surface labelled read-only against a disposable
HOME and requires the tree to be byte-for-byte unchanged, with
autopilot enable as the positive control that proves the observation would
notice a write.
Two surfaces are narrower on purpose. An unrecognized top-level command gets a
pointer rather than the manual — it previously reprinted the aggregate usage of
all thirteen commands, a 13-line, 146-column wall in answer to one mistyped
word. And a usage error inside a command prints only that command’s usage, so
getting a free flag wrong no longer tells you about daemon.
Byte counts: 1024-based, with KB/MB/GB labels
Every byte count this product renders — free_human, total_human,
logical_human, reclaimable_human, expected_reclaimed_human,
actual_reclaimed_human, and the same numbers in text output — comes from one
function, reporting::human_bytes. It divides by 1024 and labels the result
B/KB/MB/GB/TB/PB. So 2147483648 renders as 2.0 GB, where a
1000-based formatter would say 2.15 GB. Counts below 1024 are a bare integer
and B, with no decimal point.
The labels are not the IEC KiB/MiB/GiB spelling that strictly matches the
arithmetic. That is a deliberate, documented inconsistency rather than an
oversight: these are the units du -h and df -h print for the same
arithmetic, and the alternative is renaming every unit in every report and
fixture to spell out a distinction most readers of a storage tool do not draw.
--target parses the same way, so what you type and what you read back agree:
glomeris free --target 5GB means 5 × 1024³ bytes. A bare number or a B
suffix is raw bytes; a trailing % is a percentage of total capacity instead.
Clients should render the *_human string rather than scale the byte count
themselves. The menu-bar app did the latter with ByteCountFormatter, which
is 1000-based, so a 2 GiB cleanup appeared as a 2.0 GB estimate and a
2.15 GB result in the same panel (fixed in HORO-1312). Where a report offers
both fields, the raw *_bytes value is for arithmetic and the *_human string
is for display. A *_human field is null, never "0 B", when the underlying
probe produced no number — an aborted execution reclaimed nothing, which is a
different claim from having freed zero bytes.
glomeris daemon <subcommand>
macOS only (exits 1 with an error message on other platforms).
| Subcommand | Effect |
|---|---|
install | Writes a per-user launchd plist and launchctl load -ws it. |
uninstall | launchctl unload -ws the agent (best-effort) and removes the plist file. |
status | Prints whether the plist is installed, its path, and whether launchctl reports it loaded. |
run | Runs the polling loop in the foreground (this is what the installed agent actually executes). |
No subcommand, or an unrecognized one, prints usage to stderr and exits 2. See Daemon Lifecycle for details.
daemon status --json (HORO-1045) prints a DaemonStatusReport.
loaded (launchd-reported) and heartbeat_age_secs (derived from the poll
loop’s own last-write) are deliberately kept as two separate fields, never
collapsed into one healthy boolean — a loaded-but-wedged daemon and an
actually-polling one must stay distinguishable:
glomeris daemon status --json
{
"plist_installed": true,
"plist_path": "/Users/dev/Library/LaunchAgents/dev.glomeris.daemon.plist",
"loaded": true,
"heartbeat_age_secs": 42
}
heartbeat_age_secs is null when no heartbeat file exists yet (the
daemon has never run).
glomeris status [--json]
macOS only (exits 1 with an error message on other platforms). Not part of
daemon — this is the same one-shot disk-pressure reading daemon run’s
poll loop evaluates each tick, available on demand without needing the
daemon installed at all (HORO-955).
glomeris status --json
{
"total_bytes": 500000000000,
"free_bytes": 125000000000,
"used_percent": 75.0,
"free_human": "116.4 GB",
"total_human": "465.7 GB",
"pressure_state": "WARN"
}
pressure_state is one of "OK", "WARN", "CRITICAL" — see
Pressure Model.
glomeris scan [path] [top_k]
Not macOS-gated. Both arguments are positional and optional:
path— root directory to scan. Defaults to..top_k— number of largest entries to report. Defaults to20if omitted or unparseable asusize.
Prints a summary line (files visited, stop reason, incomplete-entry count) followed by one line per candidate: size, depth, path.
glomeris detect [--project-root <path>]... [--json] [--progress-json]
Not macOS-gated. Runs every detector in DetectorRegistry::builtin() exactly
once per invocation and prints, per detector: found (<N> evidence),
tool_absent, or failed: <reason>. The candidate report printed below those
lines comes out of that same single pass, so the two halves of the output
cannot describe different probes of a filesystem that changes between them
(HORO-1487).
--project-root <path> is optional and repeatable — pass it once per
project directory you want the cargo/node detectors to check for a
target//node_modules/ dir. Without it, those two detectors have no
project roots to scan and always report tool_absent.
--json prints a DetectReport — one candidate line per discovered
resource, including the already-computed
executable/offered_actions/refusal_reason triple (HORO-1053) a caller
(e.g. the menu-bar app) reads to decide what it can offer, without ever
re-deriving that from policy_label/reasons itself:
Candidates are returned biggest reclaimable size first (HORO-1307).
Before that they came back in detector-registration order, so a 40 GB Cargo
target/ could be printed below a 2 MB npm cache. Ties are broken first by
measurement quality — an exact size outranks a ≥ lower bound of the same
number, because the exact one is the claim you can act on — and then by
resource_id, so two runs over an unchanged machine produce the same order.
Candidates whose size could not be measured at all sort last rather than
being treated as zero. Ordering is applied at the single point where the
report is assembled, so --json and the human-readable output can never
disagree about it.
impact_tier is "unknown", "normal", "notable" or "large": a
pre-computed magnitude band, so a UI does not have to invent thresholds of
its own. It escalates on either an absolute size (≥ 1 GB is notable,
≥ 10 GB is large) or a share of remaining free space (≥ 5% is
notable, ≥ 20% is large) — measured against free rather than total space,
because the problem someone opens Glomeris with is “I am running out of
room”. A 400 MB cache is large on a machine with 1.6 GB left.
impact_tier is a size signal and nothing else. It is not a safety
signal, and it must never be read as one: a large candidate can be
PROTECTED, and a normal one can be AUTO_SAFE. executable,
offered_actions and refusal_reason remain the only statement about what
Glomeris is permitted to do.
glomeris detect --project-root ~/dev/myproject --json
{
"candidates": [
{
"resource_id": "cargo_target_dir:/Users/dev/proj/target",
"kind": "cargo_target_dir",
"reclaimable_bytes": 2147483648,
"reclaimable_human": "2.0 GB",
"reclaimable_bytes_is_lower_bound": false,
"impact_tier": "notable",
"policy_label": "AUTO_SAFE",
"reasons": ["no_active_use_observed"],
"executable": true,
"offered_actions": [
{
"action_id": "cargo.clean.target_dir",
"requires_confirmation": false
}
],
"refusal_reason": null
},
{
"resource_id": "docker_build_cache:docker",
"kind": "docker_build_cache",
"reclaimable_bytes": 10737418240,
"reclaimable_human": "10.0 GB",
"reclaimable_bytes_is_lower_bound": true,
"impact_tier": "large",
"policy_label": "UNKNOWN_INCOMPLETE",
"reasons": ["evidence_incomplete"],
"executable": false,
"offered_actions": [],
"refusal_reason": "no registered cleanup action for this resource kind"
}
],
"detectors": [
{"detector": "cargo_target_dir", "status": "found", "candidates_found": 1},
{"detector": "docker_build_cache", "status": "found", "candidates_found": 1},
{"detector": "node_modules", "status": "tool_absent", "candidates_found": 0},
{
"detector": "homebrew_cache",
"status": "failed",
"candidates_found": 0,
"reason": "brew --cache exited with status exit status: 1"
}
],
"discovery_complete": false
}
detectors reports one entry per detector that ran, in registry order, and
status is found, tool_absent or failed — the same three tokens
--progress-json uses below. reason is present only on failed, and is the
detector’s own account of what went wrong.
discovery_complete (HORO-1484) is false when at least one detector failed,
and it is derived from detectors rather than tracked separately, so the
summary cannot disagree with the array it summarises. When it is false,
candidates is not a complete account of what could be reclaimed, and no
consumer may present it as one — “nothing worth reclaiming” is a claim about
the machine, and a search that did not finish has not established it. A
tool_absent detector does not make discovery incomplete: a tool that is
not installed has nothing to report, which is a different fact from a tool that
was asked and could not answer.
--progress-json (HORO-1052) emits one NDJSON-encoded ProgressEvent line
to stderr per detector start/finish while discovery runs — a way for a
spawning UI to distinguish “still working” from “hung” on a slow/contended
host (v0.2.0 founder-dogfood measured a single discovery pass up to 3m40s).
Stdout is completely unaffected either way, --progress-json output can be
combined with --json, and nothing is emitted at all unless the flag is
passed:
glomeris detect --progress-json 2>&1 1>/dev/null
{"phase":"detector_started","detector":"cargo_target_dir"}
{"phase":"detector_finished","detector":"cargo_target_dir","candidates_found":1,"outcome":"found"}
{"phase":"detector_started","detector":"docker_images"}
{"phase":"detector_finished","detector":"docker_images","candidates_found":0,"outcome":"tool_absent"}
{"phase":"detector_started","detector":"homebrew_cache"}
{"phase":"detector_finished","detector":"homebrew_cache","candidates_found":0,"outcome":"failed","reason":"brew --cache exited with status exit status: 1"}
outcome (HORO-1484) is found, tool_absent or failed, and reason is
present only on failed. The last two lines above are why it exists: a
detector whose probe failed reports candidates_found: 0, exactly like one
that looked and found nothing, so a consumer reading only the count shows
“0 found” for a check that never ran. docker_images being absent is normal
and expected; homebrew_cache failing is not, and the two must not be
presented alike.
detect, explain, llm-plan, and execute all share this same
discovery phase and all support --progress-json identically.
glomeris explain <resource_id_or_path> [--project-root <path>]... [--json] [--progress-json]
Not macOS-gated. Runs the same discovery-and-classification pipeline as
detect, then prints the full evidence-and-policy picture for exactly the
one resource matching <resource_id_or_path> (a detect report’s
resource_id, or a filesystem path) — evidence provenance, size (logical
vs. reclaimable, explicitly labeled as different), active-use signals, and
policy classification (HORO-955). Exits 1 if no discovered candidate
matches the query.
--project-root <path> and --progress-json mean exactly what they mean
for detect above.
glomeris explain cargo_target_dir:/Users/dev/proj/target --json
{
"resource_id": "cargo_target_dir:/Users/dev/proj/target",
"kind": "cargo_target_dir",
"detector": "cargo_target_dir",
"sources": ["cargo metadata: target-dir"],
"logical_bytes": 2147483648,
"logical_human": "2.0 GB",
"reclaimable_bytes": 2147483648,
"reclaimable_human": "2.0 GB",
"reclaimable_bytes_is_lower_bound": false,
"completeness": "complete",
"confidence": "high",
"active_use_signals": [],
"regenerability": "regenerable_by_rebuild",
"policy_label": "AUTO_SAFE",
"reasons": ["no_active_use_observed"],
"native_cleanup_available": true,
"native_cleanup_action_id": "cargo.clean.target_dir",
"fingerprint_token": "<opaque token — copy verbatim, never hand-construct>",
"executable": true,
"offered_actions": [
{
"action_id": "cargo.clean.target_dir",
"requires_confirmation": false
}
],
"refusal_reason": null
}
fingerprint_token (HORO-1051) is an opaque, wire-safe encoding of the
resource’s identity fingerprint — null for a resource with no dev/inode/
mtime identity (e.g. Docker’s build cache). This is the exact token
execute --observed-fingerprint later expects back for an ASK-classified
resource; see execute’s section below.
glomeris clean --dry-run [--target <resource_id_or_path>] [--project-root <path>]...
Not macOS-gated. Renders what would be cleaned, without executing
anything — --dry-run is required; there is no non-dry-run execution path
on this subcommand (real destructive execution is glomeris free --target
or glomeris execute’s job). Without --target, every discovered
candidate is considered; with it, only the one matching resource is.
glomeris clean --dry-run
Human-readable output only — clean --dry-run has no --json mode. One
line per considered resource, either its rendered ActionPlan.explain text
or a skip_reason (e.g. PROTECTED, no registered action).
glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json] [--progress-json] [--schema] [--print-payload]
Not macOS-gated. ADVISORY, NON-EXECUTING (HORO-1008) — never constructs a
policy::Approval and never calls policy::approval::authorize or
executor::execute. See BYOK LLM Planner for the full
configuration and safety-property writeup.
- Without
--plan-file, credentials are read only fromGLOMERIS_LLM_API_KEY/GLOMERIS_LLM_BASE_URL/GLOMERIS_LLM_MODEL(all three required, no default base URL) — never from a CLI flag.GLOMERIS_LLM_BASE_URLis the API root:/chat/completionsis appended to it verbatim, so it usually ends in/v1(https://gateway.example.com/v1). See BYOK LLM Planner. - A failed provider call is reported on the
provider error:line (or theprovider_errorJSON field) as one secret-free sentence naming the HTTP status, API style, request path, request id, and a bounded excerpt of the provider’s error body — never theAuthorizationheader, the key, the host, or the request payload. --plan-file <path>reads the file’s raw bytes as if they were the model’s raw response text, through the same validation pipeline the live provider uses — no network call, no API key required.--project-root <path>is optional and repeatable, same meaning asdetect’s flag above.--api-key/--key/--tokenare explicitly rejected (not accepted and ignored) — the error names$GLOMERIS_LLM_API_KEYinstead.--jsonprints the report as JSON. Each item carries the model’s ownpriority/model_reasonalongside the machine’spolicy_label,requested_action_id,explain,skip_reason,completeness,confidence, and a nestedcandidate— the byte-for-byte same projectiondetect --jsonprints for that resource, includingexecutable/offered_actions/refusal_reason(HORO-1308). A consumer deciding what may be done readscandidate;priority/model_reasonare the only two fields a provider chose, and the order ofitemsis the provider’s too —plan_with_llmnever sorts, it only drops.--jsonoutput is printed before aprovider_errorexit, so an exit1report is still complete and readable.--progress-jsonstreams the same NDJSON discovery progress asdetect, on stderr, leaving stdout a single clean JSON document. This is the exact pair (--json --progress-json) the menu-bar app’s AI Plan card spawns; see Menu Bar App.--print-payload(HORO-1298) runs discovery, prints the exact request a live run would send — the system prompt, the user prompt, and the local wire-id-to-real-resource table under a heading marking it as not sent — then returns before any provider is constructed. Requires no credential and makes no network call, so it cannot send what it displays. The outbound prompts identify resources only by positional alias (resource_1, …); absolute paths appear in the alias table and nowhere else. See BYOK LLM Planner.--schema(HORO-1048) prints an example, syntactically validLlmPlanJSON document to stdout and exits — a distinct, self-contained mode that never runs discovery, never reads--plan-file, and never checks live-mode credentials, regardless of what else is passed alongside it. See BYOK LLM Planner for the full example and field table.
glomeris llm-plan --schema
{
"items": [
{
"resource_id": "cargo_target_dir:/path/to/project/target",
"action_id": "cargo.clean.target_dir",
"priority": 1,
"reason": "stale build artifacts, not modified in 30 days"
}
]
}
This exact output round-trips unchanged through glomeris llm-plan --plan-file <path> — see the linked BYOK page for the full field table and
the round-trip test that proves it. The resource_id form shown here is
the one a human writes by hand in a fixture; a live model is given
positional wire aliases and answers with those, and both forms resolve.
Human-readable output always opens with LLM SUGGESTION — advisory only, nothing is executed by this command.
Exit codes for this subcommand specifically:
0— success, including zero suggestions or every suggestion beingPROTECTED.1— the provider call or response parsing failed,--plan-filenamed an unreadable path, or--print-payloadcould not serialize the request.2— usage error: an unrecognized argument,--plan-filewith no value, missing live-mode environment configuration, or an--api-key/--key/--tokenflag.
glomeris llm-check [--json]
Not macOS-gated. Tests the configured BYOK setup and reports whether the endpoint, credential and model work (HORO-1309). Runs no detectors, reads no project roots, collects no evidence, and consults no policy — it is not a planning command, and there is nothing it could execute.
- Sends the two fixed prompts in
actions::llm::CONNECTION_TEST_SYSTEM_PROMPT/CONNECTION_TEST_USER_PROMPT(“You are a connection test. Reply with the single word: ok.” / “ok”) through the sameLlmProvider::completea real plan uses. Same code path, so a pass means a plan will route and authenticate — not that a cheaper probe succeeded. - The prompts are constants with no interpolation, so a connection test
describes nothing about this machine: no path, no home directory, no account
name (
connection_test_prompts_describe_nothing_local). - Configuration comes only from
GLOMERIS_LLM_API_KEY/GLOMERIS_LLM_BASE_URL/GLOMERIS_LLM_MODEL, all three required, exactly as forllm-plan.--api-key/--key/--tokenare rejected by flag name, and the error names the flag only — never the value beside it. --jsonprints anLlmCheckReport:outcome,model,endpoint_path,error,response_excerpt. Human-readable output opens withLLM CONNECTION OKorLLM CONNECTION FAILED (<outcome>).outcomeis one of five tokens, produced by the singleactions::llm::llm_check_outcomemapping so the CLI, the book and the menu-bar app cannot disagree about what a failure was:
outcome | What happened | Where the fix is |
|---|---|---|
ok | The provider answered and its reply is excerpted in response_excerpt | — |
unreachable | No HTTP response at all — DNS, TLS, refused connection, timeout | Network, host name, or VPN |
rejected | The provider answered with a non-2xx status | Credential, or the base-URL path — read the path in error |
unusable_response | A 2xx response that was empty, not JSON, or missing choices[0].message.content | Model name, or a gateway not actually speaking the OpenAI shape |
misconfigured | A base URL that cannot work — validate_base_url refused it before anything was sent | The base URL itself; see BYOK LLM Planner |
The fifth is the odd one out: LlmError::InvalidConfiguration can only come
from constructing the provider, which happens before any report exists, so
glomeris llm-check never prints a report whose outcome is
misconfigured — it exits 2 with one line on stderr instead. The token is
in the vocabulary because that is the name for what happened, and a caller
branching on exit 2 (the menu-bar app does) classifies it that way itself
rather than inventing a sixth word for the same condition.
errorcarries the same secret-free sentencellm-plan’sprovider_errordoes — status, API style, request path,x-request-idwhen the provider sends one, and a bounded excerpt of the provider’s own error body, with the configured key scrubbed out of it. Never theAuthorizationheader, the key, the scheme, the host, or the query string.
glomeris llm-check --json
Exit codes for this subcommand specifically:
0— the provider answered andoutcomeisok.1— a check ran and did not pass (unreachable,rejected,unusable_response). The report is printed first, so an exit1still carries a complete, readable diagnosis on stdout.2— nothing was sent and there is no report at all: an unrecognized argument, an--api-key/--key/--tokenflag, a base URLvalidate_base_urlrefused, or missing configuration (the message names the three variable names, never a value). Stdout is empty; the reason is one line on stderr prefixedglomeris llm-check:.
The split matters for a caller branching on the code without parsing output,
which is exactly what the menu-bar app’s connection test does: 2 means “fix
your invocation or setup”, never “the network is having a bad day”. A 2 with
no report is the one case where a UI has to quote the CLI’s own sentence rather
than render a report.
glomeris execute --action-id <id> --resource-id <id> [--project-root <path>]... [--confirm-ask --observed-fingerprint <token>] [--json] [--progress-json]
macOS only (exits 1 with an error message on other platforms). The sole
interactive destructive-execution subcommand (HORO-1055) — the only place
in this CLI where a caller can trigger one specific, real destructive
action against one specific, real resource. --action-id and
--resource-id are both required; the caller supplies ONLY these
selectors (plus, for ASK, an observed fingerprint token) — there is no
flag to pass a PolicyClass, a raw filesystem path as a direct target, a
shell string, or any --force/override.
Internal flow, in one process:
- Acquires the same HORO-1054 execution lock
free/emergencyuse — held for the whole call, released on exit. - Runs the same discovery-and-classification pipeline
detect/explainuse (--project-root <path>is optional and repeatable, same meaning as elsewhere). - Resolves
--resource-idagainst the discovered candidates and--action-idagainst that resource’s own registered action — refusing if either does not resolve, or if the resolved action’s id differs from--action-id. - For an
ASK-classified resource, builds consent ONLY from decoding--observed-fingerprint(via the same token formatexplain --json’sfingerprint_tokenfield emits) — never from a fingerprint freshly observed by this same process, which would defeat the whole fingerprint-pinning purpose.--confirm-askand--observed-fingerprintmust be passed together or not at all. - Calls the real, unmodified
policy::approval::authorize, then — only if it returns an approval — the real, unmodifiedexecutor::execute.PROTECTEDrefuses unconditionally regardless of any flag combination;execute’s own deletion-time revalidation can still abort a plan that was authorized a moment earlier if the resource changed in between.
--json prints an ExecuteReport (action id, resource id, outcome,
failure/abort detail, expected vs. actual reclaimed bytes — the latter is
a real measurement, taken after execution, not an estimate) on the
Executed path. Every refusal/not-found/busy path instead prints an
ExecuteRefusalReport ({"reason": "...", "message": "..."}) to stdout
before exiting with the matching code below, so a --json caller never
gets silent stdout on a non-Executed outcome. reason is one of
resource_not_found, action_not_found, action_mismatch, protected,
ask_no_consent, ask_consent_mismatch, auto_safe_contract_violation,
or busy (HORO-1056: the lock-contention case below, the one refusal
that happens before discovery/resolution even runs) — each a distinct,
machine-readable value naming exactly which refusal/abort path fired,
never a generic error string.
An AUTO_SAFE resource needs no confirmation flags at all:
glomeris execute --action-id cargo.clean.target_dir \
--resource-id cargo_target_dir:/Users/dev/proj/target --json
{
"action_id": "cargo.clean.target_dir",
"resource_id": "cargo_target_dir:/Users/dev/proj/target",
"outcome": "succeeded",
"failure_message": null,
"abort_reason": null,
"expected_reclaimed_bytes": 2147483648,
"actual_reclaimed_bytes": 2147483648,
"expected_reclaimed_human": "2.0 GB",
"actual_reclaimed_human": "2.0 GB"
}
The two *_human strings are new in HORO-1312 and are what a UI should
display — see Byte counts.
An ASK-classified resource requires --confirm-ask plus the exact
--observed-fingerprint token captured from a prior explain --json call
on that same resource (never a fingerprint freshly observed by execute
itself) — $FINGERPRINT_TOKEN below is that call’s fingerprint_token
field, copied verbatim, never hand-constructed:
glomeris execute --action-id cargo.clean.target_dir \
--resource-id cargo_target_dir:/Users/dev/proj/target \
--confirm-ask --observed-fingerprint "$FINGERPRINT_TOKEN" \
--json --progress-json
If the resource’s identity changed between the explain call and this
execute call, revalidation aborts the plan rather than proceeding:
{
"action_id": "cargo.clean.target_dir",
"resource_id": "cargo_target_dir:/Users/dev/proj/target",
"outcome": "aborted_by_revalidation",
"failure_message": null,
"abort_reason": "ResourceIdentityChanged",
"expected_reclaimed_bytes": 2147483648,
"actual_reclaimed_bytes": null,
"expected_reclaimed_human": "2.0 GB",
"actual_reclaimed_human": null
}
Note both actual_* fields are null rather than 0/"0 B". Nothing was
deleted, which is not the same report as a cleanup that freed no bytes.
Every refusal path (e.g. PROTECTED, no consent supplied, a stale
fingerprint) prints an ExecuteRefusalReport instead, with --json:
{
"reason": "protected",
"message": "refused — this resource is PROTECTED; no flag combination can authorize executing against it"
}
Exit codes for this subcommand specifically:
0— the action executed and succeeded.1— the action executed but failed (a step of the plan errored).2— usage error: an unrecognized/missing argument,--confirm-askwithout--observed-fingerprint(or vice versa), or a malformed--observed-fingerprinttoken.3— refused by policy:PROTECTED(unconditional),ASKwith no consent supplied, orASKwith a supplied consent that did not match the freshly observed fingerprint.4— aborted byexecute’s own deletion-time revalidation (a TOCTOU- style guard: the resource’s identity or policy classification changed between authorization and execution).5—--resource-idmatched no discovered candidate, the resource had no registered action, or the resolved action’s id did not match the supplied--action-id.75— the execution lock is already held by anotherglomerisinvocation (seefree’s exit codes above;executereuses the exact sameEXIT_EXECUTION_LOCK_BUSYconstant — this ticket’s own AC described this case as exit6, but the already-established lock convention from HORO-1054 is kept rather than introducing a second, conflicting “busy” code). With--json, this prints anExecuteRefusalReportwithreason: "busy"to stdout (HORO-1056) — previously this path was silent on stdout even under--json.
glomeris emergency
macOS only (exits 1 with an error message on other platforms). Takes no
arguments, and rejects any with exit 2 — until HORO-1311 it silently discarded
them, so glomeris emergency --dry-run performed a real recovery run. See
Emergency Mode.
glomeris history [--json] [--limit <N>]
Reads back a bounded, oldest-first tail of the monitor’s history.tsv
(HORO-1046) — the same append-only file daemon run’s poll loop already
writes via PersistenceBackend::record. No new persistence format; this is
a read path only.
--limit <N> is optional and defaults to 20. It bounds how many of the
most recent pressure transitions are returned — a malformed line in
history.tsv is skipped rather than failing the whole read, and a missing
history file (the daemon has never run, or never recorded a transition)
renders as an empty list rather than an error.
With --json, prints a HistoryReport ({"events": [...]}); each event
has unix_time_secs, from, to, used_percent, free_bytes, and
free_human. Without --json, prints one line per event as plain text.
glomeris history --json --limit 2
{
"events": [
{
"unix_time_secs": 1700000000,
"from": "OK",
"to": "WARN",
"used_percent": 82.5,
"free_bytes": 80000000000,
"free_human": "74.5 GB"
},
{
"unix_time_secs": 1700000600,
"from": "WARN",
"to": "CRITICAL",
"used_percent": 95.1,
"free_bytes": 20000000000,
"free_human": "18.6 GB"
}
]
}
glomeris actions <list [--json]|history [--json] [--limit <N>]>
Neither subcommand is macOS-gated — list touches no filesystem/launchd
state at all, and history only reads a plain file.
glomeris actions list [--json]
Enumerates every action currently registered in ActionRegistry::builtin()
(HORO-1047) — the read path that replaced having to read
src/actions/homebrew.rs source directly to find a real action id string.
applies_to is a direct projection of each action’s own Action::applies_to,
never a hand-maintained list.
glomeris actions list --json
{
"actions": [
{ "action_id": "cargo.clean.target_dir", "applies_to": ["cargo_target_dir"] },
{ "action_id": "node.clean.node_modules", "applies_to": ["node_modules"] },
{ "action_id": "homebrew.cleanup.cache", "applies_to": ["homebrew_cache"] }
]
}
glomeris actions history [--json] [--limit <N>]
Reads back a bounded, oldest-first tail of actions.jsonl (HORO-1057) — the
real-execution audit trail that execute, free, emergency and
autopilot run each append to, best-effort, after their own outcome is
already decided. Unlike history.tsv (which records pressure transitions
only), this is the audit trail of what was actually executed: action id,
resource id, the policy label it was authorized under, outcome, abort reason
(when applicable), actual reclaimed bytes, and which real-execution path
produced it.
--limit <N> is optional and defaults to 20, same bounding/malformed-line-
skip/missing-file-empty contract as glomeris history. An audit-write
failure never affects the execution it was trying to record — the write is
best-effort and its result is never surfaced to the caller.
With --json, prints an ActionHistoryReport ({"events": [...]}); each
event has timestamp, action_id, resource_id, policy_label,
outcome, abort_reason, actual_reclaimed_bytes, actual_reclaimed_human,
and source. Without --json, prints one line per event as plain text.
source is one of five values, produced by ActionSource::as_str in
src/monitor/persistence.rs:
source | The path that executed it |
|---|---|
execute | glomeris execute, one action against one named resource |
free | glomeris free --target, the recovery loop |
emergency | glomeris emergency, machine-wide, AUTO_SAFE only |
autopilot_auto_safe | glomeris autopilot run, action policy allowed on its own |
autopilot_preauthorized_ask | glomeris autopilot run, action attempted only because autopilot enable --preauthorize-ask had already named that kind and reason |
The last two are deliberately distinct rather than one autopilot value: the
question an audit trail has to answer is not just what ran but who
permitted it, and a pre-authorized ASK was permitted by the operator
naming that resource kind, not by policy alone.
Note “attempted” in that last row. A record is written for a failed or aborted
attempt too — not for a refused one, which never reached the filesystem and
so appears in autopilot run’s own report instead. Today every
pre-authorized ASK aborts at
deletion-time revalidation — so autopilot_preauthorized_ask currently only
ever appears alongside outcome: "aborted_by_revalidation". That is a known
limitation with a named cause and a pinning test, not the intended end state:
see Known Limitations.
glomeris actions history --json --limit 2
{
"events": [
{
"timestamp": 1700000000,
"action_id": "cargo.clean.target_dir",
"resource_id": "cargo_target_dir:/Users/dev/proj/target",
"policy_label": "AUTO_SAFE",
"outcome": "succeeded",
"abort_reason": null,
"actual_reclaimed_bytes": 2147483648,
"actual_reclaimed_human": "2.0 GB",
"source": "execute"
},
{
"timestamp": 1700000600,
"action_id": "node.clean.node_modules",
"resource_id": "node_modules:/Users/dev/proj/node_modules",
"policy_label": "ASK",
"outcome": "aborted_by_revalidation",
"abort_reason": "ResourceIdentityChanged",
"actual_reclaimed_bytes": null,
"actual_reclaimed_human": null,
"source": "free"
}
]
}
glomeris free (--goal-used-percent <N> | --target <N%|NB>) [--dry-run] [--json] [--progress-json] [--stop-file <path>] [--autopilot [--unattended]] [--project-root <path>]...
macOS only (exits 1 with an error message on other platforms). Exactly one of the two goal flags is required; any other argument is rejected (usage printed to stderr, exit 2).
The two axes
The goal can be stated on either axis, and the two flags are not interchangeable wordings of one thing:
| Flag | Axis | Meaning |
|---|---|---|
--goal-used-percent <N> | disk used | Stop when the volume is at most N percent used. |
--target <N%|NB> | free space | Stop when at least this much of the volume is free. |
--goal-used-percent 60 and --target 40% request the same end state.
--goal-used-percent is the product-facing form — it is the number a person
reads off a status bar, and it is what the Glomeris app sends. --target is
the raw free-space floor the recovery loop itself works in, and its meaning
and value formats are unchanged.
Passing both exits 2 rather than picking one: they are different numbers about how much of a disk to delete, so there is no safe precedence between them.
Accepted --target value formats:
- A percentage: a number followed by
%, in the range0–100(e.g.--target 15%). Target is “at least this percent of total capacity free.” - An absolute byte amount: a number optionally followed by
B,KB,MB,GB, orTB(case-insensitive; no suffix means raw bytes). Multipliers are binary/1024-based —--target 5GBmeans5 * 1024^3bytes free, not5 * 10^9.
--goal-used-percent takes a plain number from 0 to 100; a trailing %
is tolerated. It is additionally checked against the current reading and
refused with exit 2 if it is not an improvement on current usage, because
recovering toward a goal you already satisfy would delete nothing and still
print a report that reads like a successful cleanup. That check runs before
the execution lock is taken and before anything is deleted. --target
deliberately keeps its older, looser behaviour: a free-space floor you already
exceed is a legitimate no-op probe.
Other flags
--dry-runprints the pre-flight for the goal — current usage, free bytes still needed, and the estimated reclaimable opportunity split by what policy would actually permit — then exits. It takes no execution lock and mutates nothing. A--targetpre-flight is rendered on the used axis too, so no surface has to show a percentage whose axis is unstated.--jsonprints a machine-readable report instead of prose. With--dry-runthat is aRecoveryPreviewReport; without it, aRecoveryRunReport. A goal refused before the run starts prints aRecoveryGoalRejectionReport(reason,message,goal_used_percent,current_used_percent) and still exits 2. A busy execution lock prints the same{"reason": "busy", ...}refusalexecute --jsondoes, and exits 75.--progress-jsonstreams NDJSON progress on stderr — one complete object per line, never on stdout, so a--jsonreport stays parseable as exactly one document. With--dry-runit is the discovery-scan streamdetectemits. Without it, it reports the recovery loop itself; see Watching a run below.--stop-file <path>asks the loop to stop after the action it is currently running, oncepathexists. See Stopping a run.--autopilotbounds the run by the stored Autopilot envelope instead of by this command’s own limits. See Running inside the grant.--unattendeddeclares that nobody asked for this run. Only valid with--autopilot, and it needs a second permission the grant states separately. See Runs nobody asked for.--project-root <path>is optional and repeatable, same meaning asdetect’s flag above — it feeds the sameDiscoveryContextthe recovery loop discovers candidates from.
Watching a run
With --progress-json, a real run emits one JSON object per line on stderr as
it proceeds. Every line carries a phase and the 1-based iteration it belongs
to:
phase | When | Also carries |
|---|---|---|
measured | The volume was read (loop step 1). The only statement of fact about free space in the stream. | total_bytes, free_bytes, used_percent, free_human, bytes_freed_so_far |
discovering | Detectors are being asked what exists now. Re-entered every iteration — no pass reuses an earlier one’s list. | — |
discovered | That pass finished. | candidates (how many have a resolvable action), detectors_failed |
revalidating | Evidence is being re-collected and reclassified before anything is chosen. | — |
action_started | A real mutation is about to run. | resource, action, policy_label, estimated_bytes |
action_finished | It finished. | resource, action, outcome, reclaimed_bytes, bytes_freed_so_far |
stop_requested | A stop was observed, between actions. | — |
Two rules hold across every line:
estimated_bytesis the only estimate.reclaimed_bytesandbytes_freed_so_farare measured from the filesystem. Never accumulate the estimate as progress — it is what the candidate claimed, not what happened.- A count that could not be measured is absent, not zero. An
action_finishedline with noreclaimed_bytesmeans the size was not determinable;"0 B"there would be a measurement nobody took.
These phase names never collide with detect’s discovery stream, so a client
reading both cannot mistake a scan for a run.
Stopping a run
--stop-file <path> is a sentinel: create the file and the loop stops after the
action it is currently running, finishing with
stop_reason: "stopped_by_user" and a complete report.
- The path must not exist yet (exit 2 otherwise). A sentinel left behind by an earlier run would stop the next one before it did anything, and the report would truthfully say the user stopped it while the user had done nothing.
- Cooperative, deliberately. Nothing here can interrupt a deletion mid-flight. A signal delivered partway through one would leave the filesystem in a state neither the loop nor the audit log could describe, so the loop asks between actions instead.
- Not valid with
--dry-run, which performs no actions to stop.
Running inside the grant
--autopilot runs this same loop under the standing grant
glomeris autopilot enable wrote. It starts no second loop and introduces no
second policy: every candidate is still classified, revalidated and executed
exactly as it would be otherwise, and the envelope can only withhold one.
- Purely subtractive. Anything an
--autopilotrun does, a run you started yourself would also have done. Nothing the envelope says can make aPROTECTEDresource executable, admit anASKwhose exact kind and reason were not pre-authorized, or reach past a scoped-path or revalidation check. - No grant, no run. Defaults grant nothing, so with Autopilot revoked or
never enabled the command exits 3 having attempted nothing — the same code
autopilot runuses, so a script can tell “not authorized” apart from a usage error (2) and from a failed run (1). A revoked envelope refuses exactly as an absent one does;revokekeeps the limits on file, and onlyenabledstands between them and a run. - Its limits replace this command’s, and are the stricter of the two. The
run-level action count and wall-clock budget come down to the envelope’s own
figures. Left at their defaults, the loop’s 600-second ceiling could cut short
a run the grant had authorized for the full 15 minutes and report
budget_exceeded— a true sentence about the wrong budget. - The envelope is printed first, on stderr, before anything is discovered. Output that may end in deletions opens with the authority it acted under rather than asking you to go and look it up afterwards. stdout still carries exactly one report.
- A stop the envelope caused says so.
stop_reason: "envelope_refused"with anenvelope_refusaltoken naming which limit it was — an exhausted action, byte or time budget, a kind outside the grant, a pressure floor not met, a grant revoked mid-run. That is deliberately notsafe_exhausted: “this disk has nothing safe left” and “Autopilot reached the limit you set” call for opposite next steps, and only one of them is a finding about the disk. - Not valid with
--dry-run(exit 2). An envelope is authority to execute; a preview executes nothing, so the flag would have nothing to narrow and the candidate list would look envelope-filtered while being the whole of it. Read the grant withglomeris autopilot show.
Runs nobody asked for
An envelope answers “what may a run do”. --unattended is about a different
question — “may a run start when nobody pressed anything” — and that is a
second permission, granted by
glomeris autopilot enable --respond-to-alerts and off by default.
- Two consents, not one. An enabled envelope alone does not authorize an
unprompted run, and
--unattendedwithout that permission exits 3 having attempted nothing. The reason is what the alternative would mean for a grant already on disk: somebody wrote it to bound a run they intended to start, and reading it as consent to act while they are away would widen a grant they already read and approved. - The check is here, in this binary.
--unattendedis the caller’s own statement about itself, and it exists so that whatever launched the process cannot be the thing that decides it was allowed to. Without the flag, a client starting an--autopilotrun on its own initiative would be indistinguishable from a person typing the same command. - It narrows nothing else. Every budget, the kind allowlist and the ASK
pre-authorizations apply to an unprompted run exactly as they do to one you
started.
--unattended --autopilotis not a mode with different limits; it is the same bounded run with one additional permission checked before it begins. revokestops it along with everything else, without clearing the preference — so revoking never looks as though it had silently reset what you chose.- Only with
--autopilot(exit 2 otherwise). A run nobody asked for is precisely the one that must be bounded by a grant somebody read, so the combination that would be unbounded is refused at the boundary rather than later.
Reported figures
Every percentage in the output names its axis (60% used (40% free),
20% free). Every byte figure in a finished run is measured, and
target_met is decided from the re-measured free space after the run, never
from the sum of what the detectors estimated. The reclaimable figures in a
--dry-run preview are estimates and are labelled as such; protected space is
counted but never added into any opportunity total.
The prose output prints a RecoveryReport: the goal or target, stop reason
(as both a token and a sentence — a run that stopped because nothing safe
remained says so rather than printing a success word), iterations run, actions
executed, actions declined/skipped, bytes freed, and free space before/after.
A run that stopped because nothing safe remained also reports what is still
there, as four separate counts rather than one total: candidates awaiting your
confirmation, candidates whose action cannot run against the resource as it
currently stands (a live tool, work in progress), protected resources, and — for
an --autopilot run — candidates this grant was not authorized to take. Each is
a different next step, so they are never summed, and the fourth is kept apart
from the other three for a further reason: those are facts about what is on the
disk, and that one is a fact about the run’s authority. Folded into the protected
count it would tell you something is off limits when it is one setting away. They
are counts of candidates, not bytes, because space the run was not permitted to
take is not an opportunity.
Under --json this is the remaining object, present only for
stop_reason: "safe_exhausted" and for an envelope_refused stop that had
already discovered candidates: no other stop concluded anything about the
candidates it never reached, so zeros there would be a claim the run did not
make. An envelope refused before discovery — revoked, or a pressure floor not
met — omits the object entirely for exactly that reason, rather than publishing a
breakdown of zeros that would read as a completed search.
See Safety Model and
Known Limitations for what this loop can and cannot
currently do end to end.
glomeris autopilot <show|enable|revoke|run>
The subcommand is optional and defaults to show, so a bare
glomeris autopilot reads the grant rather than acting on it. run is macOS
only (exits 1 with an error message on other platforms); the other three verbs
work anywhere.
| Verb | What it does | Can it delete? |
|---|---|---|
show | Prints the stored envelope and its path. Does not create the file it reads. | No |
enable | Writes a new envelope from this command line’s flags and turns Autopilot on. Requires --kinds. | No |
revoke | Turns Autopilot off, keeping the limits. Effective for the next run; nothing to restart. | No |
run | Considers discovered candidates within the envelope. Deletes unless --dry-run. Holds the execution lock. | Yes |
The flags, their defaults and their hard ceilings are documented in Autopilot, which is also where the argument for why an LLM plan file cannot expand authority lives. What this page adds:
--max-bytestakes raw bytes only — unlikefree --target, there is noGB/MBsuffix parsing here, because an envelope is written once and read many times and an exact number is easier to audit than a rounded one.--min-pressureaccepts anyPressureStatename case-insensitively (healthy,warn,pressured,critical,emergency) plus the literalnone. An unobservable reading fails any floor you set, rather than passing it.--respond-to-alertsbelongs toenableand takes no value. It grants the one thing the limits above cannot express: that a run may begin without being asked, in answer to a disk-pressure alert (free --autopilot --unattended). Off by default and deliberately not implied byenableitself; see Runs nobody asked for for why those are two consents. It changes nothing about what a run may do once it starts, andrevokesuspends it with the rest of the grant while keeping it on file.--preauthorize-askis repeatable and takeskind:reasonusing the same tagsglomeris actions listandglomeris explainprint. Like the four limit flags above it, it belongs toenable—runaccepts only--dry-run,--plan-fileand--project-root, and exits 2 on anything else, so a run cannot widen its own grant on the command line.--jsonbelongs toshow,enableandrevoke, and prints one shape from all three: the grant (enabled,allowed_kinds,ask_preauthorizations, the four limits,min_pressure, and the two unprompted-run fields below), the hardceilings, every choiceenablewould accept (allowlistable_kinds,preauthorizable_reasons,pressure_states), what no envelope can authorize (never_allowlistable_kinds,never_preauthorizable_reasons,never_executable_labels), theai_authoritysentences, andstored_at. It always describes what is in force after the command ran —enablereports from what reached the file, not from what it was about to write — so a client reads one shape to learn one thing.rundoes not take it: the one reason to add it would be to make triggering a run from a GUI easier, and Menu Bar App explains why there is no such button.- The JSON carries two fields about unprompted runs, and they answer
different questions.
respond_to_alertsis the setting as you left it, in force or not — the one a settings toggle binds to, so that revoking Autopilot does not read as having silently cleared the preference.starts_unpromptedis whether an unprompted run is authorized right now: this grant is in force and it says so. Anything that acts consults the second, and is not the thing that computes it —enabled && respond_to_alertsevaluated in a client would be that client deciding its own authority. runprints anAutopilotReport: one line per candidate with its kind, policy label, model rank and outcome, then the run totals (actions attempted, actions succeeded, bytes freed, whether a budget stopped it early) and the envelope it ran under. Refusals appear here and nowhere else —glomeris historyis a log of what happened to the filesystem, and a refusal did not touch it.
Exit codes: 0 success, including a run that found nothing it was allowed to
do; 1 an action failed, or the envelope file could not be read or written;
2 usage error, including an unknown resource kind, a limit above its ceiling,
or enable without --kinds; 3 Autopilot is not enabled, so nothing was
attempted; 75 another invocation holds the execution lock.
3 exists so that “no grant” is distinguishable from “granted, ran, found
nothing” in a script — both of which are quiet, and only one of which means
the user has something to configure.
glomeris settings <show|set> [--notify-at-used-percent <N>] [--default-goal-used-percent <N>] [--json]
The subcommand is optional and defaults to show, so a bare
glomeris settings reads your preferences rather than changing them.
| Verb | What it does | Can it delete? |
|---|---|---|
show | Prints both preferences and their path. Does not create the file it reads. | No |
set | Stores one or both preferences. Requires at least one flag. | No |
Two preferences live here, and the whole point of the surface is that they are not the same number:
| Preference | Question it answers | Range | Default |
|---|---|---|---|
--notify-at-used-percent | When should Glomeris call my attention to disk usage? | 1–99 percent used | 75 |
--default-goal-used-percent | Where should recovery stop? | 0–100 percent used | 70 |
Both are percent of capacity USED, never free — the same axis
free --goal-used-percent takes, converted to the core’s percent-free
--target by the single adapter documented in the glomeris free section
above. Every line this command prints names its axis, in prose and in JSON,
because 75 beside 70 with no label is the one misreading this feature
cannot afford.
What this page adds:
- The goal must be below the threshold. A goal at or above the point that
raised the alert would be satisfied the moment it was announced, so the pair
is refused with the reason
goal_not_below_notify_threshold. - Both flags are validated as a pair, in one step. Moving from
(75, 70)to(60, 55)is a valid destination that no single-field order can reach — whichever field moved first would be momentarily invalid against the old value of the other. Sosetrefuses you for the destination you asked for, never for the order your flags happened to appear in. - A trailing
%is accepted on either value, because a user typing what they read on screen has not made a mistake. setwith neither flag is a usage error rather than a successful no-op: a command line that did not say what it wanted should not report that it did it.- The four disk-pressure states (
healthy,warn,pressured,critical,emergencyboundaries) are not configurable and this command cannot reach them. That is deliberate: if thewarnboundary were a preference, lowering it would change what awarnrecorded last week meant, and raising it would make the next poll report a transition the disk never made. The alert threshold is a separate scalar layered over that machine, and its default tracks thewarnboundary so the product is no quieter than it is today. - Neither preference is an authorization. Raising a goal cannot make a
PROTECTEDresource deletable, cannot bypass anASK, and cannot widen an Autopilot envelope. These two numbers decide when the product speaks and where recovery aims; every gate between a candidate and its deletion is elsewhere. --jsonbelongs to both verbs and prints one shape from both:notify_at_used_percentandnotify_at_descriptionfor the threshold, adefault_goalobject carryingused_percent,free_percentanddescriptionfor the goal, aboundsobject carrying the four limits above asnotify_at_minimum_used_percent,notify_at_maximum_used_percent,goal_minimum_used_percentandgoal_maximum_used_percent,stored_at, andloaded_from_file.boundsis there so a settings screen can offer a control that cannot compose a value this command would refuse, without knowing the limits independently and drifting from them; the cross-field rule is deliberately absent from it, because a limit that moves as the other number moves is not a bound on either one. It always describes what is in force after the command ran.loaded_from_fileis reported rather than anis_defaultflag, because a user may legitimately store the default numbers and a client must be able to tell “nothing chosen yet” from “these were chosen”. A refusedsetprints a rejection object on stdout instead —reason(a stable token),message, and only whichever ofnotify_at_used_percent/goal_used_percentthe refusal actually involved.- The file is
~/Library/Application Support/Glomeris/settings.conf, beside the Autopilot envelope and in the same versioned format: aversionkey, onekey = valueper line, unknown keys an error rather than noise, and loading routed through the validating constructor so a hand-edited file cannot install a pair the CLI would have refused.
Exit codes: 0 the preferences were printed, or the change was stored; 1 the
settings file could not be read or written, or what it contains is not valid;
2 usage error, including set with neither flag, or a value on this command
line that was refused.
A refused stored file is 1, not 2: it is not this command line’s mistake,
and a script must be able to tell “you typed 120” from “the file on disk
disagrees with itself”.
glomeris external-context [--json]
What optional remote context may reach the model, and what never will. Read-only and off by default: with nothing configured it prints that nothing leaves this machine, which is a fact about the shipped product rather than a hypothetical.
Glomeris can optionally read a pull request’s state from GitHub and a work item’s state from Jira, as supporting evidence for ranking a working tree that looks stale. Neither is deletion authority — see the list at the end of this section — and neither is on until you write the configuration file.
This command asks nothing. Answering “what may leave” by leaving is the one shape it must not have: you would be making the very requests you are trying to understand, against a service that logs them, before deciding whether you want that. Everything it prints comes from your configuration file and from the value vocabularies the code itself serializes.
Three kinds of line, and the difference between them is the point:
| Section | What it describes |
|---|---|
| Providers | Your local setup. Whether each provider is configured, whether it is usable, the environment variable its credential is read from, the endpoint, and for GitHub the one repository host in scope. |
| Fields that may be sent | What could actually leave, one line per serialized key, each with its complete set of possible values. |
| Never sent | The explicit negative. |
What this page adds:
- “Not configured” and “configured but not usable” are different rows. A
provider whose credential environment variable is unset is reported as
configured with a refusal naming the variable, not as absent. The same
distinction runs all the way through the feature: a provider that could not
answer never reports an absence, and
no pull request was foundis a different value fromGitHub could not be reached. - The field list is the complete set, and it is derived rather than written.
Each state vocabulary comes from the enum the projection serializes, so a
variant added later cannot widen what really travels while this page keeps
understating it. A remote state reaches a model as a bounded token —
merged,in_progress,none_observed— with how many days ago it was read, and with the kind of failure when a provider could not answer. - Nothing identifying goes. Not the repository, the branch, the issue key, the pull request’s number, title, body or author, your account, or any path. The command prints that negative explicitly, because a list of what travels does not answer “did you send my branch name”.
- Your Jira account address is not echoed, even though the configuration file holds it. It identifies a person rather than a setting anybody debugs, and this report gets pasted into bug reports. The site is shown, because you need to see which one is asked.
- No credential value appears anywhere, in prose or in JSON. The variable is named; the value is never read by this command at all.
- Jira has no host scope and does not claim one. Correlation there requires
an explicit, deterministic issue key in the branch name —
HORO-1234, not a fuzzy match fromfix-storage-stuff— so there is no remote to compare against. GitHub’srepository_hostis the most privacy-relevant line in the report: a working tree whose remote is anywhere else is never named to that provider at all. - None of it is permission. A merged pull request and a closed ticket are
context for ranking. They cannot make a working tree deletable: a dirty
tree, an untracked file, a local commit made after the merge, or a process
holding the directory open each keep it protected or uncertain whatever a
remote service says. That is enforced in the policy engine, which never sees
external context, and pinned by
tests/external_context_grants_no_authority.rs. - Read-only by construction, not by convention. The providers can name
exactly one transport, whose whole vocabulary is a single GET;
scripts/check-external-context-is-read-only.shproves the module names no mutating verb, no second transport, and nothing that prints or writes. --jsonprints the same report — each provider’s configured and usable state, its credential variable, endpoint and host scope, every field that may be sent with its vocabulary, and thenever_sentlist. This is what the menu-bar app’s privacy preview reads.- The file is
~/Library/Application Support/Glomeris/external-context.conf, besidesettings.confand in the same versioned format.
Exit codes: 0 the report was printed; 2 usage error — this command takes no
arguments besides --json.
An unparseable configuration file is 0, not 1, which is the opposite of
settings show. The two commands answer different questions. settings show is
asked “what are my preferences”, and an unreadable file means it has no answer.
This one is asked “what could leave this machine”, and an unreadable file has a
complete and reassuring answer — nothing is configured, so nothing leaves — that
a non-zero exit and a bare stderr line would throw away. The refusal is reported
as the first line of the report instead.
glomeris pressure <show|notified|respond> [<answer>] [--json]
The seam between the background monitor and the menu-bar app. You are unlikely to type it; the app runs it every poll.
It exists because of a constraint neither side can work around alone: only an
app bundle can put buttons on a macOS notification, and the monitor
(glomeris daemon) is a bare launchd process, not a bundle. So the monitor
records that a notification is owed and the app raises it. This command is
how the app asks what is owed and reports which button was pressed.
The subcommand is optional and defaults to show, so a bare
glomeris pressure reads rather than writes.
| Verb | What it does | Can it delete? |
|---|---|---|
show | Prints the current episode, whether a notification is owed, your alert threshold, your recovery goal, and the answers that can be given. | No |
notified | Records that the notification was put on screen, so it is not raised again for the same episode. | No |
respond | Records which button was pressed. Takes one of the three answers below. | No |
What an episode is
One continuous stretch of disk usage being at or above
settings --notify-at-used-percent. It opens the first time an observation
reaches that threshold and closes only once usage has fallen three percentage
points below it. That gap is the hysteresis, and it is the whole reason one
spell of a full disk produces one notification rather than one per poll: a
volume hovering exactly on the threshold stays inside one episode, and it takes
a real recovery rather than a rounding wobble to end it.
Episode ids are monotonic and never reused, because the app uses the id as the notification’s identifier — two notifications about one episode coalesce, and two episodes never do.
The three answers
| Answer | What it means |
|---|---|
review_and_recover | Open the recovery screen for this episode. Opens it — it does not start recovery, and deletes nothing. |
remind_later | Say nothing for two hours, then notify again if usage is still above the threshold. The episode stays open, so this is a delay, not a dismissal. |
ignore_episode | Say nothing more about this spell of disk pressure. Monitoring continues and the next episode notifies again. |
Hyphens are accepted too (respond review-and-recover), because that is what a
command line looks like. Nothing else is: an unrecognized answer is refused
rather than mapped onto the nearest match, since the nearest match to a
misspelled ignore_episode is the one that opens recovery.
ignore_episode cannot switch notifications off. It is stored on the
episode, so it expires with it — there is no answer here, and no flag anywhere,
that silences monitoring indefinitely. That is the point of storing the answer
on the episode rather than in settings.
What this command does not do
- It does not observe.
shownever opens or closes an episode and never decides that a notification is owed; only the monitor’s own poll does. On a machine where the monitor is not installed there is therefore no episode at all, andshowsays so while still reporting whether usage is above your threshold. - It does not delete anything, and it cannot widen what may be deleted. An
episode decides when you are spoken to, never what is permitted. Answering
review_and_recoveropens a screen. Every gate between a candidate and its deletion is elsewhere and unchanged. - It does not start recovery, automatically or otherwise. Crossing a threshold raises a question; a human answers it.
--json
Belongs to all three verbs and prints one shape from all three, always describing the state after the command ran. This is what the menu-bar app reads.
current carries the disk reading with both percentages under separate names
and the byte figures; notify_at_used_percent/notify_at_description the
threshold; a default_goal object the recovery goal, in the same shape
settings --json prints it; threshold_crossed and notification_due the two
facts the app acts on; episode the episode’s identity, its opened/peak/latest
readings, how many notifications have been raised, the answer if one was given,
and the snooze deadline if one is pending; responses the answers that can be
given, so the app offers buttons it did not invent.
A refusal prints a rejection object on stdout instead — reason as a stable
token (no_open_episode, no_notification_due) and message — rather than
writing to stderr, because a refusal here is an ordinary outcome that the app
has to read and display.
Exit codes: 0 the episode was printed, or the answer was recorded; 1 the
episode state or your settings could not be read or written, or filesystem usage
could not be measured; 2 usage error, including an answer that is not one of
the three; 3 there was nothing to record.
3 rather than 1 for that last case, and this is the load-bearing part of the
contract: no episode is open, or no notification was owed. It is not a
failure. The usual cause is that the disk recovered while the notification was
on screen, and the app polls this command every thirty seconds — a client that
treated it as an error would retry at poll speed forever.
Exit codes
Also available as glomeris help exit-codes, which is the copy to trust — it
renders from the same table as the per-command help, so a command’s exit codes
cannot drift between its own --help and the summary.
0— success.1— a macOS-only command was run on a non-macOS platform, a platform-level operation (e.g. reading thelaunchdplist path) failed, or (forllm-planspecifically) the LLM provider call/response parsing failed, or--plan-filenamed an unreadable path, or (forllm-checkspecifically) the check ran and did not pass.2— usage error: unknown top-level command, unknowndaemonsubcommand, missing/unrecognizedfreearguments, or (forllm-plan/llm-checkspecifically) an unrecognized argument, a missing flag value, missing live-mode LLM environment configuration, a base URL that cannot work, or an--api-key/--key/--tokenflag.3— (autopilot runandfree --autopilotonly) this run was not authorized, so nothing was attempted: either Autopilot is not enabled, or--unattendedwas given and the grant does not allow starting a run unasked. One code for both because to a script they are the same fact.75— (free/emergency/execute/autopilot runonly) the HORO-1054 execution lock is already held by anotherglomerisinvocation.
glomeris execute has its own, more specific set of exit codes (0–5
plus 75) — see its own section above for the full table; a couple of
those codes (1, 2) overlap this list’s meanings but are worth reading
in full since execute is the one subcommand with real destructive
consequences.
See glomeris llm-plan’s and glomeris llm-check‘s own sections above for
those subcommands’ exit codes in full detail — llm-check’s 1/2 split in
particular carries a meaning this list cannot: whether a report exists.
Pressure Model
The monitor module implements pure disk-pressure state logic plus the
polling loop that drives it. All platform-specific I/O (real statvfs,
real notifications) lives under platform::macos, which the monitor’s
traits abstract away.
PressureState
Defined in src/monitor/pressure.rs, in ascending urgency order:
#![allow(unused)]
fn main() {
pub enum PressureState {
Healthy,
Warn,
Pressured,
Critical,
Emergency,
}
}
Rendered (Display) as HEALTHY, WARN, PRESSURED, CRITICAL,
EMERGENCY.
Thresholds
ThresholdConfig::default() (src/monitor/config.rs) defines both a
percentage-used bound and an absolute free-bytes bound per state:
| State | used % ≥ | free bytes ≤ |
|---|---|---|
| Warn | 75.0 | 50 GiB |
| Pressured | 85.0 | 20 GiB |
| Critical | 92.0 | 10 GiB |
| Emergency | 97.0 | 3 GiB |
A threshold is “crossed” when either bound is crossed (percentage OR
absolute bytes) — not both. Classification checks most-urgent-first
(Emergency → Critical → Pressured → Warn → Healthy), so one reading that
crosses several boundaries lands on the single most urgent one. These are
library defaults; ThresholdConfig is passed as a parameter to monitor::run,
so a caller could supply different values, but nothing in this codebase does
so today.
Debounce / hysteresis
PressureStateMachine::observe() (src/monitor/state_machine.rs) requires
confirm_after consecutive matching observations of a candidate state
before confirming a transition. The production default is confirm_after = 2
(two consecutive one-minute polls, at the default 60-second poll interval).
- If the newly observed state equals the currently-confirmed state, any in-flight candidate is discarded.
- If a different candidate than the pending one appears, its counter resets to 1 — candidates are not averaged.
- Once the counter reaches
confirm_after, the transition confirms and aTransition { from, to }is emitted exactly once. - The same logic and the same
confirm_aftervalue apply in both directions — there is no asymmetry between escalating and recovering (no “escalate fast, recover slow” rule exists in this code).
This exists specifically to prevent a single noisy sample from producing a notification on every poll.
The polling loop
monitor::run() (src/monitor/poller.rs) starts the state machine assuming
Healthy, then loops: stat the watched path, classify the reading, feed it
to the state machine, and — only on a confirmed transition — call the
notifier and the persistence backend. Sleep happens between iterations.
max_iterations: Option<u64> bounds the loop for tests; production passes
None to run indefinitely.
Error handling is deliberately non-fatal at every layer:
- A failed
FsStat::stat()call is reported to theon_iterationcallback as an outcome-less iteration and does not stop the loop. - A notifier failure or a persistence-write failure is captured into the
poll outcome’s
notify_error/persist_errorfields — the loop keeps running either way. This is proven by a test that drives 20 successful loop iterations against an always-failing persistence backend.
Persistence
FilePersistence (src/monitor/persistence.rs) appends one tab-separated
line per confirmed transition (unix_time_secs, from, to,
used_percent, free_bytes) to a plain text file — not a database. Its own
module doc states it is “explicitly failure-tolerant: any error here must
never stop the monitor loop from continuing to observe and notify.” No
feature in this crate depends on persistence succeeding. Structured (SQLite)
persistence is deferred to a later ticket.
What the monitor does not do
The monitor is a cheap, O(1) capacity check (statvfs) — it never walks the
directory tree. Discovering what is reclaimable is the scanner’s and
detectors’ job (see Architecture), a deliberately separate
and more expensive path.
Evidence Model
Evidence (src/evidence/model.rs) describes what a detector observed
about a resource — never what should be done about it. That judgment
belongs entirely to the policy layer.
Evidence
Key fields:
#![allow(unused)]
fn main() {
pub struct Evidence {
pub resource: ResourceId,
pub fingerprint: ResourceFingerprint,
pub detector: DetectorId,
pub logical_bytes: ProbeOutcome<u64>,
pub physical_bytes: Option<u64>, // always None in the MVP
pub reclaimable_bytes: ProbeOutcome<u64>,
pub last_modified: ProbeOutcome<SystemTime>,
pub last_accessed: ProbeOutcome<SystemTime>,
pub regenerability: Regenerability,
pub recoverability: Recoverability,
pub native_cleanup: NativeCleanup,
pub open_by_process: ProbeOutcome<Vec<ProcessRef>>,
pub process_cwd_match: ProbeOutcome<Vec<ProcessRef>>,
pub git_state: ProbeOutcome<Option<GitState>>,
pub tool_liveness: ProbeOutcome<bool>,
pub collected_at: SystemTime,
pub sources: Vec<String>,
}
}
physical_bytes is always None today — subtree st_blocks summation is
out of scope for the current milestone. git_state’s Observed(None) means
the probe ran successfully and determined the resource genuinely isn’t in a
git working tree — a complete, legitimate answer, not a missing one; only
Unavailable(reason) means the probe itself failed.
ProbeOutcome<T> — the core honesty guard
#![allow(unused)]
fn main() {
pub enum ProbeOutcome<T> {
Observed(T),
Unavailable(ProbeReason),
}
pub enum ProbeReason {
ToolAbsent,
ToolNotRunning,
PermissionDenied,
TimedOut,
Failed,
NotAttempted,
}
}
This is deliberately not Result<T, E>: there is no Default impl and
no unwrap_or_default() escape hatch. A caller that wants the observed value
must explicitly branch on this enum — there is no way to silently coerce an
unavailable probe into a zero/empty/false value. This is the structural
guard against “a failed probe becomes a safe default.”
Completeness and Confidence
#![allow(unused)]
fn main() {
pub enum Completeness {
Complete,
Partial { missing: Vec<EvidenceField> },
Failed,
}
pub enum Confidence {
High,
Medium,
Low,
}
}
Evidence::completeness() checks every field ResourceKind::required_evidence()
lists for that resource’s kind: Complete if none are missing, Failed if
all required fields are missing, Partial otherwise. Confidence is
derived from completeness (Complete → High, Partial with ≤1 field missing
→ Medium, everything else → Low) and is advisory only — the policy layer
does not consume it; it exists to give a human or LLM something short to
point at.
Resource kinds whose owning tool has no persistent daemon to check
(Cargo/Npm/Pnpm/Yarn/Homebrew) do not require ToolLiveness for
Complete, since that field would otherwise be structurally unreachable —
see Known Limitations for a related caveat about
Docker build cache never reaching Complete at all.
ProbeOutcome never masquerades as “safe”
Nothing in this codebase treats missing or failed evidence as evidence of
“nothing to clean up.” A detector’s own status type
(DetectorStatus::Failed(reason)) is documented as meaning “we don’t know,”
never “nothing found,” and ResourceKind::Unknown is a deliberate fail-closed
sink that the policy layer maps to PROTECTED unconditionally rather than
falling through to any default treated as safe. The actual “therefore refuse
to delete” enforcement lives in policy::classify, not in
this module — this module only guarantees the data can’t fabricate a
misleadingly complete picture.
Correlation: open_by_process, process_cwd_match, git_state, tool_liveness
Detectors (detectors/) only ever populate the discovery-stage fields
(size, mtime, regenerability). The four correlation fields above are always
Unavailable(NotAttempted) straight out of a detector — a separate
collector, evidence::correlate::DefaultEvidenceCollector, is responsible
for actually attempting correlation. This split is what keeps freshly
discovered evidence from ever reporting Completeness::Complete before
correlation has actually run.
DefaultEvidenceCollector wires in the real, subprocess-backed probes:
| Field | Backed by | Command |
|---|---|---|
open_by_process | LsofOpenFileProbe | lsof -F pcn +D <path> |
process_cwd_match | LsofProcessCwdProbe | lsof -a -d cwd -F pcn +D <path> |
git_state | GitCliProbe | git -C <path> rev-parse --show-toplevel, then git status --porcelain and git rev-parse --git-dir --git-common-dir on the resolved repo root |
tool_liveness | PgrepToolLivenessProbe | pgrep -x <daemon-name> (only for tools with a real daemon: Xcode.app, Docker’s com.docker.backend; other tools skip the subprocess entirely and report Unavailable(ToolNotRunning)) |
Each probe runs under a shared per-call timeout (ProbeBudget) enforced by
polling try_wait() and killing the child on timeout to avoid leaving a
zombie process. A collect() call can spend up to roughly four times that
timeout in the worst case, since up to four subprocesses each get their own
full budget.
Absence and failure handling is per-tool, not global: a missing lsof,
git, or pgrep binary degrades that specific field to
Unavailable(ProbeReason::ToolAbsent) — it does not fail the whole
correlation pass, and it is never silently treated as “nothing found.” A
non-zero exit with no output (e.g. lsof finding nothing) is treated as a
legitimate empty result; a non-zero exit with stderr content is treated as
a genuine probe failure (Unavailable(ProbeReason::Failed)).
Docker’s correlation gap
DockerBuildCache/DockerImageCache resources use ResourceLocator::Tool
(a tool-native id, not a filesystem path) because there is no single
canonical path to a Docker build cache. DefaultEvidenceCollector can only
run tool_liveness for a Tool-locator resource; open_by_process,
process_cwd_match, and git_state are always Unavailable(NotAttempted)
for it. Since required_evidence() still requires all three, Docker build
cache cannot reach Completeness::Complete today — see
Known Limitations.
Safety Model
The three-class decision
policy::classify() (src/policy/engine.rs) is the single deterministic
module that decides AUTO_SAFE / ASK / PROTECTED for a resource. It is
pure: no I/O, no ambient clock (now is always a parameter) — which is what
makes fail-closed behavior testable and lets the executor call the exact
same function again on freshly re-collected evidence at deletion time.
#![allow(unused)]
fn main() {
pub enum PolicyClass {
AutoSafe,
Ask,
Protected,
}
}
There is no separate “Unknown” fourth outcome. Missing, stale, or failed
evidence maps inside classify to Ask or Protected with a specific
reason code — it never becomes a state a caller might misread as “not
Protected, therefore fine.”
classify()’s fail-closed ordering
classify checks conditions in this exact order, and the order is itself
part of the safety guarantee:
- Unknown resource kind → unconditionally
Protected. Never falls through to any evidence-based judgment. - Path/pattern-based protected check (see below) →
Protected, checked before any freshness/completeness logic. AProtectedclassification never depends on evidence freshness — a stale-but-protected resource is stillProtected, never “upgraded” by fresher evidence. - Staleness — evidence older than
PolicyConfig::max_evidence_age(default 5 minutes), or whosecollected_atis somehow in the future →Ask+EvidenceStale. - Completeness —
Failed→Ask+EvidenceProbeFailed;Partial→Ask+EvidenceIncomplete. - Active-use signals (most-significant first) — an open file handle or
matching process cwd →
ResourceInActiveUse; a dirty or linked git worktree →GitWorktreeDirty; a live owning-tool daemon →OwningToolLive. Any of these →Ask. - Per-instance regenerability —
NotRegenerable→Ask+RebuildCostHigh, regardless of how clean the rest of the evidence looks. - Otherwise →
AutoSafe, with reasonsEvidenceFreshAndComplete(+RegenerableByToolif applicable) +NoActiveUseObserved.
The PROTECTED matcher
policy::protected::protected_reason() is an explicitly conservative,
MVP-scope, non-exhaustive denylist, not a claim of completeness. It
performs no filesystem I/O (no canonicalize, no read_link, no
metadata) — it only inspects the path/locator text already carried by the
resource, which is what keeps classify pure. What it actually matches
today:
- Docker image cache — every
DockerImageCacheresource is treated as a possible persistent volume and classifiedProtectedunconditionally, because detectors do not yet distinguish a persistent volume from disposable image cache. (DockerBuildCacheis not covered by this rule.) - Credential material — any path component
.sshor.gnupg, or a filename ending.pem/.key, or containingcredentials. - Git internals — any path with a
.gitpath component (not merely “inside a git-tracked project,” which is the separate, evidence-drivenGitWorktreeDirtyreason). - Infra state — a
.terraformpath component, or a filename ending.tfstate/.tfstate.backup. - System paths —
/System,/usr(except/usr/local),/bin,/sbin,/private/var/db. - Unsafe mount/symlink targets —
/Volumes,/dev,/Network,/net.
This is a starting denylist to extend, not a guarantee that every dangerous path is covered.
ASK and consent
policy::approval::authorize() is the only way to construct an Approval
— the type an executor is required to hold before acting. It is built with a
private, unconstructible-outside-the-module marker type, so nothing outside
approval.rs can fabricate one.
AutoSafe→ always authorizes, no consent needed.Protected→ never authorizes, unconditionally, with no override parameter that can change that.Ask→ authorizes only if a matchingUserConsentis supplied: the consent’s resource identity and fingerprint must match exactly. Consent granted for one resource instance does not carry over to a different instance at the same path, or to the same resource after its underlying fingerprint has changed.
There is no TTY prompt, but consent is no longer unreachable. Three
surfaces supply a UserConsent today, and none of them is a prompt:
glomeris execute --confirm-ask --observed-fingerprint <token>— you pass back the exactfingerprint_tokena priorexplain --jsonprinted on that same resource. Mismatched or missing, it is a usage error rather than a silent approval.- The menu-bar app’s Clean button, which builds precisely those two flags
(
CandidateDetailView.buildExecuteArguments) and nothing else. The GUI is the confirmation step; the consent still travels as a fingerprint. glomeris autopilot enable --preauthorize-ask <kind>:<reason>— narrow advance consent for oneAskreason on one resource kind, recorded in the envelope. It is a flag onenable, not onrun: the consent is written down before the run andautopilot runrejects the flag outright, so the run cannot grant itself anything the stored envelope does not already say. See Autopilot, and the limitation on that path in Known Limitations.
What has not changed is glomeris free’s recovery loop: it wires
auto_approve_ask: false (src/main.rs), so inside that loop Ask
candidates are still reported as declined/skipped and never executed. That is
the loop’s own choice, not an absence of machinery.
Deletion-time TOCTOU revalidation
executor::execute() never trusts a previously computed PolicyDecision at
face value. Immediately before mutating anything, it:
- Snapshots the resource’s live filesystem identity via
symlink_metadata(nevermetadata, so a symlink is detected as a symlink, never resolved through) — captured before the slower correlation re-probe runs, so a symlink swap during that window can’t backdate the anchor. - Re-collects evidence from scratch (re-probes size/mtime/fingerprint
directly, re-runs the same
EvidenceCollector). - Aborts (
ResourceIdentityChanged) if the fresh fingerprint doesn’t match the one the approval was granted against. - Re-runs
classify()on the fresh evidence. - Aborts (
PolicyClassDowngraded) if the class no longer matches what was approved. - Aborts (
PolicyReasonsWidened) if any new reason appears that wasn’t present at approval time —Askis a heterogeneous bucket, so consent granted againstRebuildCostHighdoes not cover a freshly observedResourceInActiveUse. - Aborts (
EvidenceDegraded) if completeness got worse since approval — a defensive, forward-looking guard. - Only then builds the actual plan from the same fresh evidence just
validated, and runs it. Immediately before any
RunTool/DeletePathstep actually mutates the filesystem, the resource’s identity is re-verified one more time against the early snapshot from step 1.
A plan with more than one step is refused outright rather than executed, because partial-deletion byte accounting has no way to report a correct total if a later step fails after an earlier one already succeeded. No registered action emits more than one step today.
AUTO_SAFE is end-to-end reachable through real execution
Detectors populate reclaimable_bytes at discovery time (HORO-992), and the
deletion-time revalidation path (executor::build_fresh_evidence) reuses
the same bounded recursive size estimate (HORO-1016) for reclaimable_bytes
that it already used for logical_bytes (HORO-994) — it no longer hardcodes
the field back to Unavailable. A real AutoSafe approval built from a detector’s
evidence genuinely survives revalidation and executes for real, proven by
tests/golden_chain_execute.rs. See Known Limitations
for what’s still out of scope (Docker build cache’s own completeness gap).
Actions never receive raw commands
Every registered Action produces a typed ActionPlan/ActionStep from
Evidence — never a caller-supplied string. ActionStep::RunTool invokes
one of a closed set of binaries (ToolBinary::{Brew,Cargo,Npm,Pnpm,Yarn}) via
argument-array Command::new(tool).args(args) calls — never a shell string.
ActionPlan/ActionStep deliberately never derive Deserialize, so no
external input (network payload, LLM text) can ever materialize one directly;
the only way external input reaches execution is by selecting a
pre-registered ActionId, whose real Action::plan implementation then
decides the actual steps from Evidence it independently trusts. See
BYOK LLM Planner for how this applies to the optional LLM path
specifically.
An envelope narrows this model; it never widens it
glomeris autopilot adds the one thing the rest of this page does not
describe: a grant that outlives the moment you typed it. It changes nothing
above. An Autopilot candidate is classified by the same classify(),
authorized by the same policy::approval::authorize, and revalidated by the
same deletion-time TOCTOU check, in that order.
What the envelope adds is a filter in front of authorization, not an
alternative to it. Its allowlist can only remove kinds from consideration;
PROTECTED is refused whatever it says; UNKNOWN_INCOMPLETE is refused
whatever it says; an ASK reason it has not been given by name is refused. The
budgets — actions, bytes, wall clock, and an optional disk-pressure floor —
bound a run’s total effect, so the worst case of an unattended run is a number
you read before you granted it.
The clause this adds to the invariant is the last one: AI can recommend. Policy decides. Executor verifies. Filesystem reality wins — and the envelope bounds the outcome. See Autopilot.
Emergency Mode
glomeris emergency (src/emergency/mod.rs, macOS only) is a degraded
recovery path that must produce a useful result even when SQLite/history/log
writes fail, there is no network, there is no LLM provider configured, and no
GUI is running. It is not glomeris free --target with a flag — its scope
and limitations are deliberately narrower.
The menu-bar app does not expose this command, deliberately: emergency takes
no arguments and acts machine-wide, which is not a thing to put behind a
one-click control. It is CLI-only.
The invariant is unchanged
Emergency pressure is not permission to weaken the safety model. This module
reuses policy::classify/policy::approval::authorize and
executor::execute exactly as they are elsewhere — it does not define any
separate, looser emergency-only rule set, and it never references anything
network- or LLM-related (enforced by a test that scans the module for such
references).
What it does, in order
- Frees its own disposable state first, unconditionally. Before
anything else, it deletes the tool’s own best-effort pressure-history
file (the same
history.tsvthe daemon appends to) — never gated on policy, because this file is not a developer resource under policy’s purview; it’s this tool’s own append-only log, and deleting it is safe by construction. A missing file is not an error: the function silently does nothing rather than fabricating work. - Discovers candidates via the same bounded detector registry used
elsewhere (
DetectorRegistry::discover_all) — never the scanner’s slower, unbounded filesystem walk. - Classifies and, for
AutoSafecandidates only, attempts execution through the sameclassify/authorize/executepipeline as normal operation.
AUTO_SAFE-only, no interactive Ask
A degraded, no-time, no-mechanism-for-consent path can only auto-execute
AutoSafe candidates. Everything else (Ask, Protected) is refused and
counted in EmergencyReport::denied_candidates — never escalated to
interactive consent. This is a deliberate MVP scope decision documented in
the module itself, not an oversight.
No LLM, no network, no required database
run_emergency never calls an LLM and never makes a network request — this
is enforced by a dedicated test, not just a comment. Persistence-write
failures (a plain write failure, or a failing directory creation) do not
stop the run: run_emergency still completes and still frees the
self-owned-state fixture, per its own fault-injection tests.
EmergencyReport
Every field is a plain, dependency-free primitive (u32/u64/String), so
the report stays printable even if every other subsystem in the process has
already failed:
actions_attempted/actions_succeededtotal_bytes_freeddenied_candidates— every candidate whoseclassify()result was notAutoSafe. Never incremented for an unrelated wiring gap (e.g. no registered action for anAutoSafekind) — those go toerrorsinstead.errors— non-fatal error strings, bounded to at most 8 entries so an adversarial run can’t grow this field without bound.
Documented limitations (from the module’s own doc comments)
- The preallocated emergency-reserve-file idea is deferred, not implemented. The originating ticket named it as an experimental idea to reject or defer without reliable measured evidence; this module does not build it.
- Resolved (HORO-994): the candidate loop now frees real bytes via a real
detector-produced candidate.
executor::build_fresh_evidence’s deletion-time revalidation reuses the same bounded recursive size estimate (HORO-1016) forreclaimable_bytesthat it already used forlogical_bytes, so a realAutoSafeapproval survives revalidation and executes. Both the self-owned-disposable-state step (step 1 above) and real detector-found candidates can now reclaim bytes. - Near-zero-real-disk-space testing was not performed. Driving a real machine’s free space to near zero to test this path was judged too destructive to be worth the risk; fault injection via fakes covers the equivalent failure modes instead.
BYOK LLM Planner
actions::llm (src/actions/llm.rs) is an optional, “bring your own key”
LLM planner. As of HORO-1008, it is wired into the glomeris binary as an
advisory, non-executing subcommand: glomeris llm-plan. This page
documents both the library module and the CLI surface over it.
Since HORO-1308 the menu-bar app has an AI Plan card over the same
subcommand. It is a thin client: it spawns glomeris llm-plan --json --progress-json, renders the report, and has no planner, no provider client
and no execution path of its own. Everything on this page — what leaves your
machine, what the validator drops, what the planner can and cannot do —
applies unchanged to the GUI, because it is the same code doing the work. See
Menu Bar App for how the card separates the model’s
words from the machine’s.
Since HORO-1309 the provider itself can be configured from that app (Settings → AI Provider) instead of only from the shell, with the key in the login keychain rather than an exported variable — see Configuring this from the menu-bar app below. The GUI is still a thin client there too: it writes three values and launches a child process with them. It does not validate the URL, does not know which models exist, and never speaks to a provider itself.
What it is
An advisory ranking suggestion over evidence the crate already collected.
LlmResourceView is an explicit, bounded projection of one Evidence
record — deliberately not a Serialize derive on Evidence itself, so
adding a field to Evidence later has no effect on what an LLM would see
unless a human explicitly adds it here too. plan_with_llm is the entry
point: it serializes a set of these views, calls a provider, and defensively
validates the response.
What leaves your machine
Two strings: a fixed system prompt and a JSON array of LlmResourceView
values. LlmRequestPayload is that request, and it is the whole outbound
surface — nothing about the local machine is added downstream of it.
Each view identifies its resource by a positional wire alias —
resource_1, resource_2, … — not by its real ResourceId. That matters
because a ResourceId for a path-backed resource kind renders as an
absolute path, and an absolute path under $HOME carries your OS account
name and your directory layout. Since HORO-1298, no path, home directory
or account name is transmitted; what the model gets is the resource’s
kind, reclaimable size, age, regenerability, evidence completeness and the
action ids offered for it, which is what a ranking judgement is actually
made from.
The alias is positional rather than a hash of the resource. A hash would be
stable across runs, which sounds better until you count the keyspace: a
path like /Users/<name>/Library/Developer/Xcode/DerivedData has one
unknown segment, so anyone holding a candidate account-name list can
confirm it against the hash offline. A stable pseudonym is also, by
construction, a handle for correlating your machine across requests. An
index is neither, and costs nothing — the model is ranking resources it was
just handed, not recognizing them from last week.
The wire-id → real-ResourceId table stays in the process’s memory and is
never serialized. It is how a response’s resource_id regains meaning
locally.
Checking this yourself
glomeris llm-plan --print-payload [--json]
Runs discovery, prints the exact request a live run would send plus the local alias table, and stops — before any provider is constructed. No API key is needed, and no network call is made, so the command cannot send the thing it is showing you. The two prompt sections are what leaves; the alias section is explicitly labelled as not sent, and is the only place the real absolute paths appear.
What it is not, and never will be
plan_with_llm’s output is a ranking suggestion only. It never calls
policy::classify/authorize, and it must not: every surviving
(ResourceId, ActionId, priority) tuple still has to go through the exact
same policy classification any other candidate would, in the caller, before
anything executes. A dedicated test
(llm_plan_item_never_bypasses_policy) demonstrates this concretely:
constructing evidence for SSH key material, getting an LLM response that
“approves” deleting it, confirming the item survives plan_with_llm’s
validation — and then showing policy::classify still returns Protected
for it regardless.
glomeris llm-plan (the CLI layer built on top of this module) keeps the
same guarantee: crate::cli::build_llm_plan_report never constructs a
policy::Approval and never calls policy::approval::authorize or
executor::execute. Nothing this subcommand prints is ever executed —
there is no --execute/--yes flag, and there never will be one on this
subcommand.
glomeris llm-plan usage
glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json]
glomeris llm-plan --print-payload [--project-root <path>]... [--json]
glomeris llm-plan --schema
- Without
--plan-file, calls a real OpenAI-compatible endpoint viaactions::llm::provider_from_env. --plan-file <path>reads the file’s raw bytes and treats them exactly as if they were the model’s raw response text, through the identicalextract_plan/LlmPlan/plan_with_llmvalidation pipeline — useful for reproducing a scenario deterministically, with no network call and no API key.--print-payloadprints the outbound request without sending it, and returns before a provider is constructed — see “What leaves your machine” above. Requires no API key.--jsonprints theLlmPlanReport(or, with--print-payload, theLlmPayloadReport) as JSON instead of the human-readable form.
Human-readable output always opens with:
LLM SUGGESTION — advisory only, nothing is executed by this command
See CLI Reference for the full flag/exit-code table.
glomeris llm-check — does the configuration actually work?
glomeris llm-check [--json]
The narrowest command in the CLI: no project roots, no detectors, no evidence,
no policy, no actions. It sends two fixed prompts — "You are a connection test. Reply with the single word: ok." and "ok" — through the same
LlmProvider::complete a plan uses, and reports whether the endpoint, the
credential and the model name work.
Two properties are worth being explicit about, because a connection test that lacked either would be worse than none:
- It proves the real path. A cheaper probe — a
GET /models, a HEAD request, a DNS lookup — can pass while aPOST /chat/completionswith your model name fails. This sends the request a plan sends, minus the evidence. - It describes nothing about this machine. The two prompts are constants
with no interpolation, and
connection_test_prompts_describe_nothing_localpins that they contain no path separator, no home directory and no account name — an assertion that is only meaningful because they are constants. So the cost of testing a configuration is one trivial completion, not a disclosure.
The outcome is one of five tokens, and they are deliberately distinct because the fix for each is in a different place:
outcome | Exit | What it means | What to change |
|---|---|---|---|
ok | 0 | The provider answered; its reply is in response_excerpt | — |
rejected | 1 | A non-2xx status — the provider answered and said no | The credential, or the base-URL path: read the path in error |
unreachable | 1 | No HTTP response at all — DNS, TLS, refused, timeout | Network, host name, VPN |
unusable_response | 1 | A 2xx that was empty, not JSON, or missing choices[0].message.content | The model name, or a gateway not really speaking the OpenAI shape |
misconfigured | 2 | validate_base_url refused before anything was sent | The base URL — see above |
The 1/2 split carries a meaning the outcome token alone does not: on 1 a
full report was printed first, so the diagnosis is on stdout; on 2 there
is no report at all, and the single line on stderr prefixed glomeris llm-check: is the whole explanation. misconfigured is therefore never a
token you will read out of a report — InvalidConfiguration can only come from
constructing the provider, which happens before a report exists. It is in the
vocabulary because it is the right name for the condition, and a caller
branching on exit 2 applies it itself rather than inventing a sixth word. That is what lets the menu-bar app’s
connection test branch without parsing prose, and it is also why a missing
configuration is 2 rather than 1 — nothing was attempted, and the fix is
entirely local.
actions::llm::llm_check_outcome is the one producer of those five tokens, and
scripts/check-vocabulary-covers-cli-tokens.sh diffs the set of tokens the
menu-bar app has wording for against the set that function can return, in CI —
so the CLI and the GUI cannot drift into describing different sets of failures,
and adding a sixth failure mode in Rust fails a check rather than silently
reaching a user as a blank explanation.
Configuring this from the menu-bar app
Settings → AI Provider (HORO-1309) configures the same three values from
the GUI, for the common case of someone who runs the app from Finder and never
exported anything. A Finder-launched LSUIElement agent inherits no shell
environment at all, so before this screen existed the GUI’s AI Plan card could
only work for someone who had launched the app from a configured shell.
| Field | Stored in |
|---|---|
| Endpoint (API root) | UserDefaults.standard — for a bundled app, the preference domain named by its bundle identifier |
| Model | the same |
| API key | the login keychain, account llmApiKey, service = the app’s bundle identifier |
Both namespaces are derived from the running bundle rather than written out as literals (HORO-1456), so a beta, renamed or diagnostic build gets its own stored settings and its own keychain items instead of the released app’s.
The precedence rule
Per field, what this app has configured wins over what it inherited. A
field left empty here falls back to the inherited GLOMERIS_LLM_* variable,
including falling back to it being absent. Nothing is ever removed — clearing a
GUI field returns you to the environment rather than switching a working setup
off.
Stated as a rule: the settings screen never lies. If the endpoint field shows a URL, that is the URL used. The alternative — environment-wins — was rejected precisely because it lets the screen display one endpoint while silently using another, with no way to correct it from the GUI.
Each field’s row says where its effective value came from (Configured here,
From the environment, or Not set), so an inherited value is visible rather
than mysterious. An empty-or-whitespace value counts as absent on both sides,
by the same rule the Rust provider_from_parts uses, so an exported-but-empty
GLOMERIS_LLM_MODEL= does not masquerade as configuration.
The CLI is unaffected either way (AC 3). glomeris in a terminal reads that
terminal’s environment and knows nothing about this store. Nothing in this
screen changes, shadows or requires anything for a shell user.
Where the key goes, and where it does not
Into the environment of the glomeris child process, and nowhere else. It is
read from the keychain at the moment a command is spawned and not cached.
It is never a command-line argument (visible to ps, and refused by the CLI by
flag name), never UserDefaults, never a file the app writes, never a log line,
and never rendered — the field is a SecureField, there is no reveal control,
and the typed text is cleared the moment it is handed to the keychain, whether
the write succeeded or not. GlomerisLlmSettingsStore deliberately has no
getter for the key; the only code that reads it is the one function that puts
it straight into a child environment, so “show the user where their key came
from” and “put the key on screen” cannot become the same operation.
scripts/check-credential-store-uses-keychain.sh asserts that in CI over every
line in the app that touches the key.
Removing it is explicit (AC 8): a Remove key button, worded as the
destructive action it is, which deletes the keychain item. Afterwards the field
falls back to the environment if one is exported, and reads Not set if not.
Note what that fallback means, because the precedence rule cuts the other way
here. If GLOMERIS_LLM_API_KEY was exported into the environment this app was
launched from, deleting the stored key does not stop Glomeris reaching a
provider — it promotes the inherited key into use. Someone pressing a button
labelled Remove key is more likely revoking access than tidying a field, so
the app says so in that case rather than reporting a bare “Key removed”: it
names the variable and tells you to unset it too. Unsetting it means relaunching
the app from an environment without it — a running process’s environment is not
editable from the settings screen.
What protects the stored key, and which programs can read it
Honest version, because an earlier one of these notes overstated it (HORO-1455).
The key is a generic-password item in your login keychain, which is a file
keychain — ~/Library/Keychains/login.keychain-db. That means:
- It is encrypted at rest, while the keychain is locked. Normally your login keychain is unlocked for your whole session, so for practical purposes the protection during that session is the access control list below, not encryption.
- It is not synced to iCloud. True — a file keychain cannot sync at all.
- It is included in a file-level backup of your home directory: Time
Machine, a cloned disk, an rsync of
~. The item is still encrypted there and is useless without your keychain password, but it is not excluded from backups. The app sets theThisDeviceOnlyaccessibility attribute, and that attribute is inert on a file keychain — it is honoured by macOS’s data-protection keychain, which needs an entitlement an unsigned build does not have. It is set so the value is already correct if Glomeris is ever signed and entitled.
Which programs can read it. The item carries an access control list naming the binaries allowed to decrypt it. Glomeris creates the item trusting exactly one program: the copy of the app that created it. Any other program asking for the value — including a different build of Glomeris, such as one you just upgraded to — does not silently get it. macOS asks you, naming the program.
You can see and edit that list yourself: open Keychain Access, find the item whose name matches the app’s bundle identifier, and look at Access Control. Nothing about this is only inspectable from inside Glomeris.
If you press “Always Allow” on that prompt, the program you allowed is added to the list permanently. Before HORO-1455 the list only ever grew: every build you upgraded past stayed on it, so over time the key became readable by every old Glomeris binary still on the disk.
How to reset the list. Save the key again — paste it into the API key field and save. Glomeris now removes the old item and creates a fresh one rather than updating in place, so the trusted list is rebuilt from one entry and every previously allowed program is dropped. (This is also why the list cannot be narrowed without re-saving: macOS does not permit an item’s access control list to be rewritten in place.) If you would rather start from nothing, Remove key deletes the item outright — read the note above first about what that promotes.
Two consequences worth stating plainly:
- An upgrade will still prompt you once. Resetting the list is not the same as avoiding the prompt; a new build is not on the old item’s list, and macOS asks before letting it read or delete the item. What changed is that answering once no longer leaves a permanent grant behind.
- If you decline that prompt, the save fails and your previously stored key is left exactly as it was. macOS refuses the deletion and the item stays put, so nothing is lost — the app reports that the save did not happen rather than quietly falling back to an in-place update, because that fallback is what let the list grow. Try again and allow it.
- If the prompt is allowed but the re-add then fails, there is no stored key and the field falls back to the environment, as if you had removed it. That is the one regression this change accepts: previously a failed write left the old key in place. It is the safer direction — the cost is pasting the key again, against a key that stayed readable by binaries you had stopped trusting.
The privacy preview
The same screen offers a preview of the outbound payload, over glomeris llm-plan --print-payload --json with the app’s configured project roots — the
same scope a real plan would use, so the preview is of your payload and not a
generic example. It separates what leaves this Mac (the two prompts) from
what stays on this Mac (the wire-alias table, which is where the absolute
paths are).
Opening it sends nothing, and cannot: --print-payload returns before a
provider is constructed (AC 7). The preview runs without the credential in its
environment at all — only the connection test is given it — so there is no
version of this screen in which looking at the payload transmits it.
LlmPlan schema — the concrete --plan-file example (HORO-1048)
glomeris llm-plan --schema prints an example, syntactically valid
LlmPlan JSON document to stdout — the canonical way to get a starting
point for a --plan-file fixture without reading src/actions/llm.rs’s
LlmPlanItem struct directly. Its resource_id is in the
ResourceId::to_string() form because a human writing a fixture by hand
names resources the way glomeris detect prints them; a live model is
handed wire aliases instead, and answers with those:
glomeris llm-plan --schema
{
"items": [
{
"resource_id": "cargo_target_dir:/path/to/project/target",
"action_id": "cargo.clean.target_dir",
"priority": 1,
"reason": "stale build artifacts, not modified in 30 days"
}
]
}
Field meanings, all on LlmPlanItem:
| Field | Type | Meaning |
|---|---|---|
resource_id | String, required | Must match either a positional wire alias (resource_1, …) from the request the model was given, or — for a hand-written --plan-file fixture — a ResourceId::to_string() from the evidence set plan_with_llm was called with. Both are looked up in sets this process built from its own discovery, so neither can name a resource that was not found locally. An unmatched value drops just that item (dropped_unknown_resource), never the whole plan. |
action_id | String, required | Must match a registered ActionRegistry action id — an unmatched value drops just that item (dropped_unknown_action). |
priority | Option<u32> | Informational ranking hint only. |
reason | Option<String> | Human-readable explanation. Never interpreted as an instruction, a path, or anything that reaches execution. |
#[serde(deny_unknown_fields)] on LlmPlanItem means any extra field
(e.g. a smuggled "command") fails deserialization of the whole
document, not just that item — see “Safety properties” below.
This example is not a fixed, hand-maintained fixture: --schema’s output
is proven, by a real subprocess round-trip test
(tests/llm_plan_schema_round_trip.rs), to be accepted unchanged by
glomeris llm-plan --plan-file <path> — i.e. it parses and validates
through the exact pipeline above without a parse-error exit code. (The
resource_id in the shipped example is deliberately one no real
discovery run will ever produce, so a round trip against a real evidence
set still drops it as an unknown resource — that is an expected,
non-error validation outcome, not a parse failure.)
Configuration
Live mode (no --plan-file) reads three environment variables, all
required, with no default base_url:
| Variable | Purpose |
|---|---|
GLOMERIS_LLM_API_KEY | Bearer token sent as Authorization: Bearer <key> |
GLOMERIS_LLM_BASE_URL | API root of an OpenAI-compatible service — see below |
GLOMERIS_LLM_MODEL | Model name sent in the request body |
GLOMERIS_LLM_BASE_URL is the API root, not the host root
Glomeris appends /chat/completions to the base URL verbatim. It never
inserts a /v1 segment for you, and never strips one. A single trailing
slash is tolerated. So the base URL must be the path prefix your provider
serves its API under — for most providers that includes /v1:
export GLOMERIS_LLM_BASE_URL="https://gateway.example.com/v1"
# request goes to https://gateway.example.com/v1/chat/completions
Setting the host root instead is the most common misconfiguration:
export GLOMERIS_LLM_BASE_URL="https://gateway.example.com"
# request goes to https://gateway.example.com/chat/completions ← wrong path
Because Glomeris does not rewrite the path, a base URL that already ends in
/v1 produces exactly one /v1 segment — there is no /v1/v1 failure mode.
What is refused locally, before anything is sent
actions::llm::validate_base_url runs when the provider is constructed, so a
base URL that cannot possibly work fails with LlmError::InvalidConfiguration
and exit 2 — no request, no round trip to blame. The refusals, each with the
message naming what to change:
| Refused | Why |
|---|---|
| Empty | Nothing to route to; there is deliberately no default base URL |
| Leading or trailing whitespace | A pasted URL routinely carries a trailing newline, and the resulting failure is a 404 nobody can explain by looking at the field |
| An interior space | A URL cannot contain one |
No scheme, or a scheme other than https:///http:// | Nothing else can be posted to |
No host (https://, https:///v1) | Same |
| A query string or a fragment | Appending /chat/completions to …/v1?key=… puts the path inside the query string, and the diagnostic path would then read /v1 and misdescribe why. Refusing also keeps a credential a user pasted into the URL out of the request Glomeris builds |
The full endpoint URL (ends in /chat/completions) | Would produce /v1/chat/completions/chat/completions — the other half of the API-root mistake above |
The message never quotes the offending value, on purpose: of the three settings, the one a user most plausibly pastes into the wrong field is the credential.
What is not refused is the host-root form above. It is a perfectly valid URL
that some providers really do serve their API at, so only the provider can say
whether it works — which is what glomeris llm-check is for.
Do not try to identify this mistake from the status code: a gateway may
answer an unrouted path with 404, but 403, 401 and even 400 are all
things real gateways return instead. The path in the error message is the
path that was actually requested, so it tells you directly which of the two
forms you configured — and since a path of exactly /chat/completions can
only have come from a base URL with no path at all, Glomeris says so itself
rather than leaving you to notice:
provider returned HTTP 403 for openai:chat_completions POST /chat/completions:
{"error":"..."} — the configured address has no path, so this request went to
the host root; most OpenAI-compatible providers serve their API under /v1, so a
missing /v1 is the likeliest cause
(One line in reality; wrapped here to fit.)
The sentence comes after the provider’s own words, never instead of them, and
it is a hint rather than a verdict: the host-root form is valid and some
providers really do serve their API there, so this cannot be a refusal. A
rejection at a configured API root — …/v1/chat/completions — gets no such
sentence at all, because a 401 there is about the key.
A full configuration, using placeholders throughout:
export GLOMERIS_LLM_API_KEY="<token>" # never commit; never pass in argv
export GLOMERIS_LLM_BASE_URL="https://gateway.example.com/v1"
export GLOMERIS_LLM_MODEL="example-model"
glomeris llm-plan --project-root ~/code/my-project
Prefer export in your shell (or a secret manager that exports into the
process environment) over a plaintext .env file, and never pass the key on
the command line — glomeris llm-plan rejects --api-key/--key/--token
outright for that reason.
Diagnosing a provider failure
Provider failures are reported in the provider_error field (and on the
text output’s provider error: line) as one readable, secret-free sentence.
The classes are deliberately distinct, because each has its fix in a different
place — glomeris llm-check reports the same distinction as its five-token
outcome:
| Class | Meaning |
|---|---|
LLM provider configuration is invalid: … | Refused locally; no request was attempted at all |
provider unreachable: … | No HTTP response at all — DNS, TLS, refused connection, timeout |
provider returned HTTP <status> … | The provider answered with a non-2xx status |
provider response was unusable: … | A 2xx response that was empty, not JSON, or missing choices[0].message.content |
An HTTP failure names the status, the API style, the request path, the
provider’s x-request-id when it sends one, and a bounded excerpt of the
provider’s own error body:
provider returned HTTP 401 for openai:chat_completions POST /v1/chat/completions \
(request_id=req-abc-123): {"error":{"code":"invalid_api_key"}}
What is deliberately not in that message: the Authorization header,
the API key, the request payload, the scheme, the host, and the query
string. The host is omitted because it may be private infrastructure and
the query string because some gateways accept a credential there; the path
alone is what diagnoses a base-URL mistake. The body excerpt is truncated
to a bounded length and is scrubbed of the configured API key, so a
provider that echoes the credential it just rejected cannot turn Glomeris’s
diagnostics into the leak.
When that path is exactly /chat/completions, the message also names the
host-root cause described above — the one reading of the path that needs no
knowledge of the host to make.
actions::llm::provider_from_env is the only place std::env::var is
called for these — the key is held just long enough to build the
OpenAiCompatibleProvider and is never logged, printed, or returned any
other way. A missing or empty value for any of the three returns
LlmError::NotConfigured; glomeris llm-plan and glomeris llm-check report
this by naming the three variable names, never a value, and exit 2.
A value that is present but cannot work returns
LlmError::InvalidConfiguration instead, and also exits 2 (HORO-1309).
Keeping the two apart matters: “you have not set this up” and “you set it up
wrongly, here is what to change” have different fixes, and reporting the second
as the first sent users to re-export three variables that were already set.
OpenAiCompatibleProvider is the one real provider implementation — any
OpenAI-compatible /chat/completions endpoint (OpenAI itself, OpenRouter, a
self-hosted gateway):
#![allow(unused)]
fn main() {
pub struct OpenAiCompatibleProvider {
pub base_url: String,
pub api_key: String,
pub model: String,
}
}
The API key is never accepted as a CLI argument. glomeris llm-plan
explicitly rejects --api-key/--key/--token with an error pointing at
$GLOMERIS_LLM_API_KEY instead — a key on the command line would be
visible to ps and land in shell history.
Never commit an API key to source control, a Jira ticket, a commit message, or a PR description. Treat it as you would any other credential.
Safety properties that are already verified
OpenAiCompatibleProviderdeliberately does not deriveDebug— a derivedDebugwould printapi_keyverbatim. A test confirms a real failedcomplete()call’s error output never contains the key.LlmError’s variants never embed the key either — every variant’sDisplayoutput is asserted key-free byapi_key_never_appears_in_display_output_of_any_variant, including theProviderStatusbody excerpt, which actively scrubs the configured key (redact_secret_removes_every_occurrence_and_collapses_whitespace) so a provider echoing the credential back cannot leak it through Glomeris.- The diagnostic endpoint path in a provider error is the path only: two
tests (
diagnostic_endpoint_path_drops_scheme_host_and_query,diagnostic_endpoint_path_never_contains_the_host) assert the scheme, host, and query string are all dropped. - The wire protocol itself is pinned against a loopback mock HTTP server
rather than mocked at the trait: the appended path, trailing-slash
tolerance,
Authorization: Bearerscheme, model and both message roles, and the 401/403/404/empty-body/non-JSON/connection-closed failure modes each have a test that inspects the bytes actually sent. - No absolute path, home directory or account name is transmitted. Asserted
at three levels: on the payload struct
(
payload_carries_no_path_or_account_name), on the CLI report (build_llm_payload_report_separates_outbound_prompts_from_local_aliases), and on the bytes of the real HTTP request, captured by a loopback listener standing in for the provider (tests/llm_plan_egress_privacy.rs) — headers included. That last test also pins the two halves that make the proof non-vacuous: the payload demonstrably does describe a path-backed resource, and the real path was available locally and was withheld deliberately. - Wire aliases are positional, so the same request shape carries
byte-identical ids regardless of what the resources are
(
wire_ids_are_positional_and_independent_of_the_resource) — there is no stable identifier in the payload that could correlate a machine across requests. - An alias resolves only through the request’s own table, and a real-id
string only through the evidence set: a model that invents either — an
out-of-range
resource_99, or a guessed path it was never shown — names nothing (out_of_range_wire_alias_is_dropped_and_counted,model_supplied_path_cannot_name_an_undiscovered_resource). No model-provided identifier gains execution authority; policy still classifies every surviving item. - The serialized view’s key set is pinned to its seven documented fields
(
serialized_view_exposes_exactly_the_documented_fields), so adding a field toLlmResourceView— the only way to widen what leaves — fails a test rather than passing silently. LlmPlanItemuses#[serde(deny_unknown_fields)]: a model trying to smuggle an extra field (e.g. a"command") fails deserialization of the whole plan, not just that item — confirmed byunexpected_field_rejects_whole_plan.- An unknown
resource_idoraction_idin the model’s response drops just that one item (counted indropped_unknown_resource/dropped_unknown_action); it never fails or invalidates the rest of the plan. LlmPlan/LlmPlanItemremain the only two#[derive(Deserialize)]types in the entire crate —LlmPlanItemReport/LlmPlanReport(the CLI report DTOs) areSerializeonly. Everything a model response can produce is aStringresolved against real, already-in-memory data, or a plainu32/Option<String>used only for display/ranking — never a path, never a shell fragment, never anything that reachesActionStepconstruction directly.- The response parser handles a raw JSON object, a fenced
```jsonblock, or a bare fenced block, and either produces a well-formedLlmPlanor nothing — never a partial parse. - A
PROTECTEDcandidate is refused before its action is ever resolved:crate::cli::build_llm_plan_reportchecksPolicyClass::Protectedbefore callingActionRegistry::get/executor::dry_runfor that item.tests/golden_llm_plan_protected_refusal.rsproves this end to end through the exact--plan-fileinput surface the CLI uses, closing the gap Known Limitations previously described as “verified at the code level only, not end-to-end through the CLI.” - The connection test’s two prompts are constants, and
connection_test_prompts_describe_nothing_localasserts they contain no path separator, no home directory and no account name — so testing a configuration costs one trivial completion and discloses nothing. - The GUI’s provider settings are held to the same line by
macos/GlomerisMenuBar/Tests/AiProviderPreferencesViewTests.swift, which asserts several things as absences, since they are code paths that must not exist: the key is bound to aSecureFieldand never read back or revealed; the typed text is cleared before the success-or-failure branch, so a failed keychain write cannot leave it in memory behind a visible error; neither command can run without the user pressing something (no.onAppear, no.task, noTimer, no.refreshable); exactly one call site is given the credential, and it is not the payload preview; and the preview’s outcome enum has no “not configured” case at all — a preview that could say that would be a preview that had tried to use a provider. Bothllm-checkreports are additionally pinned by golden fixtures on both sides (tests/dto_golden_fixtures.rs,DtoGoldenFixturesTests.swift), asserting nobase_url, noapi_keyand no://reaches a report a UI shows and a log keeps. - The GUI half of AI Plan is held to the same line by
macos/GlomerisMenuBar/Tests/AiPlanSectionViewTests.swift: a row built from a confident, plausible recommendation to delete a credential carries the refusal and no willingness to act, and the card has no execution call site, no fingerprint token and no re-sorting of the model’s order. The repo-widescripts/check-no-policy-label-branching.shadditionally proves nothing in the app — this card included — branches on apolicy_labelstring to decide what may run.
Known limitations
- Only one provider shape is implemented — no Anthropic-native, Azure-OpenAI-specific auth, or other provider shape.
- No retry/backoff on transient network failures.
- No streaming support —
complete()returns the full response text at once. - No integration into any existing rule-only ranking loop (
glomeris free) — deliberately out of scope for this ticket. A future ticket may wire it in as an optional enhancement, falling back to rule-only ranking whenever the provider call fails. glomeris llm-planhas no--execute/--yesflag and never will on this subcommand — it is advisory-only by design, not an MVP gap.
Autopilot
Autopilot is a standing grant with limits, written down where you can read it.
Every other deleting command in Glomeris is something you type at the moment
you want it: execute names one action on one resource, free runs a bounded
recovery loop you asked for, emergency is the button you press when the disk
is nearly full. Autopilot is the one that can act without you present — so the
whole of its design is about bounding what “without you present” is allowed to
mean.
The invariant the rest of this book states as
AI can recommend. Policy decides. Executor verifies. Filesystem reality wins.
gains one clause here:
…and the envelope bounds the outcome.
The envelope
An AutopilotEnvelope is the complete statement of what Autopilot may do. It
is a small file of key = value lines at
~/Library/Application Support/Glomeris/autopilot.conf, and nothing else
grants authority — not an environment variable, not the plan you hand run,
and not the menu-bar app’s Autopilot tab, which grants by running
autopilot enable and keeps no second copy of the answer. There is one grant
and it is that file.
| Field | Default | Hard ceiling | What it bounds |
|---|---|---|---|
| enabled | revoked | — | Whether run may execute anything at all. |
| allowed kinds | none | — | Which ResourceKinds a candidate may be. An empty allowlist reaches nothing. |
| max actions | 3 | 25 | How many actions one run may perform. |
| max bytes | 5 GiB | 64 GiB | The byte total one run may reclaim. |
| max duration | 60s | 900s | Wall-clock budget for one run. |
| min pressure | none | — | A disk-pressure floor below which the run refuses. |
| pre-authorized ASK | none | — | Exactly which (kind, reason) pairs an ASK classification may proceed on. |
| respond to alerts | off | — | Whether a run may begin without being asked. The one row that is not a ceiling on outcomes — see Runs nobody asked for. |
Two properties of that table matter more than the numbers in it:
The defaults grant nothing. AutopilotEnvelope::default() and
::revoked() are the same value, and a missing file reads as revoked rather
than as an error. Autopilot being off is not a setting someone has to
remember to choose; it is what the absence of a decision means.
A limit above its ceiling is refused, not clamped. Asking for 500 actions is a usage error (exit 2) and writes no envelope. Clamping would be the worse behaviour by a wide margin: you would have been told 500, be living under 25, and have no way to notice the difference.
Reading the envelope is read-only in the strict sense — glomeris autopilot
with no arguments prints the grant and does not create the file it just read.
What the envelope cannot authorize
PROTECTED is unconditional. No flag on autopilot enable can reach it, and
no ordering in a plan file can make a PROTECTED candidate executable. The
allowlist is a narrowing filter over what policy already permits, never a
widening one.
UNKNOWN_INCOMPLETE is likewise unreachable. It is the reporting label for an
ASK decision whose evidence is incomplete, stale, or came from a probe that
failed — so --preauthorize-ask refuses every evidence-quality reason.
Pre-authorizing “I accept the rebuild cost” is a judgment a person can make in
advance. “I accept that we do not know what this is” is not.
--preauthorize-ask takes one kind:reason pair at a time, e.g.
node_modules:rebuild_cost_high. There is no wildcard, because the only
purpose of a wildcard here would be to turn a narrow consent into a blanket
one.
What a model may do, exactly
autopilot run --plan-file <path> reads an LlmPlan — the same schema
BYOK LLM Planner documents, and produced by the same llm-plan
command — and uses it to order the candidates this machine already
discovered. That is its entire authority.
It cannot:
- add a candidate (a resource ID the local scan did not produce is dropped);
- choose an action (
ValidatedPlanItem::action_idis deliberately ignored by the run; the action comes from the local registry, by kind); - supply a path, a command, or a command fragment;
- raise any limit in the envelope;
- change any policy class;
- invent an action that is not registered.
The argument that this is airtight is structural rather than defensive:
ordering is a permutation of a list, and reordering a list cannot add a member
to it. The plan is consumed as a sort key over locally-derived candidates, so
the set of things that can be deleted is fixed before the plan is read. A
regression test hands the run a plan whose reason field is
"; rm -rf /Users/dev/company-repo && echo pwned" and asserts, among other
things, that the string never reaches the audit log — but the reason that test
passes is that model_reason is never read by anything but the record writer.
There is deliberately no live-provider mode. autopilot run accepts a
plan file and nothing else, so no deletion in this product waits on a network
call. Asking a model is llm-plan’s job, and the two are separate commands so
that the deleting one has no network path at all.
What one run actually does
For each candidate, in the order the plan (or the local ranking) put them:
- Admission gate — a pure function over the envelope, the candidate’s
policy decision, the observed pressure and the run’s budget ledger. It
returns either admission or a
RefusalReasonthat names which bound stopped it: kind not allowed, policy class, unauthorizedASKreason, pressure floor, action budget, byte budget, time budget. - Authorization — the normal
policy::approval::authorizepath, which is the only way anApprovalcan exist. A pre-authorizedASKsupplies a realUserConsentmatched to the resource and fingerprint; nothing forges one. - Execution — the normal
executor::execute, including deletion-time TOCTOU revalidation. An identity change, a class downgrade, widened reasons or degraded evidence aborts the action having deleted nothing. See Safety Model. - Ledger — the attempt is charged exactly once, at
max(expected, actual)bytes, so an action that reclaimed more than predicted cannot under-charge the budget.
The action budget stops the run; a byte-budget refusal does not. A candidate too large for the remaining byte budget is refused and the run continues to the next one, because the next one may well fit — whereas an exhausted action count is exhausted for everything.
A run that finds nothing it is allowed to do exits 0. “Allowed nothing” is a successful outcome, not an error.
The other consumer of the grant: free --autopilot
The envelope is not autopilot run’s private property. glomeris free --autopilot bounds the goal-driven recovery loop by the same stored grant, using
the same admission gate above as one more narrowing stage between selection and
authorization — so a goal-driven run can never reach further than a grant allows,
and the grant remains the single place that authority is written down.
Two consequences worth stating, because they are what makes the second consumer safe rather than merely convenient:
- The grant cannot widen anything. It has no power to add a candidate, choose
an action, supply a path or raise a limit — only to withhold. Everything an
--autopilotrun does, the same command without the flag would also have done. - A stop the grant caused is reported as such, never as an exhausted disk:
stop_reason: "envelope_refused"plus the refusal token.revoketherefore stops afree --autopilotrun exactly as it stopsautopilot run— both re-read the file, and neither has a cached copy.
See CLI Reference for the flag’s exit codes and its refusal
alongside --dry-run.
Runs nobody asked for
Everything above bounds what a run may do. Whether a run may begin when
nobody pressed anything — the standing answer to a disk-pressure alert — is a
separate question, and the envelope answers it separately:
autopilot enable --respond-to-alerts, off by default, exercised by
free --autopilot --unattended.
It is two consents, not one. An enabled envelope does not authorize an
unprompted run; --unattended without that permission exits 3 having attempted
nothing. Reading “enabled” as consent to act unattended would widen every
envelope already on disk — each written by somebody bounding a run they meant to
start — into authority to delete while they were away. That is not a grant those
people gave, so it is not one this reads out of their file.
It is the only field here that is not a ceiling on outcomes, which is why the table above says so rather than glossing it. It is still narrowing in the direction that matters: off, the default, means every run has a human behind it, and turning it on changes nothing about what a run may do once it starts. The budgets, the allowlist and the pre-authorizations all still apply, unchanged.
The check lives in Rust, twice over. The grant exposes
starts_unprompted — in force and set — as one value, so that no client ever
composes its own permission out of two fields; and free itself refuses an
--unattended run the grant does not cover, so the check cannot be skipped by
whatever launched the process. --unattended is the caller’s statement about
itself, and demanding it is what makes a run nobody started distinguishable from
one somebody typed.
revoke stops it with everything else, and keeps the setting on file — so
revoking never reads as having silently cleared what you chose.
The audit trail
Every executed attempt appends one JSON line to
~/Library/Application Support/Glomeris/actions.jsonl, readable with
glomeris history. Autopilot’s records carry two fields the interactive paths
do not populate:
source—autopilot_auto_safeorautopilot_preauthorized_ask, so the authority an action ran under is recoverable from the log rather than inferred. Interactive records readexecute,freeoremergency.model_rank— where the plan ranked that candidate, ornullwhen no plan was involved. This is how you tell “the model suggested it and policy allowed it” from “policy allowed it and no model was consulted.” It is 1-based, and it is the same number the run’s own report printed as[AI rank N]— a log offset by one from the report it came from would be worse than no log.
“Where the plan ranked it” means its position in the plan’s items array, not
its priority field. priority is advisory and its direction was never
specified anywhere — nothing says whether 1 means most urgent or least — so
ordering on it would mean inventing a semantics and then depending on it.
Refusals are not written to actions.jsonl — it is a log of what was done
to the filesystem, and a refusal did nothing to the filesystem. They are in the
AutopilotReport that run prints, which is where you look to find out why a
run was quiet.
Revocation
glomeris autopilot revoke takes effect on the next run, and there is nothing
to restart: every run re-reads the file. Revocation keeps the limits, so a
later enable cannot come back carrying limits you never read.
enable replaces the previous envelope rather than merging into it, for the
same reason. A grant should be one line you can read out loud, never the
accumulated union of every enable you have ever typed.
Nothing here expires on its own. The envelope is a file: it survives quitting,
restarting and logging out, and it stays in force until it is revoked. That is
a deliberate omission rather than a missing feature — a grant that lapsed on a
timer would mean the honest answer to “what may Autopilot do right now” depended
on the clock, and show would have to be read differently depending on when you
read it.
Granting it without a terminal
Everything above can also be read, granted and revoked in the menu-bar app,
under Settings → Autopilot (Cmd+,). Reading a standing deletion grant only
from --help and a config file would have meant that in practice most people
who enabled it had never read it, so the GUI is part of the feature rather than
a convenience on top of it.
It changes nothing about where authority lives. The tab renders
autopilot show --json — the same envelope, the same ceilings, the same lists
of what can never be allowlisted, pre-authorized or executed — and writes by
running autopilot enable or autopilot revoke. It holds no policy logic, and
it offers no choice that did not arrive from the CLI as data, which is the
project-wide rule for that app and is checked in CI rather than trusted.
Two things it deliberately does not do: it cannot start a run (autopilot run
has no --json for the same reason), and it cannot construct a grant the CLI
would reject. See
Menu Bar App for the screen itself,
including why its form is pre-filled from the envelope in force and why its
byte budget rounds up.
Known limitation: a pre-authorized ASK cannot complete today
A pre-authorized ASK candidate is admitted, authorized with real consent,
and then always aborts inside deletion-time revalidation with
PolicyClassDowngraded. Nothing is deleted.
The cause is upstream of Autopilot. ASK/RebuildCostHigh arises only from a
per-instance Regenerability::NotRegenerable, while
executor::build_fresh_evidence rebuilds regenerability from the resource
kind’s static default — so the fresh classification lands on AutoSafe and
the class comparison trips.
This is left as-is deliberately. The failure direction is the safe one (refuse,
mutate nothing), and “make build_fresh_evidence carry a per-instance
judgment” changes the TOCTOU anchor every deleting command shares. It is
asserted by a test named for it —
a_preauthorized_ask_still_aborts_at_deletion_time_revalidation — rather than
left to be discovered, so whoever fixes it upstream starts from a failing test
that points at the cause.
See Known Limitations for the rest.
Daemon Lifecycle
glomeris daemon <install|uninstall|status|run> manages an optional
background poller via macOS’s launchd, implemented in
src/platform/macos/launchd.rs.
User-level, not root
Everything here runs at the per-user level: the plist is written to
~/Library/LaunchAgents/com.glomeris.monitor.plist, and it is loaded with
plain launchctl load/unload (gui/<uid> session semantics). Nothing
writes to a system-level LaunchDaemons location, and no root/sudo access is
required for any of these commands.
glomeris daemon install [--force]
Writes a plist that runs <program_path> daemon run every 60 seconds
(StartInterval) by default, with RunAtLoad set, and stdout/stderr
redirected to ~/Library/Logs/Glomeris/monitor.log/.err.log. Then runs
launchctl load -w <plist_path>. If launchctl fails, the plist file is
left in place (so status/a manual retry can still see it) — only the
final launchctl result is treated as a hard failure.
What this owns, preserves, and refuses to clobber (HORO-1021): the
plist at ~/Library/LaunchAgents/com.glomeris.monitor.plist is a wholly
Glomeris-owned artifact — no other tool reads or writes it. Even so,
install never silently discards a hand-edited value:
- Idempotent. Re-running
installagainst an already-up-to-date plist (matching a fresh default install) writes nothing. - Preserves a hand-edited
StartInterval. If you’ve changed the poll interval by hand, re-runninginstallkeeps your value — only the program path is refreshed (the actual point of reinstalling after rebuilding/moving the binary). - Refuses any other unrecognized customization without
--force. If the on-disk plist has been changed in a way this tool doesn’t manage (a hand-added key, a flippedRunAtLoad, etc.),installmakes zero changes and exits with status2, explaining the refusal. Pass--forceto overwrite anyway — a backup of the current file is taken first, at<plist_path>.plist.bak. - Atomic, verified writes. Every real write goes to a temp file in the
same directory, is renamed into place, and is read back and compared
before
installreports success — a crash or concurrentinstallmid-write can’t leave a corrupt/partial plist. - Aborts on concurrent modification. If the plist changes on disk
between
install’s read and its write (e.g. twodaemon installs racing), the later write aborts with zero mutation of the live plist and exit status2rather than risk clobbering the concurrent change — rerun to retry. This is checked twice: once before anything is written, and again right before the live plist itself is overwritten (after the backup copy has been taken) — narrowing the window where a concurrent writer could land and be silently clobbered down to the instant between that second check and the write itself. That last sliver can’t be closed without holding a lock across the whole read-check-write section, which is out of scope here. If the second check is the one that catches a race, a.plist.bakbackup of the pre-race content may already be on disk; that stray backup is harmless (never the concurrent writer’s data) and is not treated as a mutation of the live plist for the purposes of this guarantee.
glomeris daemon uninstall
Runs launchctl unload -w <plist_path> (best-effort — an already-unloaded
agent reporting an error from launchctl is not treated as fatal here),
backs up the current plist to <plist_path>.plist.bak, then removes the
plist file if present. Idempotent — uninstalling an already-missing plist
is not an error.
glomeris daemon status
Prints whether the plist file exists, its path, and whether launchctl list <label> currently reports the job loaded. This never errors: an unreachable
launchctl (e.g. a non-macOS or sandboxed CI environment) is reported as
loaded: false, not surfaced as a failure.
glomeris daemon run
Runs the polling loop in the foreground — this is the exact command the
installed launchd agent invokes on its own schedule. It watches /,
using ThresholdConfig::default() (see Pressure Model),
the real macOS statvfs-backed FsStat, and the real osascript-backed
notifier, appending confirmed pressure transitions to
~/Library/Application Support/Glomeris/history.tsv.
Known CI limitation
Only plist generation and path/file logic are unit tested in this
repository’s CI. Actually asking launchd to load/run the agent needs a
real macOS user session and is not exercised by cargo test.
Architecture
Module map
From src/lib.rs, the crate’s pub mod declarations:
#![allow(unused)]
fn main() {
pub mod actions;
pub mod detectors;
pub mod emergency;
pub mod evidence;
pub mod executor;
pub mod monitor;
pub mod platform;
pub mod policy;
pub mod scanner;
}
platform::macos is itself gated #[cfg(target_os = "macos")] — it does
not exist in a build on any other OS. src/main.rs is a thin CLI shim over
this library; the pure logic in monitor/scanner and their unit tests do
not depend on being reachable from a binary at all.
| Module | Responsibility |
|---|---|
monitor | Disk-pressure state machine and the polling loop that drives it. Cheap, O(1) capacity checks only (statvfs) — never walks a directory tree. See Pressure Model. |
scanner | A generic, bounded, streaming filesystem walker that reports the top-K largest entries by size. Unrelated to the detectors below — it has no notion of resource kind, regenerability, or policy. |
detectors | Bounded, per-tool probes (Cargo, Homebrew, Node, Xcode, Docker) of known roots, with a bounded recursive size estimate under each root — not a full filesystem walk. Produces discovery-stage Evidence with the four correlation fields always Unavailable(NotAttempted). |
evidence | The Evidence/ProbeOutcome/Completeness domain model (evidence::model, evidence::probe), plus the correlate submodule that fills in the four correlation fields via lsof/git/pgrep. See Evidence Model. |
policy | The single deterministic classify()/authorize() pair that decides AUTO_SAFE/ASK/PROTECTED. See Safety Model. |
actions | Typed, pre-registered cleanup actions (ActionRegistry) that turn Evidence into a closed ActionPlan/ActionStep — never a caller-supplied string. Also home to the optional BYOK LLM planner (actions::llm). |
executor | Dry-run, real execution with deletion-time TOCTOU revalidation, and the bounded closed-loop recovery orchestration (executor::recovery_loop, glomeris free). |
emergency | The degraded, AUTO_SAFE-only, no-LLM/no-network recovery path (glomeris emergency). |
platform::macos | The only place real OS-level I/O happens: statvfs (disk stats), osascript (notifications), launchd (daemon lifecycle). |
Data flow
Observe (monitor)
-> statvfs, classify pressure, debounce, notify/persist (best-effort)
Discover (scanner | detectors)
-> scanner: generic top-K-by-size walk (no resource semantics)
-> detectors: bounded per-tool probes -> discovery-stage Evidence
(correlation fields still Unavailable(NotAttempted))
Correlate (evidence::correlate)
-> lsof / git / pgrep, per-field timeout and failure handling
-> fills open_by_process, process_cwd_match, git_state, tool_liveness
Decide (policy::classify / policy::approval::authorize)
-> AUTO_SAFE | ASK | PROTECTED, fail-closed ordering
-> Approval is the only thing an executor may act on
Plan (actions::Action::plan)
-> typed ActionPlan/ActionStep from Evidence
-> optional: actions::llm ranks candidates, never authorizes them
Execute & revalidate (executor::execute)
-> re-collect evidence, re-run classify, abort on any disagreement
-> re-verify filesystem identity immediately before mutating
-> re-measure actual reclaimed bytes after the fact
executor::recovery_loop (the orchestration behind glomeris free) and
emergency are both thin callers that wire the above stages together,
one candidate at a time, without reimplementing or loosening any of them —
neither introduces its own execution or authorization primitive.
Why the scanner and the detectors are separate
The scanner (scanner::walker) is a generic, dependency-free (std-only),
non-recursive directory walker with no concept of what it finds — it reports
size and depth only. The detectors are the opposite: narrow, tool-aware
probes of specific known locations (e.g. ~/Library/Developer/Xcode/DerivedData,
brew --cache’s reported path, <project>/target) that produce
policy-relevant Evidence. glomeris scan exercises the former; glomeris detect, glomeris free, and glomeris emergency all exercise the latter.
Troubleshooting
External tools Glomeris shells out to
| Tool | Used by | Purpose | If absent |
|---|---|---|---|
docker | detectors::docker | docker system df --format '{{json .}}' to report build/image cache size | ToolAbsent — normal, expected, not an error |
brew | detectors::homebrew | brew --cache to find Homebrew’s cache path | ToolAbsent — normal, expected |
lsof | evidence::correlate::open_files, ::process | Find processes with a resource open, or with it as their cwd | That correlation field becomes Unavailable(ToolAbsent) for the affected resource; other fields are unaffected |
git | evidence::correlate::git | Determine repo root, dirty/untracked state, worktree-ness | git_state becomes Unavailable(ToolAbsent) for the affected resource |
pgrep | evidence::correlate::tool_liveness | Check whether Xcode.app or Docker’s backend daemon is running | tool_liveness becomes Unavailable(ToolAbsent) for the affected resource |
launchctl | platform::macos::launchd | Load/unload/query the background daemon | daemon status reports loaded: false; install/uninstall report an error but the plist file itself is still written/removed |
osascript | platform::macos::notify | Post a macOS notification on a confirmed pressure transition | Notification failure is captured as notify_error on the poll outcome; the monitor loop keeps running |
A missing tool is treated as a normal, expected state everywhere in this
codebase, not an error — every detector’s own doc comment says so
explicitly (e.g. “not every detector’s tool is installed on every machine”).
The one nuance: for Docker specifically, a running-but-unreachable daemon
(e.g. Docker Desktop not started) is also folded into ToolAbsent, not a
separate Failed state.
Cargo, Node, and Xcode detectors never shell out to cargo/node/npm/
xcodebuild at all — they only check the filesystem (known_project_roots
for a target/node_modules directory, or
~/Library/Developer/Xcode/DerivedData). Their ToolAbsent really means
“expected resource not present” (no project roots configured, or the
directory doesn’t exist), not “binary missing from PATH.”
“Why does glomeris detect show tool_absent for something I have installed?”
For Cargo/Node, tool_absent also appears if no known_project_roots were
configured for the DiscoveryContext used — these two detectors never
search the filesystem on their own; they only check specific roots handed to
them. Check how the caller (CLI/daemon) constructed the DiscoveryContext.
“Why did a probe come back Unavailable(Failed) instead of ToolAbsent?”
ProbeReason::Failed means the tool ran but something about its output or
exit status wasn’t a recognized “nothing found” or “tool absent” shape — for
example lsof exiting non-zero with stderr content, or brew --cache
succeeding but reporting a path this process can’t canonicalize. This is
deliberately never coerced into a safe default; treat it the same as “we
don’t know,” not “nothing to clean up.”
“Part of the scan failed — where do I see which detector?”
A failed detector and one whose tool is absent both contribute zero candidates, so a count alone cannot tell them apart while they mean opposite things: “we don’t know what is there” versus “there is nothing there.” Every surface therefore reports the outcome and not just the count.
| Surface | Where a failure appears |
|---|---|
glomeris detect | With no candidates at all, the line reads no candidates discovered by the detectors that succeeded — N failed, so this is not a clean bill of health rather than the bare no candidates discovered |
glomeris detect --json | A detectors array — one entry per detector, in registration order, with status (found/tool_absent/failed), candidates_found, and a reason present only on failed — plus a derived discovery_complete |
glomeris detect --progress-json | Each detector_finished event carries outcome, and reason when it failed |
glomeris free, glomeris emergency | A discovery incomplete: N detector(s) failed block naming each one; if the run stopped at SafeExhausted, it also says in so many words that this is not a finding that nothing safe is left |
| Menu-bar app | “Nothing found where Glomeris could look” in place of the all-clear, or “This list may be incomplete” below a non-empty list — either way naming the checks that did not finish |
An absent tool appears as tool_absent and does not make
discovery_complete false. That is a real answer, not a missing one, and
flagging it would make the caveat permanent on any machine without Docker —
which is the fastest way to teach people to ignore it.
“The menu-bar app says the glomeris CLI was not found, but I installed it”
The app lists the locations it searched in the same message. If your binary
is not in one of them, that is the mismatch — move or symlink it into
/opt/homebrew/bin or /usr/local/bin, or launch the app from a shell
whose PATH contains its directory. A GUI app launched from Finder gets
only PATH=/usr/bin:/bin:/usr/sbin:/sbin, so a PATH that works in your
terminal is not visible to the app. See
Menu Bar App for the
full resolution order.
You do not need to restart the app after installing the CLI: it re-resolves on every invocation, so the next poll picks it up.
macOS permissions
Glomeris’s own I/O runs as the invoking user; it does not request or use
Full Disk Access, and nothing in this codebase currently prompts for or
checks any macOS privacy permission (TCC). If a probe or detector needs to
read a location gated by such a permission on your system, expect a
PermissionDenied-flavored Failed/Unavailable outcome rather than a
silent empty result — check the affected detector/probe’s specific error
message.
glomeris free/glomeris emergency doesn’t actually delete anything
This is expected today, not a bug you need to work around — see Known Limitations for exactly why, and Safety Model for the revalidation logic responsible.
Non-macOS platforms
daemon, emergency, and free print an error to stderr and exit 1 on
any OS other than macOS, because platform::macos (real statvfs,
osascript, launchd) does not exist in that build at all. scan and
detect are not gated this way and should work cross-platform, though the
project’s CI only exercises macos-14.
Security & Privacy
No telemetry, no cloud requirement
Glomeris has no backend, no account system, and no telemetry. Every command
documented in this book runs entirely locally. The only network access anywhere
in the codebase is to whatever BYOK endpoint you configure yourself, from
exactly two commands — glomeris llm-plan (the planner) and glomeris llm-check (the connection test, which sends two fixed words and nothing about
this machine). See BYOK LLM Planner. Nothing else in the crate makes
a network request.
No raw filesystem inventory, and no filesystem paths, sent to any LLM
As of HORO-1008, glomeris llm-plan is wired into the binary (see
BYOK LLM Planner). No filesystem data is sent anywhere unless you
explicitly run that subcommand without --plan-file and have all three
provider settings configured — every other command in this book either makes no
network request at all, or (glomeris llm-check) sends two fixed words that
describe nothing local. Since HORO-1309 those three settings may come from the
menu-bar app’s Settings rather than from GLOMERIS_LLM_* in a shell, which
changes where they are stored and nothing about what is sent: the GUI’s AI Plan
card spawns the same subcommand with the same bounded payload. Even then, what is sent is bounded and
explicit: LlmResourceView is a hand-maintained projection of one
Evidence record containing a resource’s kind, size estimate, age,
regenerability, completeness, and the action ids offered for it — never raw
file contents, never a directory listing, never anything beyond that fixed
set of fields. Adding a field to Evidence later has no effect on what a
model sees unless a human explicitly adds it to LlmResourceView too, and
adding a field to LlmResourceView itself fails a test that pins its
serialized key set.
Nor are resource paths sent. Until HORO-1298 each view identified its
resource by its real ResourceId, which for the five path-backed resource
kinds renders as an absolute path — and an absolute path under $HOME
discloses the OS account name and the machine’s directory layout. Note that
--project-root never bounded this: it scopes the cargo and node detectors
only, while the Xcode detector is $HOME-bounded and the Homebrew and
Docker detectors shell out and are bounded by neither. Each view now
carries a positional wire alias — resource_1, resource_2, … — and the
table mapping an alias back to a real ResourceId has no Serialize
derive and stays in the process’s memory. The aliases are positional rather
than hashed on purpose: a hashed path would be both brute-forceable (one
unknown segment in /Users/<name>/Library/...) and a stable handle for
correlating your machine across requests.
Run glomeris llm-plan --print-payload to see the exact request a live run
would send, including the local alias table, without sending it and without
configuring a credential. Three tests enforce the property, the strongest
of them (tests/llm_plan_egress_privacy.rs) by capturing the real HTTP
request with a loopback listener and asserting the transmitted bytes
contain no path separator at all.
BYOK secret handling
OpenAiCompatibleProviderdeliberately does not deriveDebug, so an accidental{:?}-log of the provider value cannot leak the API key.- Verified by test: neither
LlmError’sDebugoutput nor a real failedcomplete()call’s error output contains the key. glomeris llm-planandglomeris llm-checkread the key only fromGLOMERIS_LLM_API_KEYviaactions::llm::provider_from_env—std::env::varfor this value is called nowhere else in the crate. The key is never accepted as a CLI flag:--api-key/--key/--tokenare explicitly rejected with an error pointing at the environment variable instead, so a key never appears inpsoutput or shell history. The rejection message names the flag only, never the value beside it, so--api-key=<secret>does not get echoed back. See the BYOK page for the full configuration contract.- In the menu-bar app (HORO-1309) the key lives in the login keychain, not
in
UserDefaults, not in a file the app writes, and not in a log line. It is read at the moment aglomerischild process is spawned, placed in that child’s environment, and not cached.GlomerisLlmSettingsStoredeliberately exposes no getter for it — the only code that can read the value is the one function that builds a child environment — and the settings screen binds it to aSecureFieldwith no reveal control and no read-back. Removing it is an explicit destructive control, not a side effect of clearing a field.scripts/check-credential-store-uses-keychain.shenforces all of that in CI over every line in the app that touches the key, so a future shortcut — stashing it in a plist “just for now”, printing it in a debug view — fails a check rather than shipping.
Subprocess invocation
Every external tool Glomeris shells out to — docker, brew, lsof,
git, pgrep, launchctl, cargo, npm/pnpm/yarn, brew (as an
action) — is invoked via std::process::Command::new(program).args(args)
with a Vec<String>/&[&str] argument array. There is no sh -c anywhere
in this codebase, and no place where untrusted content is interpolated into
a shell command string.
The one place a string is built and passed to an external tool is
platform::macos::notify’s osascript invocation, which constructs an
AppleScript source snippet via format!. That string’s only two
substituted values are the notification title and body produced by
monitor::notifier::notification_text, which only ever emits fixed template
text derived from the closed PressureState enum — never external or
attacker-controlled input — and the whole script is still passed to
osascript as a single argument in an argument array, never through
sh -c.
No arbitrary LLM-composed commands reach execution
ActionPlan/ActionStep deliberately never implement Deserialize, so no
externally-sourced input (a parsed LLM response, a network payload) can ever
construct one directly. The only way external input reaches execution is by
selecting a pre-registered ActionId string, which a real Action::plan
implementation then interprets against Evidence this process already
collected and trusts independently. See Safety Model for
how this interacts with policy classification.
Bounded, printable failure reporting
EmergencyReport caps the number of retained error strings at 8, and no
type in this crate’s execution/reporting path derives Serialize in a way
that would let an unbounded or adversarial input grow a report without
limit.
Known Limitations
This page collects specific, verified limitations pulled from the actual code and from the PRs that introduced each piece — not a generic disclaimer. This is an experimental MVP; treat every claim elsewhere in this book as scoped by what’s on this page.
Resolved: real detector-produced candidates now complete real deletions
Fixed in HORO-994. Detectors populate Evidence::reclaimable_bytes at
discovery time (HORO-992) — Xcode/Cargo/Node/Homebrew via a real size
estimate, Docker via docker system df’s own reported figure — and
executor::execute()’s deletion-time TOCTOU revalidation
(executor::build_fresh_evidence) now reuses the same size-estimate
computation for reclaimable_bytes that it already used for
logical_bytes, rather than hardcoding it back to Unavailable. A golden
end-to-end integration test (tests/golden_chain_execute.rs) proves the
full chain — real detector → AutoSafe classification → Approval →
execute() → real deletion → real re-measured freed bytes — against a
disposable fixture. glomeris free --target and glomeris emergency can
both now actually free real bytes, not just report a dry-run plan.
Resolved (HORO-1016): the size estimate was non-recursive and badly under-counted nested trees
Before this fix, detectors::shallow_logical_bytes (the shared size probe
behind every detector’s logical_bytes/reclaimable_bytes and behind
executor::build_fresh_evidence’s revalidation) summed only a directory’s
immediate entries — for a subdirectory entry it counted the directory
inode’s own size, never its contents. On a real machine this reported a
4.8 GB Xcode DerivedData tree as 34.9 KB, a 1.0 GB Homebrew cache as 1.3 MB,
and a 16 GB Cargo target/ directory as 4.6 KB. Fixed:
detectors::estimate_logical_bytes walks the full subtree via an explicit
stack (never real recursion, so an arbitrarily deep tree cannot overflow
the stack), bounded by a shared 200,000-entry / 750ms budget
(detectors::size_estimate_budget) used identically at discovery time and
at deletion-time revalidation. Budget exhaustion always yields a truthful
partial-sum lower bound — never Unavailable — so a truncated walk can
never look like a probe failure to Evidence::completeness() or trip a
spurious abort in executor::execute’s TOCTOU revalidation; when the walk
does stop early, a provenance note is attached via Evidence::push_source
(advisory only, never a policy input). tests/reclaimable_bytes_reaches_auto_safe.rs
and the new executor::tests::estimate_matches_total_size_best_effort_on_the_same_tree
lock the estimate’s byte semantics to executor::total_size_best_effort’s
existing convention (files and symlinks counted by their own size,
directories contribute 0). On a real developer machine the 750ms deadline
truncates the Xcode DerivedData walk at roughly 42,000 entries, so that
resource’s reported bytes are normally a lower bound, by design — no
extrapolation is attempted.
Docker build cache can never reach Completeness::Complete
DockerBuildCache/DockerImageCache resources use ResourceLocator::Tool
(a tool-native id, not a filesystem path) because Docker’s build cache has
no single canonical path. DefaultEvidenceCollector can only run
tool_liveness for a Tool-locator resource — open_by_process,
process_cwd_match, and git_state always come back
Unavailable(NotAttempted) for it, and required_evidence() still requires
all three for Complete. AutoSafe for Docker build cache is out of reach
until a future ticket gives this resource kind a real path-based or
tool-native correlation strategy.
Docker image cache is unconditionally Protected; Docker has no registered cleanup action
Detectors do not yet distinguish a persistent, named Docker volume from
disposable image cache, so every DockerImageCache resource is classified
Protected unconditionally (see Safety Model). Separately,
Docker (both build cache and image cache) has no registered Action at all
— it stays detect-only, per an accepted design cut. Neither of these is
accidental; both are documented design decisions, not bugs.
Resolved (HORO-957): Homebrew’s cleanup action is registered but never actually executes
HomebrewCleanupCache’s ActionStep::RunTool has scoped_path: None
because there is no narrower, safe way to ask brew to clean only one
thing — real brew cleanup -s has no path argument to scope to. HORO-957’s
independent golden-scenario evaluation found that, before this fix, that
meant executor::execute() ran this step with zero identity/TOCTOU
guard, and confirmed the real Homebrew cache classifies AutoSafe on a
real developer machine today — a live risk that emergency/free --target
could silently trigger a real, irreversible brew cleanup -s. Fixed:
execute() now refuses any RunTool step with scoped_path: None
unconditionally, fail-closed, before ever spawning the tool. dry_run/
clean --dry-run still render this action’s plan; real execution stays
permanently refused unless a scoped equivalent becomes available upstream.
HORO-1358 carried that refusal upstream into reporting: detect --json and
explain --json now report this candidate as executable: false with no
offered action, because reporting puts the action’s own plan to the same
pre-mutation structural rule execute() applies. HORO-1360 moved that rule
into one shared predicate (actionability) and pointed two more surfaces at
it: the BYOK LLM prompt no longer names homebrew.cleanup.cache among a
resource’s offered action ids, and autopilot reports such a candidate as
ineligible without spending one of its bounded attempts, rather than
charging an attempt and reporting an execution failure that was certain in
advance. clean --dry-run, free and emergency still resolve actions
independently and do not read the predicate, so they can still nominate
homebrew.cleanup.cache — where it remains refused at execution. That
remainder is HORO-1359.
See HORO-1005 for a related, non-blocking follow-up (scoped_path isn’t
yet structurally tied to what a RunTool step’s args actually mutate —
not currently exploitable, since this was the only unscoped action and it’s
now refused outright).
Resolved (HORO-957): Cargo target/ and node_modules are now discoverable
Before this fix, DiscoveryContext::known_project_roots defaulted to an
empty list and was never populated by any real CLI code path — so the two
most canonical “developer storage hotspots” named in the epic were
structurally undiscoverable no matter what was actually on disk. Fixed: a
repeatable --project-root <path> flag is now wired into detect/
explain/clean/free (deliberately not emergency, which takes no
arguments by design).
glomeris free still declines every ASK candidate
This section used to say Ask had no handling at all, because no interactive
prompt existed. Narrower than that now: glomeris execute --confirm-ask --observed-fingerprint <token>, the menu-bar app’s Clean button, and
Autopilot’s --preauthorize-ask all supply a real UserConsent — see
Safety Model.
What remains is specific to one command. glomeris free’s recovery loop wires
RecoveryConfig::auto_approve_ask: false (src/main.rs), so every
Ask-classified candidate inside that loop is reported as declined/skipped
and never executed, however much of the target it would have reclaimed. There
is still no TTY prompt anywhere in this codebase; a free run cannot ask you
mid-loop, so it does not ask at all. To act on an Ask candidate, use
explain --json to read its fingerprint_token and then execute.
A pre-authorized ASK under Autopilot cannot complete (HORO-1310)
Autopilot’s --preauthorize-ask <kind>:<reason> grants narrow advance consent
for one ASK reason on one kind. The gate admits such a candidate,
policy::approval::authorize issues a real Approval for it, and then
executor::execute’s deletion-time revalidation always aborts it with
AbortReason::PolicyClassDowngraded. Nothing is deleted.
The cause is upstream of Autopilot and shared by every deleting command.
Ask/RebuildCostHigh arises only from a per-instance
Regenerability::NotRegenerable, while executor::build_fresh_evidence
rebuilds regenerability from the resource kind’s static default — so the
fresh classification lands on AutoSafe and the class comparison trips.
Left as-is deliberately: the failure direction is the safe one (refuse, mutate
nothing), and changing build_fresh_evidence changes the TOCTOU anchor
execute, free, emergency and autopilot run all depend on. It is pinned
by src/autopilot/run.rs’s
a_preauthorized_ask_still_aborts_at_deletion_time_revalidation, so the fix
starts from a failing test that names the cause. See
Autopilot.
Correlation depends on lsof/git/pgrep being present and stable
Runtime correlation is subprocess-based. CI validates this against
macos-14 only, where all three tools are confirmed present; behavior on
other macOS versions is not separately verified. tool_liveness for
non-daemon tools (Cargo, Npm, Pnpm, Yarn, Homebrew) is structurally
always Unavailable(ToolNotRunning) — there is no persistent process whose
presence would mean “this tool is active” for them, which caps their
Evidence::completeness() at Partial/Confidence::Medium, never
Complete/High. This is a deliberate fail-closed consequence, not an
anomaly to “fix.”
launchd real-scheduling behavior is not exercised by cargo test
Only plist generation and path/file logic are unit tested. Actually asking
launchd to load and run the agent on a schedule needs a real macOS user
session and is not covered by CI.
The protected-path matcher is conservative, not exhaustive
See Safety Model — the denylist covers a specific, named set of categories (credential material, git internals, infra state, system paths, unsafe mounts) and is explicitly documented in its own module comment as a starting point to extend, not a completeness guarantee.
Resolved (HORO-1008): the BYOK LLM planner now has a CLI surface, and PROTECTED refusal is proven end-to-end
glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json]
is an advisory, non-executing subcommand wired on top of actions::llm —
see BYOK LLM Planner and CLI Reference. It
never constructs a policy::Approval and never calls
policy::approval::authorize or executor::execute.
--plan-file <path> feeds a fixture response through the exact same
extract_plan/LlmPlan/plan_with_llm pipeline the live provider uses,
with no network call. tests/golden_llm_plan_protected_refusal.rs uses
this to prove, through the real CLI-facing crate::cli::build_llm_plan_report
function, that a plan request to delete SSH key material is refused —
closing the gap the previous version of this section described:
HORO-943’s golden acceptance scenario step 6 (“prove a protected resource
cannot be deleted even if an LLM plan requests it”) was previously verified
at the code level only (llm_plan_item_never_bypasses_policy), never
through an actual CLI input surface.
Remaining, deliberate scope cuts for this subcommand specifically:
- Advisory-only, by design — no
--execute/--yesflag exists or is planned for this subcommand. - No retry/backoff, and no streaming — mirrors
actions::llm’s own existing limitations (see BYOK LLM Planner). - Only one provider shape (
OpenAiCompatibleProvider, any OpenAI-compatible/chat/completionsendpoint) — no Anthropic-native or Azure-OpenAI-specific auth. - Not wired into
glomeris free’s recovery loop — that remains a future ticket’s optional enhancement, peractions::llm’s own module docs.
Separately, and independently of the LLM planner: a previous version of this
section claimed that no live detector could ever emit a resource whose path
matches a policy::protected pattern, and that PolicyClass::Protected was
therefore enforced at the code level but not reachable by an evaluator
driving only the shipped product’s detectors. The first half of that is
wrong, and HORO-1313’s release-gate pass disproved it on a disposable
fixture.
What a protected pattern matches is a path component, not an installed
location, so anything a detector can discover underneath one is protected.
A real node_modules tree created under a path containing an .ssh
component was found by the live Node detector, classified Protected with
reason protected_credential_material and executable: false, and then
refused by the real executor twice — once by a plain execute, once with a
valid --confirm-ask and matching --observed-fingerprint — exiting 3 both
times with nothing deleted.
So the accurate statement is narrower, and the safety conclusion is
stronger rather than weaker. The conventional locations of the tool caches
Glomeris knows about do not normally sit under a protected path, which is
why Protected is uncommon in day-to-day use; a project root that does sit
under one reaches it through ordinary discovery, and the refusal holds when
it happens. --project-root is the usual way to get there, deliberately or
by accident.
tests/golden_llm_plan_protected_refusal.rs still reaches Protected
through a hand-built Evidence fixture, and that remains the right shape
for a hermetic test — it does not depend on a tree existing on the machine
running CI.
Resolved (HORO-957): prebuilt release artifacts
cargo-dist packaging produces macOS artifacts for aarch64-apple-darwin
and x86_64-apple-darwin, published as GitHub Release assets with
checksums. Building from source remains fully supported.
Resolved (HORO-1305 through HORO-1309): there is a GUI
This page used to say a SwiftUI app was not part of this MVP. There is one:
a LSUIElement menu-bar app, documented in Menu Bar App.
What has not changed is where authority lives. The app is a thin client over
the same CLI — it shells out to glomeris and renders what comes back. It
classifies nothing, decides nothing, and holds no policy logic, which a CI
guard (scripts/check-no-policy-label-branching.sh) enforces mechanically
rather than by convention. Every screen maps back to a named CLI invocation;
that table is at the end of the Menu Bar App page.
What the app offers no way to invoke is running something without being
asked: glomeris emergency has no button, and neither does glomeris autopilot run. Both still show up in the app’s history when run from a terminal, which
is the point of a shared audit trail.
Autopilot’s grant is a different question, and this page used to get it
wrong too: it said the envelope could only be granted or revoked on the command
line. That is no longer true. Settings → Autopilot reads
autopilot show --json and writes through autopilot enable/autopilot revoke, so the one authorization in this product that lets something delete
without asking again can be read and withdrawn by someone who never opens a
terminal — see Menu Bar App. The
distinction the app keeps is between authorizing and acting, not between the
terminal and the GUI.
What holds this book to the code
Three mechanical checks, because the drift this page is about was found by reading rather than by CI:
tests/help_golden.rspins every rendered help surface byte for byte against committed fixtures, so a command’s own help text cannot change silently.scripts/check-docs-cover-cli-commands.shrequires CLI Reference to have a section for every command insrc/cli/help.rs’s singleCOMMANDStable. A new subcommand now fails CI until it is documented.scripts/check-vocabulary-covers-cli-tokens.shcompares the CLI’s JSON tokens against the menu-bar app’s wording for them, in both directions.
None of that can catch prose that goes stale, which is what the rest of this
page is for. Where this book and src/main.rs disagree, the source is right
and the disagreement is a bug on this page.