Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

Glomeris is an evidence-first, policy-constrained developer storage autopilot for macOS. When a machine enters disk pressure, Glomeris discovers meaningful reclaimable storage (Cargo target dirs, node_modules, Homebrew’s cache, Xcode DerivedData, Docker’s reported build cache), explains why a resource is or is not safe to remove, and executes only policy-approved cleanup actions — re-measuring actual freed bytes rather than trusting an estimate.

The canonical safety invariant

AI can recommend. Policy decides. Executor verifies. Filesystem reality wins.

Concretely, as implemented today:

  • An optional BYOK LLM planner (glomeris::actions::llm) may rank and explain candidates, but it can only select from a closed, typed set of pre-registered action IDs and real resource IDs the crate already found — it can never compose a raw command or an arbitrary path (see BYOK LLM Planner).
  • A single deterministic module (glomeris::policy::engine::classify) decides AUTO_SAFE / ASK / PROTECTED for every resource. Nothing upstream of it, including any LLM output, can bypass that decision (see Safety Model).
  • The executor (glomeris::executor::execute) never trusts a previously computed decision at face value: it re-collects evidence and re-runs classify immediately before mutating anything, and aborts rather than acts if the fresh read disagrees with what was approved.
  • Every execution re-measures actual reclaimed bytes after the fact rather than assuming the estimate was correct.

What Glomeris is not

  • Not cross-platform. The MVP targets macOS only. There is no Windows or Linux support, and the platform::macos module is compiled out entirely on other operating systems.
  • Not a cloud service. There is no backend, no telemetry, and no account system. Everything runs locally: a CLI binary, an optional per-user launchd agent, and an optional menu-bar app. The one thing that can leave your machine is a BYOK LLM request you configure yourself, to an endpoint you name — bounded and previewable (see BYOK LLM Planner).
  • Not a generic system optimizer. Glomeris only understands a fixed, named set of developer-tool-owned resource kinds (Cargo, npm/pnpm/yarn, Homebrew, Xcode, Docker). It does not attempt to clean arbitrary “junk” files or ordinary user documents.
  • Not an automatic deleter of ordinary user documents. Every resource kind Glomeris knows about is a build/tool cache with a defined owning tool; an unrecognized resource kind (ResourceKind::Unknown) is classified PROTECTED unconditionally rather than falling through to any default treated as safe.
  • Not a GUI with authority of its own. There is a SwiftUI menu-bar app (see Menu Bar App) — an earlier version of this page said there was not. It is a thin client: it shells out to the same glomeris binary and renders what comes back. It classifies nothing and decides nothing; a CI guard (scripts/check-no-policy-label-branching.sh) fails the build if Swift code branches on a policy label, so the one way the GUI could quietly grow its own safety opinion is checked rather than trusted.

Two meanings of “autopilot”

Both are used in this book, so they are worth separating once:

  • The product is a storage autopilot in the sense above — it discovers, explains and verifies on its own rather than asking you to audit paths by hand. That is what the first paragraph means.
  • glomeris autopilot is one specific command: an unattended run inside a policy envelope you set on the command line — byte, action, time and resource-kind limits, AUTO_SAFE only unless you pre-authorise otherwise. See Autopilot.

glomeris autopilot run is the only thing that deletes unattended. detect, explain, clean, execute, free and emergency all need you to invoke them. The optional launchd agent does run unattended, but it only polls disk pressure and notifies — it executes no action and never touches the filesystem it is watching (see Daemon Lifecycle).

Project status

This is an experimental MVP, not a stable release — see Known Limitations for a specific, current accounting of what is and is not implemented.

The command set described here is held to the code mechanically: tests/help_golden.rs pins every help surface byte for byte, and scripts/check-docs-cover-cli-commands.sh fails CI if a command ships without a section in CLI Reference. Prose can still go stale, which is what Known Limitations is for.

Installation

Every tagged release is built and published automatically by cargo-dist (.github/workflows/release.yml), which attaches the binary tarballs directly to each GitHub Release with checksums. Download the one matching your Mac’s architecture, extract it, and copy the glomeris binary onto your PATH (e.g. /usr/local/bin).

Two macOS targets are built for every release: aarch64-apple-darwin (Apple Silicon) and x86_64-apple-darwin (Intel, cross-compiled) — see dist-workspace.toml. Pick by architecture; nothing selects it for you.

Homebrew tap

brew tap Chisanan232/tap
brew trust --tap chisanan232/tap   # Homebrew 7 and later only; see below
brew install glomeris

brew upgrade glomeris picks up later releases the same way, and Homebrew selects the right architecture automatically.

The tap is Chisanan232/homebrew-tap, and dist-workspace.toml points cargo-dist at it so each release can publish a regenerated formula there (HORO-1069, HORO-1320).

Three things about this path are worth knowing before you use it.

Homebrew 7 will not load a third-party tap until you trust it. On Homebrew 7.0 and later, brew install glomeris straight after brew tap fails with Refusing to load formula chisanan232/tap/glomeris from untrusted tap. That is Homebrew protecting you from arbitrary Ruby in a tap you have not vouched for, not a broken formula. brew trust --tap chisanan232/tap records the decision in ~/.homebrew/trust.json (or under $XDG_CONFIG_HOME/homebrew/) and only needs doing once. Older Homebrew versions have no brew trust and do not need this step.

A manually copied binary already on your PATH blocks the symlink. If you previously followed the prebuilt-archive instructions above and copied glomeris into /opt/homebrew/bin or /usr/local/bin, Homebrew installs into its Cellar but refuses to link over your file, and warns that the Homebrew copy is shadowed. Remove your manual copy, or let Homebrew take ownership:

brew link --overwrite glomeris --dry-run   # lists exactly what it would remove
brew link --overwrite glomeris

The formula currently tracks v0.2.0 and is updated by hand. The release workflow’s publish-homebrew-formula job needs a credential for the tap repository that does not exist yet, so no release has regenerated the formula automatically. That credential is the remaining half of HORO-1320. Until it is configured, the tap can lag the newest tagged release — check the releases page if you need the very latest, or use a prebuilt archive above.

Build from source

Glomeris requires Rust/Cargo (2021 edition). There is no other runtime dependency.

git clone https://github.com/Chisanan232/glomeris.git
cd glomeris
cargo build --release

The resulting binary is at target/release/glomeris. Copy it onto your PATH (e.g. /usr/local/bin) if you want to run it as glomeris directly.

Installing the menu-bar app

See Menu Bar App for what GlomerisMenuBar.app actually shows and lets you do once it’s installed.

The app runs the glomeris CLI for everything it displays. A release build ships its own copy inside the bundle, so it works with no separate CLI install; otherwise it looks along PATH and then in /opt/homebrew/bin and /usr/local/bin, which covers both Homebrew prefixes and the manual-copy instructions above. Menu Bar App documents the exact order and why PATH alone is not enough for a GUI app.

From the next tagged release onward, GlomerisMenuBar.app.zip is published as an extra asset on that GitHub Release alongside the CLI tarball, by a separate macos-app-release.yml workflow that fires once the release is published. That workflow landed after the current latest release was tagged, so releases published so far carry only the CLI tarballs — to get the app before the next tag, build macos/GlomerisMenuBar/GlomerisMenuBar.xcodeproj yourself. It’s built and ad-hoc signed automatically (codesign --force --deep --sign -, identity -) — this seals the bundle well enough to run, but it is not signed with an Apple Developer ID and not notarized. Developer ID notarization is tracked as a future follow-up, not done today.

That matters the moment you download the zip through a browser: macOS tags the extracted app with a “quarantine” flag, and Gatekeeper will refuse to just open it — double-clicking shows “Apple could not verify that GlomerisMenuBar is free of malware” with no direct way to proceed. This is expected for an ad-hoc-signed app and does not mean the download is broken or unsafe; it’s the same warning any non-notarized app gets. Follow these one-time steps to open it:

  1. Unzip GlomerisMenuBar.app.zip and move GlomerisMenuBar.app wherever you want to keep it (e.g. /Applications).
  2. Double-click it. You’ll see the “could not verify” warning — click Done (or Cancel) to dismiss it. This first attempt is expected to fail; it’s what unlocks the next step.
  3. Open System Settings → Privacy & Security, scroll down to the Security section, and you’ll see a line like “GlomerisMenuBar” was blocked to protect your Mac with an Open Anyway button next to it. Click it, then authenticate with your password/Touch ID when prompted.
  4. Double-click GlomerisMenuBar.app again. A dialog reappears asking if you’re sure — click Open. macOS remembers this decision, so every later launch works with a plain double-click, no repeat of these steps.

This System Settings path is Apple’s own documented method for opening software from an unidentified developer and is the one to use on current macOS (Sequoia and later, including Tahoe) — recent macOS versions have locked down the older shortcut of control-clicking the app and choosing Open, so that shortcut may not offer an Open option at all anymore. If it does work for you, it’s faster: control-click the app, choose Open from the menu, then click Open in the dialog that appears — but treat the System Settings steps above as the reliable path.

Terminal fallback (last resort). If neither of the above surfaces an Open Anyway/Open option — for example the security exception was already dismissed, or you’re scripting the install — remove the quarantine flag directly:

xattr -d com.apple.quarantine /path/to/GlomerisMenuBar.app

Only run this against an app you actually intend to trust: it deletes the flag that tells Gatekeeper to check the app at all, so use it deliberately, not as a routine habit.

Platform support

Glomeris targets macOS only. daemon, emergency, and free are gated on #[cfg(target_os = "macos")] and print an error and exit non-zero on any other OS; scan and detect do not have this restriction, since they only touch std::fs/std::process and don’t depend on the macOS-only platform::macos module.

Verifying the build

cargo test
cargo clippy --all-targets -- -D warnings

Both are part of this repository’s CI (.github/workflows/ci.yml, macOS runner) and should pass on a clean checkout.

Menu Bar App

GlomerisMenuBar.app is a native SwiftUI menu-bar app (Epic HORO-1043, “Glomeris MVP 2.0”) that gives the glomeris CLI a graphical front end. See Installation for how to download and open it.

Thin client, not a second decision-maker

This is the project’s core invariant for the app, stated verbatim in the Swift source (macos/GlomerisMenuBar/Sources/GlomerisMenuBarApp.swift):

GlomerisMenuBar is a THIN CLIENT ONLY, over the existing Rust glomeris CLI’s --json output. This Swift codebase renders and lets a human approve/decline what the Rust CLI already decided.

NO policy classification (AUTO_SAFE/ASK/PROTECTED), NO evidence correlation, NO action planning, and NO filesystem execution logic may EVER live here.

In practice this means every screen below is a formatting step over one of glomeris’s own --json reports (status, daemon status, detect, explain, history, actions history, execute, autopilot show|enable|revoke), spawned as a subprocess. The app never re-derives “is this safe to clean” from policy_label or reasons — it reads the already-computed executable, offered_actions, and refusal_reason fields the CLI provides for exactly this purpose (see CLI Reference). If you trust the CLI’s policy decisions on the command line, you’re trusting the exact same decisions in the menu bar — the app has no separate opinion.

How the popover reads

Every section is a titled card in a scrolling column, ordered by the questions you open the panel with: what the disk is doing now, what could be reclaimed, what has already happened. Three presentation rules hold across all of them:

  • Plain language first, the CLI’s own token second. PRESSURED reads as “Running low”, aborted_by_revalidation as “Stopped safely” — and the raw token stays visible beside or beneath it, because this is an evidence-first product and you have to be able to match what the panel says against --json and against the docs.
  • Three separate axes, never merged. Storage impact (how much space), safety (what the policy allows) and evidence quality (how sure the CLI is) are shown as distinct badges. A large AUTO_SAFE candidate is an opportunity, not a hazard; a small PROTECTED one is still protected. Size is therefore drawn in a neutral tone at every magnitude.
  • Colour is never the only signal. Every badge carries a symbol and a word as well as a tint, so the state survives greyscale, a colour-vision deficiency and a tinted wallpaper. A refusal (PROTECTED) is deliberately not painted like a failure: it is the policy working, and there is nothing for you to fix.

Disk space, and Background monitor

Two cards, polled on appear and every 10 seconds while the popover is open. The two facts are deliberately never collapsed into a single “healthy” indicator:

  • Current disk pressure (glomeris status --json) — used percent, free space, and pressure state (one of HEALTHY, WARN, PRESSURED, CRITICAL, EMERGENCY).
  • Daemon health (glomeris daemon status --json) — whether the launchd agent is loaded, and how long ago the poll loop last recorded a heartbeat ("Last heartbeat: 42s ago", or "no heartbeat recorded").

A launchd-loaded-but-wedged daemon and an actually-polling one stay visibly distinguishable — the app never merges loaded and heartbeat_age_secs into one boolean.

Reclaimable space

Shows the most recent glomeris detect --json scan (“Last scanned: <time>”) plus a Refresh button. Nothing is scanned automatically: there is no appear-triggered scan, no timer, and no background polling loop for candidates — detect runs only when you explicitly tap Refresh, streaming live per-detector progress via --progress-json while it works (button label switches to “Scanning…”).

Five states that are easy to conflate are kept distinct, because each is a different claim about your disk:

StateWhat it says
“No scan yet”Nothing has been looked at. Not a clean bill of health.
“Nothing worth reclaiming”Scanned, and there is genuinely nothing — good news.
“Nothing found where Glomeris could look”Scanned, found nothing, but part of the search never answered — so whatever is there is unknown rather than absent. Names which checks did not finish.
“No candidates match this filter”Things were found; you are just not looking at them.
A scan failureSays what failed. An empty list is never shown in its place.

The third state also has a non-empty counterpart (HORO-1484): when a detector fails but others still found candidates, the rows are shown as normal with “This list may be incomplete” added below them. The rows are real and stay on screen; what they may not do is look like the whole account. A detector whose tool is simply not installed is not a failure and produces neither message — docker being absent is normal, brew being asked and failing is not.

Each row shows the candidate’s kind, what cleaning it would free, and its safety verdict in words. Tapping a row opens its detail view — there is no inline “Clean” button in this list.

Rows arrive in the CLI’s own order — biggest reclaimable size first — and are shown exactly as detect returned them. The app does no ranking of its own: deciding what matters most is a judgment, and it belongs next to the evidence in Rust rather than in a thin client that would then drift from it.

The View options menu (the funnel next to Refresh) offers two things:

  • Order — “Biggest first”, which is the CLI’s order untouched, or “Path (A–Z)”, an alphabetical index for finding a resource whose path you already know. There is deliberately no third “largest first” option: that is “Biggest first”.
  • Show — All, Safe to reclaim, Asks first, Protected, or Not enough evidence. This hides rows and does nothing else; it cannot enable, authorise or perform anything, and the list has no action affordance for it to unlock. Protected items are a first-class filter value rather than something hidden by default: “what on this machine is off-limits, and why” is a reasonable question, and quietly omitting them would teach you that protection means invisibility.

When a filter is active the header reads “N of M items” so a shortened list can never be mistaken for a smaller problem.

Size is not safety

A third badge appears on the rows worth pausing on — “Biggest wins” or “Worth a look” — using the impact_tier the CLI already computed. It is an emphasis hint about magnitude only, carries no safety colour, and appears on no more rows than deserve it (ordinary and unmeasured candidates get no badge rather than a badge saying “normal”, which would be noise).

Storage impact, safety, and evidence confidence are three separate axes and the popover never collapses them. A large candidate may be protected; a small one may be perfectly safe to reclaim. Nothing is distinguished by colour alone — every badge pairs its colour with a distinct symbol and words — and VoiceOver reads each row as size, then safety, then emphasis, so the badge is never the only route to the fact.

Candidate detail

Opened by tapping a candidate row; this is the sole place a “Clean” button exists in the whole app. It calls glomeris explain --json for the exact resource, which supplies fingerprint_token — the same token execute --observed-fingerprint later pins consent to. This is also why there’s no per-row Clean button in the candidates list above: cleaning a resource always goes through one explain call that captures the fingerprint the confirmation flow needs.

The sheet is ordered by the questions you have when you open it: may I clean this, is it worth cleaning, on what evidence — and last, collapsed behind a disclosure, the literal values explain --json returned, which stay selectable so they can be quoted in a bug report.

  • The Clean button’s enabled state reads executable from the explain report directly — never a re-derived guess from policy_label.
  • If the resolved action requires_confirmation (an ASK or UNKNOWN_INCOMPLETE candidate), tapping Clean shows a confirmation alert before doing anything.
  • Confirming runs glomeris execute --action-id <id> --resource-id <id> --confirm-ask --observed-fingerprint <token> --progress-json, using the exact fingerprint token captured from the explain call above — never a fingerprint freshly re-observed at click time, which would defeat the whole point of fingerprint-pinning.
  • The outcome (succeeded, failed, aborted_by_revalidation, or a specific refusal reason like protected/ask_consent_mismatch) is rendered from the real ExecuteReport/ExecuteRefusalReport JSON the CLI printed — one specific message per outcome, not a generic success/failure toast.

AI Plan

The one card that talks to the internet, and only ever when you press its button. There is no timer behind it, nothing runs when the popover opens, and nothing retries — asking a provider costs money, so asking has to be something you did. Until you press it the card says exactly that: “Nothing has been sent anywhere.”

Pressing Ask AI for a plan runs glomeris llm-plan --json --progress-json with your project roots, streams the same per-detector progress the Reclaimable space card shows, and offers a Stop button that terminates the CLI child rather than just abandoning the result. Above the button, permanently: “The model recommends. Glomeris decides what may run.” and a note that asking costs money and that “Settings shows exactly what would be sent, without sending it” — the privacy preview described under Preferences — AI Provider below.

Each suggestion is one row, and every row is split down the middle by who said what:

  • The machine’s half — safety class, storage impact, the size tier, and the evidence/confidence pair — is rendered as badges from the same vocabulary the candidates list uses, and comes from the real policy engine. It reads identically whether or not a provider was ever contacted. Below them, in plain words, is what Glomeris is willing to do: “Glomeris is willing to run this”, “Glomeris will ask you to confirm this before anything runs”, or the refusal, verbatim.
  • The model’s half — its rationale — is an attributed, italic quotation under a “The model says” label. It is deliberately not a badge. A badge in this app means a verdict was reached; putting a model’s sentence in one would dress an opinion as a finding. A screen reader hears the machine’s verdict first, then “The model’s reason, which is advice and not a verdict: …”.

The rows appear in the order the provider returned them, and the card says so (“The model’s order. Glomeris’s own ranking is the list above.”). No rank number is drawn on a row, and the model’s own priority field is carried in --json but never shown: the list position already makes that claim once, and two competing numberings would read as though one of them were authoritative. Nothing in the app re-sorts the list either — Glomeris’s ranking is the Reclaimable space card, which is why the AI Plan card sits below it.

A recommendation cannot make anything runnable. A model that confidently describes your SSH private key as a stale build directory gets its sentence quoted, next to a PROTECTED badge and a refusal, with no action offered — the row is built from the same executable/offered_actions/refusal_reason fields the candidates list reads, which the CLI computed before the provider was contacted. Suggestions naming a resource Glomeris never found, or an action it does not have, are dropped by the CLI and the count of them is printed under the list rather than quietly swallowed.

Tapping a row opens the same candidate detail sheet as the candidates list, which issues its own explain call and owns the only Clean button in the app. A plan item carries no fingerprint token, so there is no shortcut past that — cleaning something a model suggested goes through exactly the path, and the same confirmation, as cleaning something you found yourself.

Six situations, six distinct messages: never asked; asking; no provider configured; the provider answered with nothing; the provider call failed (quoted); and output this app could not read, which means the app and the glomeris on PATH are different versions. The first three are not failures and are not coloured like failures. The no provider configured message says why a shell user can still land there — a Finder-launched menu-bar app inherits no shell environment, so GLOMERIS_LLM_* variables exported in a terminal are invisible to it — and is the one state with a Set up an AI provider… button beside it, which opens Settings. The button is beside the message rather than inside it on purpose: the copy layer is a pure token-to-words mapping, and a message that could carry a control would be a message that could act.

Disk space history, and What Glomeris has done

Two independent, already-computed lists in two cards, each polled the same way as the status section:

  • Disk space history — the last N pressure transitions from glomeris history --json (e.g. WARN -> CRITICAL, with the used-percent/free-space reading at the time). The two ends of each transition are named exactly as the status card names them, because they are the same enum.
  • What Glomeris has done — the last N real-execution records from glomeris actions history --json: which action ran against which resource, its policy label, outcome, and which of execute/free/emergency produced it. “Triggered by” is spelled out in words, because whether something was cleaned that you did not ask for is the question this card exists to answer. The AUTO_SAFE-vs-refused/aborted visual distinction reads outcome/abort_reason directly, never the policy_label text — and aborted_by_revalidation reads as “Stopped safely” rather than as a failure, because the resource changed between checking and acting and so nothing was touched.

An empty list in either card says only that it is empty. Neither claims an all-clear it cannot support: an empty pressure history could mean the disk has been steady, or it could mean nothing has been watching it, and the report does not distinguish those.

Neither list accepts a --project-root flag — both read global daemon-state files (history.tsv/actions.jsonl), not project-scoped detection state.

Preferences — Project Roots

The popover’s Project roots… footer button — or Cmd+, — opens Settings, whose Projects tab is a simple list with add/remove controls for the project-roots preference, backed by a small local store. (The footer exists because the menu-bar item opens a window rather than a menu, so the popover is the app’s only surface; Quit is there for the same reason.) These are the same paths you’d otherwise pass repeatedly as --project-root <path> on the command line — the CLI’s cargo/node detectors only look under directories they’re told about. This view is pure presentation: it edits the stored list and does not itself call detect/explain/ execute; the roots are appended as --project-root arguments the next time another card (Disk space, Reclaimable space, …) spawns the CLI.

Preferences — AI Provider

Settings’ second tab (HORO-1309) is where the BYOK provider is configured, for the common case of someone who runs the app from Finder and has never exported anything in a shell. Four cards, in the order the questions arise:

AI Provider — the API root address and the model name. Under each field, a row saying where that field’s effective value came from: Configured here, From the environment, or Not set. Per field, what you type here wins over what the app inherited; a field left empty falls back to the inherited GLOMERIS_LLM_* variable, including falling back to it being absent. Clearing a field therefore returns you to the environment rather than switching a working setup off. The rule is that the settings screen never lies: if the field shows an address, that is the address used. Neither field has a Save button, because neither has unsaved state — they write through as you type. Nothing here validates the URL or knows which models exist; the CLI does both, and Test connection is what reports it.

API Key — a SecureField, a Save button, and a Remove key button marked as the destructive action it is. The key goes into the login keychain, and from there into the environment of the glomeris child process at the moment one is spawned. It is never a command-line argument (visible to ps, and refused by the CLI by flag name), never this app’s preferences, never a file the app writes, never a log line, and never displayed: there is no reveal control, nothing reads it back, and the typed text is cleared as soon as it is handed to the keychain — whether the write succeeded or not, so a failure cannot leave it sitting in a field behind an error message.

Test Connection — one glomeris llm-check --json, described honestly as “one tiny request — two words”, disabled until all three of the address, the model and the key are available. The result is one of five outcomes in plain words, each reading differently because each has its fix in a different place: a rejected credential, an unreachable host, a reply the app could not use, a base URL that cannot work, and success — which reports the path it posted to, the model, and what came back. See BYOK LLM Planner for the outcome table. A failure is carried verbatim; nothing is paraphrased into something more reassuring than what happened.

What Gets Sent — one glomeris llm-plan --print-payload --json, with the same project roots a real plan would use, so it previews your payload rather than a generic example. It splits the result into Leaves this Mac (the two prompts, with a character count) and Stays on this Mac (the wire-alias table, which is where the absolute paths are). The second group is blue, not green: green in this app means AUTO_SAFE or succeeded, and data being withheld is a deliberate hold, not a success.

Opening that preview sends nothing and cannot — --print-payload returns before a provider is constructed — and the preview is the one command on this screen that runs without the credential in its environment. Neither button does anything until pressed: there is no timer, nothing runs when the window opens, and nothing retries.

Preferences — Autopilot

Settings’ third tab (HORO-1310). Autopilot is the one feature that grants standing permission to delete without being asked again, and the person granting it is the least likely to be reading --help — so a grant that could only be made and read on the command line would be a grant most users would never read. This tab is where AC 6 of that ticket (“see exactly what Autopilot is authorized to do before enabling it”) and AC 7 (“revoke immediately”) are met for someone who never opens a terminal.

It is still a thin client. Every choice it offers arrives as data from glomeris autopilot show --json: the checkboxes are allowlistable_kinds, the steppers’ upper bounds are ceilings, the picker’s options are pressure_states, the one question that may be answered in advance is preauthorizable_reasons, and the “never available” lists are never_allowlistable_kinds, never_preauthorizable_reasons and never_executable_labels verbatim. If the CLI cannot be reached the tab draws no form at all and says so, rather than falling back on its own idea of what Autopilot allows — a settings screen offering a limit it invented would be a policy decision made in Swift.

Eight cards, in the order the questions arrive:

  • Autopilot — on or off, then what is in force right now in the CLI’s own figures (including max_bytes_human, so the number you read is formatted by the thing that enforces it), the envelope’s path, a Revoke now button and a Refresh. An enabled grant is drawn in the caution tone, never the critical one: a standing authorization you chose is not a fault.
  • What It May Reclaim — one checkbox per allowlistable kind, in the CLI’s own order, each with its plain-language name and what it is. Below them, the kinds that are never available whatever is granted here.
  • How Much, Per Run — three steppers (actions, whole gigabytes, seconds), described as three independent budgets a run stops at whichever it reaches first, with the hard ceilings quoted underneath.
  • When It May Act — the disk-pressure floor, including “Whenever there is something to reclaim” for no floor at all.
  • Answering In Advance — the narrow ASK pre-authorization, one kind:reason pair at a time. Until a kind is ticked there is nothing to answer, and the card says that instead of offering a consent that would apply to nothing.
  • Grant This / Change This Authorization — the write. Disabled until at least one kind is ticked, and it says so.
  • What the AI Decides — the ai_authority sentences, quoted rather than paraphrased, and the labels Glomeris will never delete whatever is authorized here and whatever a model recommends.
  • See What It Would Do — a copyable glomeris autopilot run --dry-run.

Five properties of this screen are worth stating outright, because each is a way it could have been quietly wrong:

The form is pre-filled from what is in force, because enable replaces the whole envelope. glomeris autopilot enable does not merge into the previous grant (see Autopilot for why), so a form that started from defaults would let you narrow one field and silently reset the other five. Pressing Enable without touching anything re-grants exactly what was already there. For the same reason the byte budget rounds up to the next whole gigabyte: a 1.5 GiB grant shown as “1 GB” would mean that merely opening this window and saving shrank a budget nobody touched.

The floors here are this screen’s, and they only narrow. The steppers stop at 1 action, 1 GB and 5 seconds. The CLI accepts --max-actions 0 — an enabled grant that can do nothing — which is coherent as an API and pointless as a setting, so this window does not offer it. Nothing here can ask for more than the CLI’s ceilings, which is a property of the values the report carries rather than of a number typed in Swift.

Withdrawing a kind withdraws its advance consent with it. Unticking a kind prunes any kind:reason pre-authorization that named it, so consent cannot outlive the kind it applied to, or quietly come back when the kind is re-ticked.

There is no way to start a run from this window. Not a Run button, not a dry-run button, no timer, nothing on appear: the only thing that runs automatically is the read. autopilot run is also the one autopilot verb with no --json, so the thin client has nothing to parse if someone adds a button later without thinking about it. A run deletes, and the argument for keeping deletion on a command you typed is the same one that keeps glomeris emergency out of the app.

A revocation that failed is never reported quietly. The two write directions fail in opposite ways — a failed enable leaves you with less authority than you asked for, a failed revoke leaves standing deletion authority in force — so they have separate wording, and every failed revoke says that the authorization still is in force and how to withdraw it with glomeris autopilot revoke. Exit 0 with output this app cannot read is treated as a successful write in both directions, because the CLI prints its report after the envelope has been saved.

The grant does not expire on its own. It is a file, it survives quitting and restarting, and it stays in force until it is revoked here or on the command line — which the card that grants it says in those words, because an authorization the user believes is temporary would be the worst kind of quiet.

How the app finds the glomeris CLI

Every screen spawns the CLI, so the app has to decide which binary that is. It checks these locations in order, and uses the first one that exists and is executable:

  1. Inside the app bundle, at Contents/MacOS/glomeris. Release builds ship the CLI here, and it wins because it is version-matched to the app and covered by the bundle’s signature. Debug builds contain no embedded CLI, so development is unaffected.
  2. Each absolute directory in PATH, in order. Relative entries — including the empty entry that shells read as “the current directory” — are ignored, so nothing can substitute a binary by writing a file named glomeris next to the running process.
  3. /opt/homebrew/bin, then /usr/local/bin — the Homebrew prefix on Apple Silicon and on Intel respectively, and where Installation tells you to put a binary from a release tarball or a source build.

Step 3 is not redundant with step 2: a GUI app does not inherit your shell’s PATH. An app launched from Finder gets PATH=/usr/bin:/bin:/usr/sbin:/sbin, which contains neither Homebrew prefix — so PATH alone would not find a brew installed CLI, even though running glomeris in a terminal works fine.

Resolution happens on every invocation, not once at launch, so installing the CLI while the app is already running takes effect at the next poll without a restart.

If no binary is found, each section says so and lists the locations it searched, rather than failing silently or naming a path it only assumed.

Seeing which binary is in use

The Command-line tool card names the binary the app resolved: its full path, which of the three rules above chose it, and a fingerprint — the first twelve characters of the SHA-256 of its contents, with the whole hash in the tooltip.

The fingerprint is there because the version number cannot do this job. Two glomeris binaries on one machine both reported version 0.2.0 while disagreeing about whether an action with no scoped path may be offered at all: the version is stamped from the crate version, so it does not change between merges and is not a build identity. The hash is.

A release build records the hash of the CLI it ships, so the card can say whether the binary in use is that one. The four things it can say:

The card saysWhat it means
Matches this appThe resolved binary is the one this app ships.
Not the build this app shipsBoth hashes are known and differ. The app still works, and still drives that binary — what it does may differ from what the app describes.
Nothing to compare againstThis app records no expected hash. Normal for a locally built app; see below.
Cannot be checkedThe binary can be run but its contents could not be read, so no comparison was possible.

A locally built app embeds no CLI. The embed step lives in the release workflow, not in the Xcode project, so an app built with xcodebuild — or from Xcode — contains no Contents/MacOS/glomeris and always falls through to PATH or the Homebrew prefixes. Whatever is installed on the machine is what your build drives, and that may be older or newer than the tree you built from. A local build also records no expected hash, which is why the card reads “Nothing to compare against” rather than reporting a mismatch.

If you are testing a change to the CLI from a local app build, the path on the card is the one to check: cargo build alone does not put a binary anywhere the app looks.

Every screen maps back to a CLI command

CardCLI command(s)
Disk spaceglomeris status --json
Background monitorglomeris daemon status --json
Reclaimable space + Refreshglomeris detect --json --progress-json
AI Plan — Ask AI for a planglomeris llm-plan --json --progress-json
Settings → AI Provider — Test connectionglomeris llm-check --json
Settings → AI Provider — Show what would be sentglomeris llm-plan --print-payload --json (no provider contacted)
Settings → Autopilot — on appear, Refreshglomeris autopilot show --json
Settings → Autopilot — Enable / Save changesglomeris autopilot enable --kinds <tag,...> --max-actions <N> --max-bytes <N> --max-duration <secs> --min-pressure <state|none> [--preauthorize-ask <kind>:<reason>]... --json
Settings → Autopilot — Revoke nowglomeris autopilot revoke --json
Candidate detailglomeris explain <resource_id> --json --progress-json
Clean (with confirmation)glomeris execute --action-id <id> --resource-id <id> [--confirm-ask --observed-fingerprint <token>] --json --progress-json
Disk space historyglomeris history --json
What Glomeris has doneglomeris actions history --json

See CLI Reference for the full flag/exit-code/JSON-shape reference behind every one of these.

Quick Start

A short tour of the commands worth running first. The complete list, with every flag and exit code, is in CLI Reference — and in the binary itself, which is the better habit: glomeris --help, then glomeris help <command> for any one of them.

If you would rather click than type, the menu-bar app covers detect, explain, clean and AI Plan over the same binary — see Menu Bar App.

Find out where to start

glomeris

Prints the version and one line pointing at glomeris --help. Nothing else; a bare invocation reads no disks and changes nothing.

Discover what tool caches exist on this machine

glomeris detect

Runs every built-in detector (Xcode DerivedData, Homebrew cache, Cargo target dirs, node_modules, Docker build cache) once and prints one line per detector: found (<N> evidence), tool_absent, or failed: <reason>. tool_absent is a normal, expected state — it means that tool isn’t installed or has no cache yet, not an error.

Scan a directory for the largest entries

glomeris scan [path] [top_k]

Both arguments are positional, not flags, and both are optional: path defaults to ., top_k defaults to 20. This walks the tree without buffering it fully in memory and prints the top-K largest entries by logical size. It is a generic size scan, unrelated to the detectors above — see Architecture for how the two differ.

Try the bounded recovery loop (macOS only)

Set a goal for how full the disk should end up, and let the loop work toward it:

glomeris free --goal-used-percent 60

Or state the same thing as a free-space floor, which is what the loop itself works in:

glomeris free --target 10GB
# or
glomeris free --target 15%

Exactly one of the two is required, and they are different axes: --goal-used-percent is target disk used, --target is a free-space floor. --goal-used-percent 60 and --target 40% ask for the same end state. --target accepts either an absolute size (B/KB/MB/GB/TB, binary/1024-based) or a percentage of total capacity (0–100, suffixed %). A --goal-used-percent that is not an improvement on your current usage is refused before anything is deleted, rather than run and reported as a success.

This one deletes. An earlier version of this page said no real deletion could complete through this path; that stopped being true in HORO-994. Read Safety Model first, and note that free declines every ASK candidate rather than prompting (see Known Limitations) — so it reclaims only what policy classified AUTO_SAFE on its own.

To see what reaching the goal would take without touching anything:

glomeris free --goal-used-percent 60 --dry-run

That prints current usage, the free bytes still needed, and the estimated reclaimable opportunity split by what policy would actually permit. It takes no execution lock and mutates nothing. glomeris clean --dry-run remains the way to see a plan for one pass rather than a goal.

Run unattended, inside limits you grant

glomeris autopilot                                  # what am I allowing today?
glomeris autopilot enable --kinds node_modules \
    --max-actions 1 --max-bytes 1073741824
glomeris autopilot run --dry-run                    # the real bounded plan

The one command that acts without you watching, so the grant is written to a file you can read rather than inferred. It grants nothing until you enable it, --kinds is required, and --max-bytes is a plain byte count. AUTO_SAFE only unless you pre-authorise one specific ASK reason on one specific kind; PROTECTED refuses unconditionally and no flag here changes that. See Autopilot.

Try emergency mode (macOS only)

glomeris emergency

Takes no arguments by design, and acts machine-wide on everything it finds AUTO_SAFE. See Emergency Mode for exactly what it does and does not do before running it on a machine you care about.

CLI Reference

The built-in help is canonical. glomeris <command> --help renders from src/cli/help.rs, which is the one place a command’s usage, flags, safety semantics, examples and exit codes are written down, and whose output is held to golden snapshots in tests/fixtures/help/ (HORO-1311). If this page and --help ever disagree, --help is right and this page is stale.

What this page adds that --help deliberately does not: the JSON payload shapes, the report field semantics, and the cross-references into the rest of this book. It is the reference you read at a desk; --help is the one you read mid-task.

glomeris

No arguments: prints glomeris <version> (from CARGO_PKG_VERSION) and a pointer to glomeris --help. --version/-V prints the same version line on its own.

Help surfaces

Everything below renders from the single COMMANDS table in src/cli/help.rs. That table lived in src/main.rs until HORO-1311, where no test could read it — so tests/shared_command_table.rs kept a hand-written mirror of it, which drifted exactly as the pre-HORO-1050 duplication had (HORO-1034: the unrecognized-command usage omitted llm-plan; by HORO-1311 the test mirror omitted llm-check). There is now one table, read directly by both the binary and its tests.

InvocationWhat it prints
glomeris --help / -h / helpEvery command, grouped by what it does to your machine, with a one-line summary each.
glomeris <command> --help / -hThat command’s usage, safety statement, its subcommands (each with its own safety label), flags, examples, exit codes and related commands.
glomeris help <command>Identical to glomeris <command> --help.
glomeris help exit-codesThe exit-status reference (see Exit codes).
glomeris help <unknown-topic>The list of topics that exist, to stderr, exit 2.
glomeris <unknown-command>A one-line error and a pointer to --help, to stderr, exit 2.
glomeris <command> <bad-flag>That command’s usage only, to stderr, exit 2.

The groups in top-level help are INSPECT, PLAN, ACT, OBSERVE, SERVICE and CONFIGURE, ordered so that the read-only commands come before anything that can delete, and so that a first-time reader meets the product before its preferences. A group heading describes consequence, not category: INSPECT says “nothing is changed”, and ACT says its commands delete data and that every deletion is policy-gated.

Note that a command’s group and its safety line describe only whether Glomeris itself writes to the filesystem when you run it. They are not policy classifications: AUTO_SAFE, ASK and PROTECTED classify resources, are decided by the policy engine, and never appear in a help safety label. See Safety Model.

Safety is declared per verb, not only per command

Four labels exist, in ascending order of consequence:

LabelMeans
Read-only — changes nothing.Nothing is written.
Advisory — proposes, never executes.Produces a plan; executes none of it.
Writes only Glomeris's own state — never your files.Writes the launch agent plist, the monitor’s history and heartbeat, the Autopilot envelope, or your stored preferences. Nothing you own.
Can delete data — every deletion is policy-gated.Deletes.

For a command that takes subcommands, the consequence is a property of the verb, not of the command name: daemon install writes a launch agent while daemon status reads, and autopilot run deletes while autopilot show prints a config file. Until HORO-1485 one label was declared per command, so glomeris daemon --help printed “Read-only — changes nothing.” above install, uninstall and run, and glomeris autopilot --help printed “Can delete data” above show. No wording could have fixed either: the weakest label is a false reassurance and the strongest a false warning, which is why relabelling daemon as destructive was not an acceptable fix.

Each verb now carries its own label, printed beside it in the Subcommands block. The command’s own safety line is the strongest claim reachable through it, and when its verbs disagree it says so explicitly rather than picking one of them:

Safety depends on the subcommand; each is labelled below. The strongest of
them: Writes only Glomeris's own state — never your files.

Two tests keep this honest, and they are deliberately not tests about strings. src/cli/help.rs‘s command_safety_covers_every_reachable_subcommand asserts the command’s label equals the strongest of its verbs’ — equality in both directions, because under-stating and over-stating are both false — and a_single_label_banner_is_true_of_every_verb_under_it asserts the same of the rendered banner, since a correct table printed through a banner that ignored it would still print a false claim. Neither could have caught the original defect on its own: with no subcommands declared, both pass vacuously. tests/subcommand_safety_is_honest.rs closes that gap from outside the table — it reads src/main.rs’s dispatch arms, so a verb the binary accepts cannot go undeclared, and it runs every surface labelled read-only against a disposable HOME and requires the tree to be byte-for-byte unchanged, with autopilot enable as the positive control that proves the observation would notice a write.

Two surfaces are narrower on purpose. An unrecognized top-level command gets a pointer rather than the manual — it previously reprinted the aggregate usage of all thirteen commands, a 13-line, 146-column wall in answer to one mistyped word. And a usage error inside a command prints only that command’s usage, so getting a free flag wrong no longer tells you about daemon.

Byte counts: 1024-based, with KB/MB/GB labels

Every byte count this product renders — free_human, total_human, logical_human, reclaimable_human, expected_reclaimed_human, actual_reclaimed_human, and the same numbers in text output — comes from one function, reporting::human_bytes. It divides by 1024 and labels the result B/KB/MB/GB/TB/PB. So 2147483648 renders as 2.0 GB, where a 1000-based formatter would say 2.15 GB. Counts below 1024 are a bare integer and B, with no decimal point.

The labels are not the IEC KiB/MiB/GiB spelling that strictly matches the arithmetic. That is a deliberate, documented inconsistency rather than an oversight: these are the units du -h and df -h print for the same arithmetic, and the alternative is renaming every unit in every report and fixture to spell out a distinction most readers of a storage tool do not draw.

--target parses the same way, so what you type and what you read back agree: glomeris free --target 5GB means 5 × 1024³ bytes. A bare number or a B suffix is raw bytes; a trailing % is a percentage of total capacity instead.

Clients should render the *_human string rather than scale the byte count themselves. The menu-bar app did the latter with ByteCountFormatter, which is 1000-based, so a 2 GiB cleanup appeared as a 2.0 GB estimate and a 2.15 GB result in the same panel (fixed in HORO-1312). Where a report offers both fields, the raw *_bytes value is for arithmetic and the *_human string is for display. A *_human field is null, never "0 B", when the underlying probe produced no number — an aborted execution reclaimed nothing, which is a different claim from having freed zero bytes.

glomeris daemon <subcommand>

macOS only (exits 1 with an error message on other platforms).

SubcommandEffect
installWrites a per-user launchd plist and launchctl load -ws it.
uninstalllaunchctl unload -ws the agent (best-effort) and removes the plist file.
statusPrints whether the plist is installed, its path, and whether launchctl reports it loaded.
runRuns the polling loop in the foreground (this is what the installed agent actually executes).

No subcommand, or an unrecognized one, prints usage to stderr and exits 2. See Daemon Lifecycle for details.

daemon status --json (HORO-1045) prints a DaemonStatusReport. loaded (launchd-reported) and heartbeat_age_secs (derived from the poll loop’s own last-write) are deliberately kept as two separate fields, never collapsed into one healthy boolean — a loaded-but-wedged daemon and an actually-polling one must stay distinguishable:

glomeris daemon status --json
{
  "plist_installed": true,
  "plist_path": "/Users/dev/Library/LaunchAgents/dev.glomeris.daemon.plist",
  "loaded": true,
  "heartbeat_age_secs": 42
}

heartbeat_age_secs is null when no heartbeat file exists yet (the daemon has never run).

glomeris status [--json]

macOS only (exits 1 with an error message on other platforms). Not part of daemon — this is the same one-shot disk-pressure reading daemon run’s poll loop evaluates each tick, available on demand without needing the daemon installed at all (HORO-955).

glomeris status --json
{
  "total_bytes": 500000000000,
  "free_bytes": 125000000000,
  "used_percent": 75.0,
  "free_human": "116.4 GB",
  "total_human": "465.7 GB",
  "pressure_state": "WARN"
}

pressure_state is one of "OK", "WARN", "CRITICAL" — see Pressure Model.

glomeris scan [path] [top_k]

Not macOS-gated. Both arguments are positional and optional:

  • path — root directory to scan. Defaults to ..
  • top_k — number of largest entries to report. Defaults to 20 if omitted or unparseable as usize.

Prints a summary line (files visited, stop reason, incomplete-entry count) followed by one line per candidate: size, depth, path.

glomeris detect [--project-root <path>]... [--json] [--progress-json]

Not macOS-gated. Runs every detector in DetectorRegistry::builtin() exactly once per invocation and prints, per detector: found (<N> evidence), tool_absent, or failed: <reason>. The candidate report printed below those lines comes out of that same single pass, so the two halves of the output cannot describe different probes of a filesystem that changes between them (HORO-1487).

--project-root <path> is optional and repeatable — pass it once per project directory you want the cargo/node detectors to check for a target//node_modules/ dir. Without it, those two detectors have no project roots to scan and always report tool_absent.

--json prints a DetectReport — one candidate line per discovered resource, including the already-computed executable/offered_actions/refusal_reason triple (HORO-1053) a caller (e.g. the menu-bar app) reads to decide what it can offer, without ever re-deriving that from policy_label/reasons itself:

Candidates are returned biggest reclaimable size first (HORO-1307). Before that they came back in detector-registration order, so a 40 GB Cargo target/ could be printed below a 2 MB npm cache. Ties are broken first by measurement quality — an exact size outranks a ≥ lower bound of the same number, because the exact one is the claim you can act on — and then by resource_id, so two runs over an unchanged machine produce the same order. Candidates whose size could not be measured at all sort last rather than being treated as zero. Ordering is applied at the single point where the report is assembled, so --json and the human-readable output can never disagree about it.

impact_tier is "unknown", "normal", "notable" or "large": a pre-computed magnitude band, so a UI does not have to invent thresholds of its own. It escalates on either an absolute size (≥ 1 GB is notable, ≥ 10 GB is large) or a share of remaining free space (≥ 5% is notable, ≥ 20% is large) — measured against free rather than total space, because the problem someone opens Glomeris with is “I am running out of room”. A 400 MB cache is large on a machine with 1.6 GB left.

impact_tier is a size signal and nothing else. It is not a safety signal, and it must never be read as one: a large candidate can be PROTECTED, and a normal one can be AUTO_SAFE. executable, offered_actions and refusal_reason remain the only statement about what Glomeris is permitted to do.

glomeris detect --project-root ~/dev/myproject --json
{
  "candidates": [
    {
      "resource_id": "cargo_target_dir:/Users/dev/proj/target",
      "kind": "cargo_target_dir",
      "reclaimable_bytes": 2147483648,
      "reclaimable_human": "2.0 GB",
      "reclaimable_bytes_is_lower_bound": false,
      "impact_tier": "notable",
      "policy_label": "AUTO_SAFE",
      "reasons": ["no_active_use_observed"],
      "executable": true,
      "offered_actions": [
        {
          "action_id": "cargo.clean.target_dir",
          "requires_confirmation": false
        }
      ],
      "refusal_reason": null
    },
    {
      "resource_id": "docker_build_cache:docker",
      "kind": "docker_build_cache",
      "reclaimable_bytes": 10737418240,
      "reclaimable_human": "10.0 GB",
      "reclaimable_bytes_is_lower_bound": true,
      "impact_tier": "large",
      "policy_label": "UNKNOWN_INCOMPLETE",
      "reasons": ["evidence_incomplete"],
      "executable": false,
      "offered_actions": [],
      "refusal_reason": "no registered cleanup action for this resource kind"
    }
  ],
  "detectors": [
    {"detector": "cargo_target_dir", "status": "found", "candidates_found": 1},
    {"detector": "docker_build_cache", "status": "found", "candidates_found": 1},
    {"detector": "node_modules", "status": "tool_absent", "candidates_found": 0},
    {
      "detector": "homebrew_cache",
      "status": "failed",
      "candidates_found": 0,
      "reason": "brew --cache exited with status exit status: 1"
    }
  ],
  "discovery_complete": false
}

detectors reports one entry per detector that ran, in registry order, and status is found, tool_absent or failed — the same three tokens --progress-json uses below. reason is present only on failed, and is the detector’s own account of what went wrong.

discovery_complete (HORO-1484) is false when at least one detector failed, and it is derived from detectors rather than tracked separately, so the summary cannot disagree with the array it summarises. When it is false, candidates is not a complete account of what could be reclaimed, and no consumer may present it as one — “nothing worth reclaiming” is a claim about the machine, and a search that did not finish has not established it. A tool_absent detector does not make discovery incomplete: a tool that is not installed has nothing to report, which is a different fact from a tool that was asked and could not answer.

--progress-json (HORO-1052) emits one NDJSON-encoded ProgressEvent line to stderr per detector start/finish while discovery runs — a way for a spawning UI to distinguish “still working” from “hung” on a slow/contended host (v0.2.0 founder-dogfood measured a single discovery pass up to 3m40s). Stdout is completely unaffected either way, --progress-json output can be combined with --json, and nothing is emitted at all unless the flag is passed:

glomeris detect --progress-json 2>&1 1>/dev/null
{"phase":"detector_started","detector":"cargo_target_dir"}
{"phase":"detector_finished","detector":"cargo_target_dir","candidates_found":1,"outcome":"found"}
{"phase":"detector_started","detector":"docker_images"}
{"phase":"detector_finished","detector":"docker_images","candidates_found":0,"outcome":"tool_absent"}
{"phase":"detector_started","detector":"homebrew_cache"}
{"phase":"detector_finished","detector":"homebrew_cache","candidates_found":0,"outcome":"failed","reason":"brew --cache exited with status exit status: 1"}

outcome (HORO-1484) is found, tool_absent or failed, and reason is present only on failed. The last two lines above are why it exists: a detector whose probe failed reports candidates_found: 0, exactly like one that looked and found nothing, so a consumer reading only the count shows “0 found” for a check that never ran. docker_images being absent is normal and expected; homebrew_cache failing is not, and the two must not be presented alike.

detect, explain, llm-plan, and execute all share this same discovery phase and all support --progress-json identically.

glomeris explain <resource_id_or_path> [--project-root <path>]... [--json] [--progress-json]

Not macOS-gated. Runs the same discovery-and-classification pipeline as detect, then prints the full evidence-and-policy picture for exactly the one resource matching <resource_id_or_path> (a detect report’s resource_id, or a filesystem path) — evidence provenance, size (logical vs. reclaimable, explicitly labeled as different), active-use signals, and policy classification (HORO-955). Exits 1 if no discovered candidate matches the query.

--project-root <path> and --progress-json mean exactly what they mean for detect above.

glomeris explain cargo_target_dir:/Users/dev/proj/target --json
{
  "resource_id": "cargo_target_dir:/Users/dev/proj/target",
  "kind": "cargo_target_dir",
  "detector": "cargo_target_dir",
  "sources": ["cargo metadata: target-dir"],
  "logical_bytes": 2147483648,
  "logical_human": "2.0 GB",
  "reclaimable_bytes": 2147483648,
  "reclaimable_human": "2.0 GB",
  "reclaimable_bytes_is_lower_bound": false,
  "completeness": "complete",
  "confidence": "high",
  "active_use_signals": [],
  "regenerability": "regenerable_by_rebuild",
  "policy_label": "AUTO_SAFE",
  "reasons": ["no_active_use_observed"],
  "native_cleanup_available": true,
  "native_cleanup_action_id": "cargo.clean.target_dir",
  "fingerprint_token": "<opaque token — copy verbatim, never hand-construct>",
  "executable": true,
  "offered_actions": [
    {
      "action_id": "cargo.clean.target_dir",
      "requires_confirmation": false
    }
  ],
  "refusal_reason": null
}

fingerprint_token (HORO-1051) is an opaque, wire-safe encoding of the resource’s identity fingerprint — null for a resource with no dev/inode/ mtime identity (e.g. Docker’s build cache). This is the exact token execute --observed-fingerprint later expects back for an ASK-classified resource; see execute’s section below.

glomeris clean --dry-run [--target <resource_id_or_path>] [--project-root <path>]...

Not macOS-gated. Renders what would be cleaned, without executing anything — --dry-run is required; there is no non-dry-run execution path on this subcommand (real destructive execution is glomeris free --target or glomeris execute’s job). Without --target, every discovered candidate is considered; with it, only the one matching resource is.

glomeris clean --dry-run

Human-readable output only — clean --dry-run has no --json mode. One line per considered resource, either its rendered ActionPlan.explain text or a skip_reason (e.g. PROTECTED, no registered action).

glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json] [--progress-json] [--schema] [--print-payload]

Not macOS-gated. ADVISORY, NON-EXECUTING (HORO-1008) — never constructs a policy::Approval and never calls policy::approval::authorize or executor::execute. See BYOK LLM Planner for the full configuration and safety-property writeup.

  • Without --plan-file, credentials are read only from GLOMERIS_LLM_API_KEY/GLOMERIS_LLM_BASE_URL/GLOMERIS_LLM_MODEL (all three required, no default base URL) — never from a CLI flag. GLOMERIS_LLM_BASE_URL is the API root: /chat/completions is appended to it verbatim, so it usually ends in /v1 (https://gateway.example.com/v1). See BYOK LLM Planner.
  • A failed provider call is reported on the provider error: line (or the provider_error JSON field) as one secret-free sentence naming the HTTP status, API style, request path, request id, and a bounded excerpt of the provider’s error body — never the Authorization header, the key, the host, or the request payload.
  • --plan-file <path> reads the file’s raw bytes as if they were the model’s raw response text, through the same validation pipeline the live provider uses — no network call, no API key required.
  • --project-root <path> is optional and repeatable, same meaning as detect’s flag above.
  • --api-key/--key/--token are explicitly rejected (not accepted and ignored) — the error names $GLOMERIS_LLM_API_KEY instead.
  • --json prints the report as JSON. Each item carries the model’s own priority/model_reason alongside the machine’s policy_label, requested_action_id, explain, skip_reason, completeness, confidence, and a nested candidate — the byte-for-byte same projection detect --json prints for that resource, including executable/offered_actions/refusal_reason (HORO-1308). A consumer deciding what may be done reads candidate; priority/model_reason are the only two fields a provider chose, and the order of items is the provider’s too — plan_with_llm never sorts, it only drops. --json output is printed before a provider_error exit, so an exit 1 report is still complete and readable.
  • --progress-json streams the same NDJSON discovery progress as detect, on stderr, leaving stdout a single clean JSON document. This is the exact pair (--json --progress-json) the menu-bar app’s AI Plan card spawns; see Menu Bar App.
  • --print-payload (HORO-1298) runs discovery, prints the exact request a live run would send — the system prompt, the user prompt, and the local wire-id-to-real-resource table under a heading marking it as not sent — then returns before any provider is constructed. Requires no credential and makes no network call, so it cannot send what it displays. The outbound prompts identify resources only by positional alias (resource_1, …); absolute paths appear in the alias table and nowhere else. See BYOK LLM Planner.
  • --schema (HORO-1048) prints an example, syntactically valid LlmPlan JSON document to stdout and exits — a distinct, self-contained mode that never runs discovery, never reads --plan-file, and never checks live-mode credentials, regardless of what else is passed alongside it. See BYOK LLM Planner for the full example and field table.
glomeris llm-plan --schema
{
  "items": [
    {
      "resource_id": "cargo_target_dir:/path/to/project/target",
      "action_id": "cargo.clean.target_dir",
      "priority": 1,
      "reason": "stale build artifacts, not modified in 30 days"
    }
  ]
}

This exact output round-trips unchanged through glomeris llm-plan --plan-file <path> — see the linked BYOK page for the full field table and the round-trip test that proves it. The resource_id form shown here is the one a human writes by hand in a fixture; a live model is given positional wire aliases and answers with those, and both forms resolve.

Human-readable output always opens with LLM SUGGESTION — advisory only, nothing is executed by this command.

Exit codes for this subcommand specifically:

  • 0 — success, including zero suggestions or every suggestion being PROTECTED.
  • 1 — the provider call or response parsing failed, --plan-file named an unreadable path, or --print-payload could not serialize the request.
  • 2 — usage error: an unrecognized argument, --plan-file with no value, missing live-mode environment configuration, or an --api-key/--key/ --token flag.

glomeris llm-check [--json]

Not macOS-gated. Tests the configured BYOK setup and reports whether the endpoint, credential and model work (HORO-1309). Runs no detectors, reads no project roots, collects no evidence, and consults no policy — it is not a planning command, and there is nothing it could execute.

  • Sends the two fixed prompts in actions::llm::CONNECTION_TEST_SYSTEM_PROMPT / CONNECTION_TEST_USER_PROMPT (“You are a connection test. Reply with the single word: ok.” / “ok”) through the same LlmProvider::complete a real plan uses. Same code path, so a pass means a plan will route and authenticate — not that a cheaper probe succeeded.
  • The prompts are constants with no interpolation, so a connection test describes nothing about this machine: no path, no home directory, no account name (connection_test_prompts_describe_nothing_local).
  • Configuration comes only from GLOMERIS_LLM_API_KEY/GLOMERIS_LLM_BASE_URL/GLOMERIS_LLM_MODEL, all three required, exactly as for llm-plan. --api-key/--key/--token are rejected by flag name, and the error names the flag only — never the value beside it.
  • --json prints an LlmCheckReport: outcome, model, endpoint_path, error, response_excerpt. Human-readable output opens with LLM CONNECTION OK or LLM CONNECTION FAILED (<outcome>).
  • outcome is one of five tokens, produced by the single actions::llm::llm_check_outcome mapping so the CLI, the book and the menu-bar app cannot disagree about what a failure was:
outcomeWhat happenedWhere the fix is
okThe provider answered and its reply is excerpted in response_excerpt—
unreachableNo HTTP response at all — DNS, TLS, refused connection, timeoutNetwork, host name, or VPN
rejectedThe provider answered with a non-2xx statusCredential, or the base-URL path — read the path in error
unusable_responseA 2xx response that was empty, not JSON, or missing choices[0].message.contentModel name, or a gateway not actually speaking the OpenAI shape
misconfiguredA base URL that cannot work — validate_base_url refused it before anything was sentThe base URL itself; see BYOK LLM Planner

The fifth is the odd one out: LlmError::InvalidConfiguration can only come from constructing the provider, which happens before any report exists, so glomeris llm-check never prints a report whose outcome is misconfigured — it exits 2 with one line on stderr instead. The token is in the vocabulary because that is the name for what happened, and a caller branching on exit 2 (the menu-bar app does) classifies it that way itself rather than inventing a sixth word for the same condition.

  • error carries the same secret-free sentence llm-plan’s provider_error does — status, API style, request path, x-request-id when the provider sends one, and a bounded excerpt of the provider’s own error body, with the configured key scrubbed out of it. Never the Authorization header, the key, the scheme, the host, or the query string.
glomeris llm-check --json

Exit codes for this subcommand specifically:

  • 0 — the provider answered and outcome is ok.
  • 1 — a check ran and did not pass (unreachable, rejected, unusable_response). The report is printed first, so an exit 1 still carries a complete, readable diagnosis on stdout.
  • 2 — nothing was sent and there is no report at all: an unrecognized argument, an --api-key/--key/--token flag, a base URL validate_base_url refused, or missing configuration (the message names the three variable names, never a value). Stdout is empty; the reason is one line on stderr prefixed glomeris llm-check: .

The split matters for a caller branching on the code without parsing output, which is exactly what the menu-bar app’s connection test does: 2 means “fix your invocation or setup”, never “the network is having a bad day”. A 2 with no report is the one case where a UI has to quote the CLI’s own sentence rather than render a report.

glomeris execute --action-id <id> --resource-id <id> [--project-root <path>]... [--confirm-ask --observed-fingerprint <token>] [--json] [--progress-json]

macOS only (exits 1 with an error message on other platforms). The sole interactive destructive-execution subcommand (HORO-1055) — the only place in this CLI where a caller can trigger one specific, real destructive action against one specific, real resource. --action-id and --resource-id are both required; the caller supplies ONLY these selectors (plus, for ASK, an observed fingerprint token) — there is no flag to pass a PolicyClass, a raw filesystem path as a direct target, a shell string, or any --force/override.

Internal flow, in one process:

  1. Acquires the same HORO-1054 execution lock free/emergency use — held for the whole call, released on exit.
  2. Runs the same discovery-and-classification pipeline detect/explain use (--project-root <path> is optional and repeatable, same meaning as elsewhere).
  3. Resolves --resource-id against the discovered candidates and --action-id against that resource’s own registered action — refusing if either does not resolve, or if the resolved action’s id differs from --action-id.
  4. For an ASK-classified resource, builds consent ONLY from decoding --observed-fingerprint (via the same token format explain --json’s fingerprint_token field emits) — never from a fingerprint freshly observed by this same process, which would defeat the whole fingerprint-pinning purpose. --confirm-ask and --observed-fingerprint must be passed together or not at all.
  5. Calls the real, unmodified policy::approval::authorize, then — only if it returns an approval — the real, unmodified executor::execute. PROTECTED refuses unconditionally regardless of any flag combination; execute’s own deletion-time revalidation can still abort a plan that was authorized a moment earlier if the resource changed in between.

--json prints an ExecuteReport (action id, resource id, outcome, failure/abort detail, expected vs. actual reclaimed bytes — the latter is a real measurement, taken after execution, not an estimate) on the Executed path. Every refusal/not-found/busy path instead prints an ExecuteRefusalReport ({"reason": "...", "message": "..."}) to stdout before exiting with the matching code below, so a --json caller never gets silent stdout on a non-Executed outcome. reason is one of resource_not_found, action_not_found, action_mismatch, protected, ask_no_consent, ask_consent_mismatch, auto_safe_contract_violation, or busy (HORO-1056: the lock-contention case below, the one refusal that happens before discovery/resolution even runs) — each a distinct, machine-readable value naming exactly which refusal/abort path fired, never a generic error string.

An AUTO_SAFE resource needs no confirmation flags at all:

glomeris execute --action-id cargo.clean.target_dir \
  --resource-id cargo_target_dir:/Users/dev/proj/target --json
{
  "action_id": "cargo.clean.target_dir",
  "resource_id": "cargo_target_dir:/Users/dev/proj/target",
  "outcome": "succeeded",
  "failure_message": null,
  "abort_reason": null,
  "expected_reclaimed_bytes": 2147483648,
  "actual_reclaimed_bytes": 2147483648,
  "expected_reclaimed_human": "2.0 GB",
  "actual_reclaimed_human": "2.0 GB"
}

The two *_human strings are new in HORO-1312 and are what a UI should display — see Byte counts.

An ASK-classified resource requires --confirm-ask plus the exact --observed-fingerprint token captured from a prior explain --json call on that same resource (never a fingerprint freshly observed by execute itself) — $FINGERPRINT_TOKEN below is that call’s fingerprint_token field, copied verbatim, never hand-constructed:

glomeris execute --action-id cargo.clean.target_dir \
  --resource-id cargo_target_dir:/Users/dev/proj/target \
  --confirm-ask --observed-fingerprint "$FINGERPRINT_TOKEN" \
  --json --progress-json

If the resource’s identity changed between the explain call and this execute call, revalidation aborts the plan rather than proceeding:

{
  "action_id": "cargo.clean.target_dir",
  "resource_id": "cargo_target_dir:/Users/dev/proj/target",
  "outcome": "aborted_by_revalidation",
  "failure_message": null,
  "abort_reason": "ResourceIdentityChanged",
  "expected_reclaimed_bytes": 2147483648,
  "actual_reclaimed_bytes": null,
  "expected_reclaimed_human": "2.0 GB",
  "actual_reclaimed_human": null
}

Note both actual_* fields are null rather than 0/"0 B". Nothing was deleted, which is not the same report as a cleanup that freed no bytes.

Every refusal path (e.g. PROTECTED, no consent supplied, a stale fingerprint) prints an ExecuteRefusalReport instead, with --json:

{
  "reason": "protected",
  "message": "refused — this resource is PROTECTED; no flag combination can authorize executing against it"
}

Exit codes for this subcommand specifically:

  • 0 — the action executed and succeeded.
  • 1 — the action executed but failed (a step of the plan errored).
  • 2 — usage error: an unrecognized/missing argument, --confirm-ask without --observed-fingerprint (or vice versa), or a malformed --observed-fingerprint token.
  • 3 — refused by policy: PROTECTED (unconditional), ASK with no consent supplied, or ASK with a supplied consent that did not match the freshly observed fingerprint.
  • 4 — aborted by execute’s own deletion-time revalidation (a TOCTOU- style guard: the resource’s identity or policy classification changed between authorization and execution).
  • 5 — --resource-id matched no discovered candidate, the resource had no registered action, or the resolved action’s id did not match the supplied --action-id.
  • 75 — the execution lock is already held by another glomeris invocation (see free’s exit codes above; execute reuses the exact same EXIT_EXECUTION_LOCK_BUSY constant — this ticket’s own AC described this case as exit 6, but the already-established lock convention from HORO-1054 is kept rather than introducing a second, conflicting “busy” code). With --json, this prints an ExecuteRefusalReport with reason: "busy" to stdout (HORO-1056) — previously this path was silent on stdout even under --json.

glomeris emergency

macOS only (exits 1 with an error message on other platforms). Takes no arguments, and rejects any with exit 2 — until HORO-1311 it silently discarded them, so glomeris emergency --dry-run performed a real recovery run. See Emergency Mode.

glomeris history [--json] [--limit <N>]

Reads back a bounded, oldest-first tail of the monitor’s history.tsv (HORO-1046) — the same append-only file daemon run’s poll loop already writes via PersistenceBackend::record. No new persistence format; this is a read path only.

--limit <N> is optional and defaults to 20. It bounds how many of the most recent pressure transitions are returned — a malformed line in history.tsv is skipped rather than failing the whole read, and a missing history file (the daemon has never run, or never recorded a transition) renders as an empty list rather than an error.

With --json, prints a HistoryReport ({"events": [...]}); each event has unix_time_secs, from, to, used_percent, free_bytes, and free_human. Without --json, prints one line per event as plain text.

glomeris history --json --limit 2
{
  "events": [
    {
      "unix_time_secs": 1700000000,
      "from": "OK",
      "to": "WARN",
      "used_percent": 82.5,
      "free_bytes": 80000000000,
      "free_human": "74.5 GB"
    },
    {
      "unix_time_secs": 1700000600,
      "from": "WARN",
      "to": "CRITICAL",
      "used_percent": 95.1,
      "free_bytes": 20000000000,
      "free_human": "18.6 GB"
    }
  ]
}

glomeris actions <list [--json]|history [--json] [--limit <N>]>

Neither subcommand is macOS-gated — list touches no filesystem/launchd state at all, and history only reads a plain file.

glomeris actions list [--json]

Enumerates every action currently registered in ActionRegistry::builtin() (HORO-1047) — the read path that replaced having to read src/actions/homebrew.rs source directly to find a real action id string. applies_to is a direct projection of each action’s own Action::applies_to, never a hand-maintained list.

glomeris actions list --json
{
  "actions": [
    { "action_id": "cargo.clean.target_dir", "applies_to": ["cargo_target_dir"] },
    { "action_id": "node.clean.node_modules", "applies_to": ["node_modules"] },
    { "action_id": "homebrew.cleanup.cache", "applies_to": ["homebrew_cache"] }
  ]
}

glomeris actions history [--json] [--limit <N>]

Reads back a bounded, oldest-first tail of actions.jsonl (HORO-1057) — the real-execution audit trail that execute, free, emergency and autopilot run each append to, best-effort, after their own outcome is already decided. Unlike history.tsv (which records pressure transitions only), this is the audit trail of what was actually executed: action id, resource id, the policy label it was authorized under, outcome, abort reason (when applicable), actual reclaimed bytes, and which real-execution path produced it.

--limit <N> is optional and defaults to 20, same bounding/malformed-line- skip/missing-file-empty contract as glomeris history. An audit-write failure never affects the execution it was trying to record — the write is best-effort and its result is never surfaced to the caller.

With --json, prints an ActionHistoryReport ({"events": [...]}); each event has timestamp, action_id, resource_id, policy_label, outcome, abort_reason, actual_reclaimed_bytes, actual_reclaimed_human, and source. Without --json, prints one line per event as plain text.

source is one of five values, produced by ActionSource::as_str in src/monitor/persistence.rs:

sourceThe path that executed it
executeglomeris execute, one action against one named resource
freeglomeris free --target, the recovery loop
emergencyglomeris emergency, machine-wide, AUTO_SAFE only
autopilot_auto_safeglomeris autopilot run, action policy allowed on its own
autopilot_preauthorized_askglomeris autopilot run, action attempted only because autopilot enable --preauthorize-ask had already named that kind and reason

The last two are deliberately distinct rather than one autopilot value: the question an audit trail has to answer is not just what ran but who permitted it, and a pre-authorized ASK was permitted by the operator naming that resource kind, not by policy alone.

Note “attempted” in that last row. A record is written for a failed or aborted attempt too — not for a refused one, which never reached the filesystem and so appears in autopilot run’s own report instead. Today every pre-authorized ASK aborts at deletion-time revalidation — so autopilot_preauthorized_ask currently only ever appears alongside outcome: "aborted_by_revalidation". That is a known limitation with a named cause and a pinning test, not the intended end state: see Known Limitations.

glomeris actions history --json --limit 2
{
  "events": [
    {
      "timestamp": 1700000000,
      "action_id": "cargo.clean.target_dir",
      "resource_id": "cargo_target_dir:/Users/dev/proj/target",
      "policy_label": "AUTO_SAFE",
      "outcome": "succeeded",
      "abort_reason": null,
      "actual_reclaimed_bytes": 2147483648,
      "actual_reclaimed_human": "2.0 GB",
      "source": "execute"
    },
    {
      "timestamp": 1700000600,
      "action_id": "node.clean.node_modules",
      "resource_id": "node_modules:/Users/dev/proj/node_modules",
      "policy_label": "ASK",
      "outcome": "aborted_by_revalidation",
      "abort_reason": "ResourceIdentityChanged",
      "actual_reclaimed_bytes": null,
      "actual_reclaimed_human": null,
      "source": "free"
    }
  ]
}

glomeris free (--goal-used-percent <N> | --target <N%|NB>) [--dry-run] [--json] [--progress-json] [--stop-file <path>] [--autopilot [--unattended]] [--project-root <path>]...

macOS only (exits 1 with an error message on other platforms). Exactly one of the two goal flags is required; any other argument is rejected (usage printed to stderr, exit 2).

The two axes

The goal can be stated on either axis, and the two flags are not interchangeable wordings of one thing:

FlagAxisMeaning
--goal-used-percent <N>disk usedStop when the volume is at most N percent used.
--target <N%|NB>free spaceStop when at least this much of the volume is free.

--goal-used-percent 60 and --target 40% request the same end state. --goal-used-percent is the product-facing form — it is the number a person reads off a status bar, and it is what the Glomeris app sends. --target is the raw free-space floor the recovery loop itself works in, and its meaning and value formats are unchanged.

Passing both exits 2 rather than picking one: they are different numbers about how much of a disk to delete, so there is no safe precedence between them.

Accepted --target value formats:

  • A percentage: a number followed by %, in the range 0–100 (e.g. --target 15%). Target is “at least this percent of total capacity free.”
  • An absolute byte amount: a number optionally followed by B, KB, MB, GB, or TB (case-insensitive; no suffix means raw bytes). Multipliers are binary/1024-based — --target 5GB means 5 * 1024^3 bytes free, not 5 * 10^9.

--goal-used-percent takes a plain number from 0 to 100; a trailing % is tolerated. It is additionally checked against the current reading and refused with exit 2 if it is not an improvement on current usage, because recovering toward a goal you already satisfy would delete nothing and still print a report that reads like a successful cleanup. That check runs before the execution lock is taken and before anything is deleted. --target deliberately keeps its older, looser behaviour: a free-space floor you already exceed is a legitimate no-op probe.

Other flags

  • --dry-run prints the pre-flight for the goal — current usage, free bytes still needed, and the estimated reclaimable opportunity split by what policy would actually permit — then exits. It takes no execution lock and mutates nothing. A --target pre-flight is rendered on the used axis too, so no surface has to show a percentage whose axis is unstated.
  • --json prints a machine-readable report instead of prose. With --dry-run that is a RecoveryPreviewReport; without it, a RecoveryRunReport. A goal refused before the run starts prints a RecoveryGoalRejectionReport (reason, message, goal_used_percent, current_used_percent) and still exits 2. A busy execution lock prints the same {"reason": "busy", ...} refusal execute --json does, and exits 75.
  • --progress-json streams NDJSON progress on stderr — one complete object per line, never on stdout, so a --json report stays parseable as exactly one document. With --dry-run it is the discovery-scan stream detect emits. Without it, it reports the recovery loop itself; see Watching a run below.
  • --stop-file <path> asks the loop to stop after the action it is currently running, once path exists. See Stopping a run.
  • --autopilot bounds the run by the stored Autopilot envelope instead of by this command’s own limits. See Running inside the grant.
  • --unattended declares that nobody asked for this run. Only valid with --autopilot, and it needs a second permission the grant states separately. See Runs nobody asked for.
  • --project-root <path> is optional and repeatable, same meaning as detect’s flag above — it feeds the same DiscoveryContext the recovery loop discovers candidates from.

Watching a run

With --progress-json, a real run emits one JSON object per line on stderr as it proceeds. Every line carries a phase and the 1-based iteration it belongs to:

phaseWhenAlso carries
measuredThe volume was read (loop step 1). The only statement of fact about free space in the stream.total_bytes, free_bytes, used_percent, free_human, bytes_freed_so_far
discoveringDetectors are being asked what exists now. Re-entered every iteration — no pass reuses an earlier one’s list.—
discoveredThat pass finished.candidates (how many have a resolvable action), detectors_failed
revalidatingEvidence is being re-collected and reclassified before anything is chosen.—
action_startedA real mutation is about to run.resource, action, policy_label, estimated_bytes
action_finishedIt finished.resource, action, outcome, reclaimed_bytes, bytes_freed_so_far
stop_requestedA stop was observed, between actions.—

Two rules hold across every line:

  • estimated_bytes is the only estimate. reclaimed_bytes and bytes_freed_so_far are measured from the filesystem. Never accumulate the estimate as progress — it is what the candidate claimed, not what happened.
  • A count that could not be measured is absent, not zero. An action_finished line with no reclaimed_bytes means the size was not determinable; "0 B" there would be a measurement nobody took.

These phase names never collide with detect’s discovery stream, so a client reading both cannot mistake a scan for a run.

Stopping a run

--stop-file <path> is a sentinel: create the file and the loop stops after the action it is currently running, finishing with stop_reason: "stopped_by_user" and a complete report.

  • The path must not exist yet (exit 2 otherwise). A sentinel left behind by an earlier run would stop the next one before it did anything, and the report would truthfully say the user stopped it while the user had done nothing.
  • Cooperative, deliberately. Nothing here can interrupt a deletion mid-flight. A signal delivered partway through one would leave the filesystem in a state neither the loop nor the audit log could describe, so the loop asks between actions instead.
  • Not valid with --dry-run, which performs no actions to stop.

Running inside the grant

--autopilot runs this same loop under the standing grant glomeris autopilot enable wrote. It starts no second loop and introduces no second policy: every candidate is still classified, revalidated and executed exactly as it would be otherwise, and the envelope can only withhold one.

  • Purely subtractive. Anything an --autopilot run does, a run you started yourself would also have done. Nothing the envelope says can make a PROTECTED resource executable, admit an ASK whose exact kind and reason were not pre-authorized, or reach past a scoped-path or revalidation check.
  • No grant, no run. Defaults grant nothing, so with Autopilot revoked or never enabled the command exits 3 having attempted nothing — the same code autopilot run uses, so a script can tell “not authorized” apart from a usage error (2) and from a failed run (1). A revoked envelope refuses exactly as an absent one does; revoke keeps the limits on file, and only enabled stands between them and a run.
  • Its limits replace this command’s, and are the stricter of the two. The run-level action count and wall-clock budget come down to the envelope’s own figures. Left at their defaults, the loop’s 600-second ceiling could cut short a run the grant had authorized for the full 15 minutes and report budget_exceeded — a true sentence about the wrong budget.
  • The envelope is printed first, on stderr, before anything is discovered. Output that may end in deletions opens with the authority it acted under rather than asking you to go and look it up afterwards. stdout still carries exactly one report.
  • A stop the envelope caused says so. stop_reason: "envelope_refused" with an envelope_refusal token naming which limit it was — an exhausted action, byte or time budget, a kind outside the grant, a pressure floor not met, a grant revoked mid-run. That is deliberately not safe_exhausted: “this disk has nothing safe left” and “Autopilot reached the limit you set” call for opposite next steps, and only one of them is a finding about the disk.
  • Not valid with --dry-run (exit 2). An envelope is authority to execute; a preview executes nothing, so the flag would have nothing to narrow and the candidate list would look envelope-filtered while being the whole of it. Read the grant with glomeris autopilot show.

Runs nobody asked for

An envelope answers “what may a run do”. --unattended is about a different question — “may a run start when nobody pressed anything” — and that is a second permission, granted by glomeris autopilot enable --respond-to-alerts and off by default.

  • Two consents, not one. An enabled envelope alone does not authorize an unprompted run, and --unattended without that permission exits 3 having attempted nothing. The reason is what the alternative would mean for a grant already on disk: somebody wrote it to bound a run they intended to start, and reading it as consent to act while they are away would widen a grant they already read and approved.
  • The check is here, in this binary. --unattended is the caller’s own statement about itself, and it exists so that whatever launched the process cannot be the thing that decides it was allowed to. Without the flag, a client starting an --autopilot run on its own initiative would be indistinguishable from a person typing the same command.
  • It narrows nothing else. Every budget, the kind allowlist and the ASK pre-authorizations apply to an unprompted run exactly as they do to one you started. --unattended --autopilot is not a mode with different limits; it is the same bounded run with one additional permission checked before it begins.
  • revoke stops it along with everything else, without clearing the preference — so revoking never looks as though it had silently reset what you chose.
  • Only with --autopilot (exit 2 otherwise). A run nobody asked for is precisely the one that must be bounded by a grant somebody read, so the combination that would be unbounded is refused at the boundary rather than later.

Reported figures

Every percentage in the output names its axis (60% used (40% free), 20% free). Every byte figure in a finished run is measured, and target_met is decided from the re-measured free space after the run, never from the sum of what the detectors estimated. The reclaimable figures in a --dry-run preview are estimates and are labelled as such; protected space is counted but never added into any opportunity total.

The prose output prints a RecoveryReport: the goal or target, stop reason (as both a token and a sentence — a run that stopped because nothing safe remained says so rather than printing a success word), iterations run, actions executed, actions declined/skipped, bytes freed, and free space before/after.

A run that stopped because nothing safe remained also reports what is still there, as four separate counts rather than one total: candidates awaiting your confirmation, candidates whose action cannot run against the resource as it currently stands (a live tool, work in progress), protected resources, and — for an --autopilot run — candidates this grant was not authorized to take. Each is a different next step, so they are never summed, and the fourth is kept apart from the other three for a further reason: those are facts about what is on the disk, and that one is a fact about the run’s authority. Folded into the protected count it would tell you something is off limits when it is one setting away. They are counts of candidates, not bytes, because space the run was not permitted to take is not an opportunity.

Under --json this is the remaining object, present only for stop_reason: "safe_exhausted" and for an envelope_refused stop that had already discovered candidates: no other stop concluded anything about the candidates it never reached, so zeros there would be a claim the run did not make. An envelope refused before discovery — revoked, or a pressure floor not met — omits the object entirely for exactly that reason, rather than publishing a breakdown of zeros that would read as a completed search. See Safety Model and Known Limitations for what this loop can and cannot currently do end to end.

glomeris autopilot <show|enable|revoke|run>

The subcommand is optional and defaults to show, so a bare glomeris autopilot reads the grant rather than acting on it. run is macOS only (exits 1 with an error message on other platforms); the other three verbs work anywhere.

VerbWhat it doesCan it delete?
showPrints the stored envelope and its path. Does not create the file it reads.No
enableWrites a new envelope from this command line’s flags and turns Autopilot on. Requires --kinds.No
revokeTurns Autopilot off, keeping the limits. Effective for the next run; nothing to restart.No
runConsiders discovered candidates within the envelope. Deletes unless --dry-run. Holds the execution lock.Yes

The flags, their defaults and their hard ceilings are documented in Autopilot, which is also where the argument for why an LLM plan file cannot expand authority lives. What this page adds:

  • --max-bytes takes raw bytes only — unlike free --target, there is no GB/MB suffix parsing here, because an envelope is written once and read many times and an exact number is easier to audit than a rounded one.
  • --min-pressure accepts any PressureState name case-insensitively (healthy, warn, pressured, critical, emergency) plus the literal none. An unobservable reading fails any floor you set, rather than passing it.
  • --respond-to-alerts belongs to enable and takes no value. It grants the one thing the limits above cannot express: that a run may begin without being asked, in answer to a disk-pressure alert (free --autopilot --unattended). Off by default and deliberately not implied by enable itself; see Runs nobody asked for for why those are two consents. It changes nothing about what a run may do once it starts, and revoke suspends it with the rest of the grant while keeping it on file.
  • --preauthorize-ask is repeatable and takes kind:reason using the same tags glomeris actions list and glomeris explain print. Like the four limit flags above it, it belongs to enable — run accepts only --dry-run, --plan-file and --project-root, and exits 2 on anything else, so a run cannot widen its own grant on the command line.
  • --json belongs to show, enable and revoke, and prints one shape from all three: the grant (enabled, allowed_kinds, ask_preauthorizations, the four limits, min_pressure, and the two unprompted-run fields below), the hard ceilings, every choice enable would accept (allowlistable_kinds, preauthorizable_reasons, pressure_states), what no envelope can authorize (never_allowlistable_kinds, never_preauthorizable_reasons, never_executable_labels), the ai_authority sentences, and stored_at. It always describes what is in force after the command ran — enable reports from what reached the file, not from what it was about to write — so a client reads one shape to learn one thing. run does not take it: the one reason to add it would be to make triggering a run from a GUI easier, and Menu Bar App explains why there is no such button.
  • The JSON carries two fields about unprompted runs, and they answer different questions. respond_to_alerts is the setting as you left it, in force or not — the one a settings toggle binds to, so that revoking Autopilot does not read as having silently cleared the preference. starts_unprompted is whether an unprompted run is authorized right now: this grant is in force and it says so. Anything that acts consults the second, and is not the thing that computes it — enabled && respond_to_alerts evaluated in a client would be that client deciding its own authority.
  • run prints an AutopilotReport: one line per candidate with its kind, policy label, model rank and outcome, then the run totals (actions attempted, actions succeeded, bytes freed, whether a budget stopped it early) and the envelope it ran under. Refusals appear here and nowhere else — glomeris history is a log of what happened to the filesystem, and a refusal did not touch it.

Exit codes: 0 success, including a run that found nothing it was allowed to do; 1 an action failed, or the envelope file could not be read or written; 2 usage error, including an unknown resource kind, a limit above its ceiling, or enable without --kinds; 3 Autopilot is not enabled, so nothing was attempted; 75 another invocation holds the execution lock.

3 exists so that “no grant” is distinguishable from “granted, ran, found nothing” in a script — both of which are quiet, and only one of which means the user has something to configure.

glomeris settings <show|set> [--notify-at-used-percent <N>] [--default-goal-used-percent <N>] [--json]

The subcommand is optional and defaults to show, so a bare glomeris settings reads your preferences rather than changing them.

VerbWhat it doesCan it delete?
showPrints both preferences and their path. Does not create the file it reads.No
setStores one or both preferences. Requires at least one flag.No

Two preferences live here, and the whole point of the surface is that they are not the same number:

PreferenceQuestion it answersRangeDefault
--notify-at-used-percentWhen should Glomeris call my attention to disk usage?1–99 percent used75
--default-goal-used-percentWhere should recovery stop?0–100 percent used70

Both are percent of capacity USED, never free — the same axis free --goal-used-percent takes, converted to the core’s percent-free --target by the single adapter documented in the glomeris free section above. Every line this command prints names its axis, in prose and in JSON, because 75 beside 70 with no label is the one misreading this feature cannot afford.

What this page adds:

  • The goal must be below the threshold. A goal at or above the point that raised the alert would be satisfied the moment it was announced, so the pair is refused with the reason goal_not_below_notify_threshold.
  • Both flags are validated as a pair, in one step. Moving from (75, 70) to (60, 55) is a valid destination that no single-field order can reach — whichever field moved first would be momentarily invalid against the old value of the other. So set refuses you for the destination you asked for, never for the order your flags happened to appear in.
  • A trailing % is accepted on either value, because a user typing what they read on screen has not made a mistake.
  • set with neither flag is a usage error rather than a successful no-op: a command line that did not say what it wanted should not report that it did it.
  • The four disk-pressure states (healthy, warn, pressured, critical, emergency boundaries) are not configurable and this command cannot reach them. That is deliberate: if the warn boundary were a preference, lowering it would change what a warn recorded last week meant, and raising it would make the next poll report a transition the disk never made. The alert threshold is a separate scalar layered over that machine, and its default tracks the warn boundary so the product is no quieter than it is today.
  • Neither preference is an authorization. Raising a goal cannot make a PROTECTED resource deletable, cannot bypass an ASK, and cannot widen an Autopilot envelope. These two numbers decide when the product speaks and where recovery aims; every gate between a candidate and its deletion is elsewhere.
  • --json belongs to both verbs and prints one shape from both: notify_at_used_percent and notify_at_description for the threshold, a default_goal object carrying used_percent, free_percent and description for the goal, a bounds object carrying the four limits above as notify_at_minimum_used_percent, notify_at_maximum_used_percent, goal_minimum_used_percent and goal_maximum_used_percent, stored_at, and loaded_from_file. bounds is there so a settings screen can offer a control that cannot compose a value this command would refuse, without knowing the limits independently and drifting from them; the cross-field rule is deliberately absent from it, because a limit that moves as the other number moves is not a bound on either one. It always describes what is in force after the command ran. loaded_from_file is reported rather than an is_default flag, because a user may legitimately store the default numbers and a client must be able to tell “nothing chosen yet” from “these were chosen”. A refused set prints a rejection object on stdout instead — reason (a stable token), message, and only whichever of notify_at_used_percent/goal_used_percent the refusal actually involved.
  • The file is ~/Library/Application Support/Glomeris/settings.conf, beside the Autopilot envelope and in the same versioned format: a version key, one key = value per line, unknown keys an error rather than noise, and loading routed through the validating constructor so a hand-edited file cannot install a pair the CLI would have refused.

Exit codes: 0 the preferences were printed, or the change was stored; 1 the settings file could not be read or written, or what it contains is not valid; 2 usage error, including set with neither flag, or a value on this command line that was refused.

A refused stored file is 1, not 2: it is not this command line’s mistake, and a script must be able to tell “you typed 120” from “the file on disk disagrees with itself”.

glomeris external-context [--json]

What optional remote context may reach the model, and what never will. Read-only and off by default: with nothing configured it prints that nothing leaves this machine, which is a fact about the shipped product rather than a hypothetical.

Glomeris can optionally read a pull request’s state from GitHub and a work item’s state from Jira, as supporting evidence for ranking a working tree that looks stale. Neither is deletion authority — see the list at the end of this section — and neither is on until you write the configuration file.

This command asks nothing. Answering “what may leave” by leaving is the one shape it must not have: you would be making the very requests you are trying to understand, against a service that logs them, before deciding whether you want that. Everything it prints comes from your configuration file and from the value vocabularies the code itself serializes.

Three kinds of line, and the difference between them is the point:

SectionWhat it describes
ProvidersYour local setup. Whether each provider is configured, whether it is usable, the environment variable its credential is read from, the endpoint, and for GitHub the one repository host in scope.
Fields that may be sentWhat could actually leave, one line per serialized key, each with its complete set of possible values.
Never sentThe explicit negative.

What this page adds:

  • “Not configured” and “configured but not usable” are different rows. A provider whose credential environment variable is unset is reported as configured with a refusal naming the variable, not as absent. The same distinction runs all the way through the feature: a provider that could not answer never reports an absence, and no pull request was found is a different value from GitHub could not be reached.
  • The field list is the complete set, and it is derived rather than written. Each state vocabulary comes from the enum the projection serializes, so a variant added later cannot widen what really travels while this page keeps understating it. A remote state reaches a model as a bounded token — merged, in_progress, none_observed — with how many days ago it was read, and with the kind of failure when a provider could not answer.
  • Nothing identifying goes. Not the repository, the branch, the issue key, the pull request’s number, title, body or author, your account, or any path. The command prints that negative explicitly, because a list of what travels does not answer “did you send my branch name”.
  • Your Jira account address is not echoed, even though the configuration file holds it. It identifies a person rather than a setting anybody debugs, and this report gets pasted into bug reports. The site is shown, because you need to see which one is asked.
  • No credential value appears anywhere, in prose or in JSON. The variable is named; the value is never read by this command at all.
  • Jira has no host scope and does not claim one. Correlation there requires an explicit, deterministic issue key in the branch name — HORO-1234, not a fuzzy match from fix-storage-stuff — so there is no remote to compare against. GitHub’s repository_host is the most privacy-relevant line in the report: a working tree whose remote is anywhere else is never named to that provider at all.
  • None of it is permission. A merged pull request and a closed ticket are context for ranking. They cannot make a working tree deletable: a dirty tree, an untracked file, a local commit made after the merge, or a process holding the directory open each keep it protected or uncertain whatever a remote service says. That is enforced in the policy engine, which never sees external context, and pinned by tests/external_context_grants_no_authority.rs.
  • Read-only by construction, not by convention. The providers can name exactly one transport, whose whole vocabulary is a single GET; scripts/check-external-context-is-read-only.sh proves the module names no mutating verb, no second transport, and nothing that prints or writes.
  • --json prints the same report — each provider’s configured and usable state, its credential variable, endpoint and host scope, every field that may be sent with its vocabulary, and the never_sent list. This is what the menu-bar app’s privacy preview reads.
  • The file is ~/Library/Application Support/Glomeris/external-context.conf, beside settings.conf and in the same versioned format.

Exit codes: 0 the report was printed; 2 usage error — this command takes no arguments besides --json.

An unparseable configuration file is 0, not 1, which is the opposite of settings show. The two commands answer different questions. settings show is asked “what are my preferences”, and an unreadable file means it has no answer. This one is asked “what could leave this machine”, and an unreadable file has a complete and reassuring answer — nothing is configured, so nothing leaves — that a non-zero exit and a bare stderr line would throw away. The refusal is reported as the first line of the report instead.

glomeris pressure <show|notified|respond> [<answer>] [--json]

The seam between the background monitor and the menu-bar app. You are unlikely to type it; the app runs it every poll.

It exists because of a constraint neither side can work around alone: only an app bundle can put buttons on a macOS notification, and the monitor (glomeris daemon) is a bare launchd process, not a bundle. So the monitor records that a notification is owed and the app raises it. This command is how the app asks what is owed and reports which button was pressed.

The subcommand is optional and defaults to show, so a bare glomeris pressure reads rather than writes.

VerbWhat it doesCan it delete?
showPrints the current episode, whether a notification is owed, your alert threshold, your recovery goal, and the answers that can be given.No
notifiedRecords that the notification was put on screen, so it is not raised again for the same episode.No
respondRecords which button was pressed. Takes one of the three answers below.No

What an episode is

One continuous stretch of disk usage being at or above settings --notify-at-used-percent. It opens the first time an observation reaches that threshold and closes only once usage has fallen three percentage points below it. That gap is the hysteresis, and it is the whole reason one spell of a full disk produces one notification rather than one per poll: a volume hovering exactly on the threshold stays inside one episode, and it takes a real recovery rather than a rounding wobble to end it.

Episode ids are monotonic and never reused, because the app uses the id as the notification’s identifier — two notifications about one episode coalesce, and two episodes never do.

The three answers

AnswerWhat it means
review_and_recoverOpen the recovery screen for this episode. Opens it — it does not start recovery, and deletes nothing.
remind_laterSay nothing for two hours, then notify again if usage is still above the threshold. The episode stays open, so this is a delay, not a dismissal.
ignore_episodeSay nothing more about this spell of disk pressure. Monitoring continues and the next episode notifies again.

Hyphens are accepted too (respond review-and-recover), because that is what a command line looks like. Nothing else is: an unrecognized answer is refused rather than mapped onto the nearest match, since the nearest match to a misspelled ignore_episode is the one that opens recovery.

ignore_episode cannot switch notifications off. It is stored on the episode, so it expires with it — there is no answer here, and no flag anywhere, that silences monitoring indefinitely. That is the point of storing the answer on the episode rather than in settings.

What this command does not do

  • It does not observe. show never opens or closes an episode and never decides that a notification is owed; only the monitor’s own poll does. On a machine where the monitor is not installed there is therefore no episode at all, and show says so while still reporting whether usage is above your threshold.
  • It does not delete anything, and it cannot widen what may be deleted. An episode decides when you are spoken to, never what is permitted. Answering review_and_recover opens a screen. Every gate between a candidate and its deletion is elsewhere and unchanged.
  • It does not start recovery, automatically or otherwise. Crossing a threshold raises a question; a human answers it.

--json

Belongs to all three verbs and prints one shape from all three, always describing the state after the command ran. This is what the menu-bar app reads.

current carries the disk reading with both percentages under separate names and the byte figures; notify_at_used_percent/notify_at_description the threshold; a default_goal object the recovery goal, in the same shape settings --json prints it; threshold_crossed and notification_due the two facts the app acts on; episode the episode’s identity, its opened/peak/latest readings, how many notifications have been raised, the answer if one was given, and the snooze deadline if one is pending; responses the answers that can be given, so the app offers buttons it did not invent.

A refusal prints a rejection object on stdout instead — reason as a stable token (no_open_episode, no_notification_due) and message — rather than writing to stderr, because a refusal here is an ordinary outcome that the app has to read and display.

Exit codes: 0 the episode was printed, or the answer was recorded; 1 the episode state or your settings could not be read or written, or filesystem usage could not be measured; 2 usage error, including an answer that is not one of the three; 3 there was nothing to record.

3 rather than 1 for that last case, and this is the load-bearing part of the contract: no episode is open, or no notification was owed. It is not a failure. The usual cause is that the disk recovered while the notification was on screen, and the app polls this command every thirty seconds — a client that treated it as an error would retry at poll speed forever.

Exit codes

Also available as glomeris help exit-codes, which is the copy to trust — it renders from the same table as the per-command help, so a command’s exit codes cannot drift between its own --help and the summary.

  • 0 — success.
  • 1 — a macOS-only command was run on a non-macOS platform, a platform-level operation (e.g. reading the launchd plist path) failed, or (for llm-plan specifically) the LLM provider call/response parsing failed, or --plan-file named an unreadable path, or (for llm-check specifically) the check ran and did not pass.
  • 2 — usage error: unknown top-level command, unknown daemon subcommand, missing/unrecognized free arguments, or (for llm-plan/llm-check specifically) an unrecognized argument, a missing flag value, missing live-mode LLM environment configuration, a base URL that cannot work, or an --api-key/--key/--token flag.
  • 3 — (autopilot run and free --autopilot only) this run was not authorized, so nothing was attempted: either Autopilot is not enabled, or --unattended was given and the grant does not allow starting a run unasked. One code for both because to a script they are the same fact.
  • 75 — (free/emergency/execute/autopilot run only) the HORO-1054 execution lock is already held by another glomeris invocation.

glomeris execute has its own, more specific set of exit codes (0–5 plus 75) — see its own section above for the full table; a couple of those codes (1, 2) overlap this list’s meanings but are worth reading in full since execute is the one subcommand with real destructive consequences.

See glomeris llm-plan’s and glomeris llm-check‘s own sections above for those subcommands’ exit codes in full detail — llm-check’s 1/2 split in particular carries a meaning this list cannot: whether a report exists.

Pressure Model

The monitor module implements pure disk-pressure state logic plus the polling loop that drives it. All platform-specific I/O (real statvfs, real notifications) lives under platform::macos, which the monitor’s traits abstract away.

PressureState

Defined in src/monitor/pressure.rs, in ascending urgency order:

#![allow(unused)]
fn main() {
pub enum PressureState {
    Healthy,
    Warn,
    Pressured,
    Critical,
    Emergency,
}
}

Rendered (Display) as HEALTHY, WARN, PRESSURED, CRITICAL, EMERGENCY.

Thresholds

ThresholdConfig::default() (src/monitor/config.rs) defines both a percentage-used bound and an absolute free-bytes bound per state:

Stateused % ≥free bytes ≤
Warn75.050 GiB
Pressured85.020 GiB
Critical92.010 GiB
Emergency97.03 GiB

A threshold is “crossed” when either bound is crossed (percentage OR absolute bytes) — not both. Classification checks most-urgent-first (Emergency → Critical → Pressured → Warn → Healthy), so one reading that crosses several boundaries lands on the single most urgent one. These are library defaults; ThresholdConfig is passed as a parameter to monitor::run, so a caller could supply different values, but nothing in this codebase does so today.

Debounce / hysteresis

PressureStateMachine::observe() (src/monitor/state_machine.rs) requires confirm_after consecutive matching observations of a candidate state before confirming a transition. The production default is confirm_after = 2 (two consecutive one-minute polls, at the default 60-second poll interval).

  • If the newly observed state equals the currently-confirmed state, any in-flight candidate is discarded.
  • If a different candidate than the pending one appears, its counter resets to 1 — candidates are not averaged.
  • Once the counter reaches confirm_after, the transition confirms and a Transition { from, to } is emitted exactly once.
  • The same logic and the same confirm_after value apply in both directions — there is no asymmetry between escalating and recovering (no “escalate fast, recover slow” rule exists in this code).

This exists specifically to prevent a single noisy sample from producing a notification on every poll.

The polling loop

monitor::run() (src/monitor/poller.rs) starts the state machine assuming Healthy, then loops: stat the watched path, classify the reading, feed it to the state machine, and — only on a confirmed transition — call the notifier and the persistence backend. Sleep happens between iterations. max_iterations: Option<u64> bounds the loop for tests; production passes None to run indefinitely.

Error handling is deliberately non-fatal at every layer:

  • A failed FsStat::stat() call is reported to the on_iteration callback as an outcome-less iteration and does not stop the loop.
  • A notifier failure or a persistence-write failure is captured into the poll outcome’s notify_error/persist_error fields — the loop keeps running either way. This is proven by a test that drives 20 successful loop iterations against an always-failing persistence backend.

Persistence

FilePersistence (src/monitor/persistence.rs) appends one tab-separated line per confirmed transition (unix_time_secs, from, to, used_percent, free_bytes) to a plain text file — not a database. Its own module doc states it is “explicitly failure-tolerant: any error here must never stop the monitor loop from continuing to observe and notify.” No feature in this crate depends on persistence succeeding. Structured (SQLite) persistence is deferred to a later ticket.

What the monitor does not do

The monitor is a cheap, O(1) capacity check (statvfs) — it never walks the directory tree. Discovering what is reclaimable is the scanner’s and detectors’ job (see Architecture), a deliberately separate and more expensive path.

Evidence Model

Evidence (src/evidence/model.rs) describes what a detector observed about a resource — never what should be done about it. That judgment belongs entirely to the policy layer.

Evidence

Key fields:

#![allow(unused)]
fn main() {
pub struct Evidence {
    pub resource: ResourceId,
    pub fingerprint: ResourceFingerprint,
    pub detector: DetectorId,
    pub logical_bytes: ProbeOutcome<u64>,
    pub physical_bytes: Option<u64>,       // always None in the MVP
    pub reclaimable_bytes: ProbeOutcome<u64>,
    pub last_modified: ProbeOutcome<SystemTime>,
    pub last_accessed: ProbeOutcome<SystemTime>,
    pub regenerability: Regenerability,
    pub recoverability: Recoverability,
    pub native_cleanup: NativeCleanup,
    pub open_by_process: ProbeOutcome<Vec<ProcessRef>>,
    pub process_cwd_match: ProbeOutcome<Vec<ProcessRef>>,
    pub git_state: ProbeOutcome<Option<GitState>>,
    pub tool_liveness: ProbeOutcome<bool>,
    pub collected_at: SystemTime,
    pub sources: Vec<String>,
}
}

physical_bytes is always None today — subtree st_blocks summation is out of scope for the current milestone. git_state’s Observed(None) means the probe ran successfully and determined the resource genuinely isn’t in a git working tree — a complete, legitimate answer, not a missing one; only Unavailable(reason) means the probe itself failed.

ProbeOutcome<T> — the core honesty guard

#![allow(unused)]
fn main() {
pub enum ProbeOutcome<T> {
    Observed(T),
    Unavailable(ProbeReason),
}

pub enum ProbeReason {
    ToolAbsent,
    ToolNotRunning,
    PermissionDenied,
    TimedOut,
    Failed,
    NotAttempted,
}
}

This is deliberately not Result<T, E>: there is no Default impl and no unwrap_or_default() escape hatch. A caller that wants the observed value must explicitly branch on this enum — there is no way to silently coerce an unavailable probe into a zero/empty/false value. This is the structural guard against “a failed probe becomes a safe default.”

Completeness and Confidence

#![allow(unused)]
fn main() {
pub enum Completeness {
    Complete,
    Partial { missing: Vec<EvidenceField> },
    Failed,
}

pub enum Confidence {
    High,
    Medium,
    Low,
}
}

Evidence::completeness() checks every field ResourceKind::required_evidence() lists for that resource’s kind: Complete if none are missing, Failed if all required fields are missing, Partial otherwise. Confidence is derived from completeness (Complete → High, Partial with ≤1 field missing → Medium, everything else → Low) and is advisory only — the policy layer does not consume it; it exists to give a human or LLM something short to point at.

Resource kinds whose owning tool has no persistent daemon to check (Cargo/Npm/Pnpm/Yarn/Homebrew) do not require ToolLiveness for Complete, since that field would otherwise be structurally unreachable — see Known Limitations for a related caveat about Docker build cache never reaching Complete at all.

ProbeOutcome never masquerades as “safe”

Nothing in this codebase treats missing or failed evidence as evidence of “nothing to clean up.” A detector’s own status type (DetectorStatus::Failed(reason)) is documented as meaning “we don’t know,” never “nothing found,” and ResourceKind::Unknown is a deliberate fail-closed sink that the policy layer maps to PROTECTED unconditionally rather than falling through to any default treated as safe. The actual “therefore refuse to delete” enforcement lives in policy::classify, not in this module — this module only guarantees the data can’t fabricate a misleadingly complete picture.

Correlation: open_by_process, process_cwd_match, git_state, tool_liveness

Detectors (detectors/) only ever populate the discovery-stage fields (size, mtime, regenerability). The four correlation fields above are always Unavailable(NotAttempted) straight out of a detector — a separate collector, evidence::correlate::DefaultEvidenceCollector, is responsible for actually attempting correlation. This split is what keeps freshly discovered evidence from ever reporting Completeness::Complete before correlation has actually run.

DefaultEvidenceCollector wires in the real, subprocess-backed probes:

FieldBacked byCommand
open_by_processLsofOpenFileProbelsof -F pcn +D <path>
process_cwd_matchLsofProcessCwdProbelsof -a -d cwd -F pcn +D <path>
git_stateGitCliProbegit -C <path> rev-parse --show-toplevel, then git status --porcelain and git rev-parse --git-dir --git-common-dir on the resolved repo root
tool_livenessPgrepToolLivenessProbepgrep -x <daemon-name> (only for tools with a real daemon: Xcode.app, Docker’s com.docker.backend; other tools skip the subprocess entirely and report Unavailable(ToolNotRunning))

Each probe runs under a shared per-call timeout (ProbeBudget) enforced by polling try_wait() and killing the child on timeout to avoid leaving a zombie process. A collect() call can spend up to roughly four times that timeout in the worst case, since up to four subprocesses each get their own full budget.

Absence and failure handling is per-tool, not global: a missing lsof, git, or pgrep binary degrades that specific field to Unavailable(ProbeReason::ToolAbsent) — it does not fail the whole correlation pass, and it is never silently treated as “nothing found.” A non-zero exit with no output (e.g. lsof finding nothing) is treated as a legitimate empty result; a non-zero exit with stderr content is treated as a genuine probe failure (Unavailable(ProbeReason::Failed)).

Docker’s correlation gap

DockerBuildCache/DockerImageCache resources use ResourceLocator::Tool (a tool-native id, not a filesystem path) because there is no single canonical path to a Docker build cache. DefaultEvidenceCollector can only run tool_liveness for a Tool-locator resource; open_by_process, process_cwd_match, and git_state are always Unavailable(NotAttempted) for it. Since required_evidence() still requires all three, Docker build cache cannot reach Completeness::Complete today — see Known Limitations.

Safety Model

The three-class decision

policy::classify() (src/policy/engine.rs) is the single deterministic module that decides AUTO_SAFE / ASK / PROTECTED for a resource. It is pure: no I/O, no ambient clock (now is always a parameter) — which is what makes fail-closed behavior testable and lets the executor call the exact same function again on freshly re-collected evidence at deletion time.

#![allow(unused)]
fn main() {
pub enum PolicyClass {
    AutoSafe,
    Ask,
    Protected,
}
}

There is no separate “Unknown” fourth outcome. Missing, stale, or failed evidence maps inside classify to Ask or Protected with a specific reason code — it never becomes a state a caller might misread as “not Protected, therefore fine.”

classify()’s fail-closed ordering

classify checks conditions in this exact order, and the order is itself part of the safety guarantee:

  1. Unknown resource kind → unconditionally Protected. Never falls through to any evidence-based judgment.
  2. Path/pattern-based protected check (see below) → Protected, checked before any freshness/completeness logic. A Protected classification never depends on evidence freshness — a stale-but-protected resource is still Protected, never “upgraded” by fresher evidence.
  3. Staleness — evidence older than PolicyConfig::max_evidence_age (default 5 minutes), or whose collected_at is somehow in the future → Ask + EvidenceStale.
  4. Completeness — Failed → Ask + EvidenceProbeFailed; Partial → Ask + EvidenceIncomplete.
  5. Active-use signals (most-significant first) — an open file handle or matching process cwd → ResourceInActiveUse; a dirty or linked git worktree → GitWorktreeDirty; a live owning-tool daemon → OwningToolLive. Any of these → Ask.
  6. Per-instance regenerability — NotRegenerable → Ask + RebuildCostHigh, regardless of how clean the rest of the evidence looks.
  7. Otherwise → AutoSafe, with reasons EvidenceFreshAndComplete (+ RegenerableByTool if applicable) + NoActiveUseObserved.

The PROTECTED matcher

policy::protected::protected_reason() is an explicitly conservative, MVP-scope, non-exhaustive denylist, not a claim of completeness. It performs no filesystem I/O (no canonicalize, no read_link, no metadata) — it only inspects the path/locator text already carried by the resource, which is what keeps classify pure. What it actually matches today:

  • Docker image cache — every DockerImageCache resource is treated as a possible persistent volume and classified Protected unconditionally, because detectors do not yet distinguish a persistent volume from disposable image cache. (DockerBuildCache is not covered by this rule.)
  • Credential material — any path component .ssh or .gnupg, or a filename ending .pem/.key, or containing credentials.
  • Git internals — any path with a .git path component (not merely “inside a git-tracked project,” which is the separate, evidence-driven GitWorktreeDirty reason).
  • Infra state — a .terraform path component, or a filename ending .tfstate/.tfstate.backup.
  • System paths — /System, /usr (except /usr/local), /bin, /sbin, /private/var/db.
  • Unsafe mount/symlink targets — /Volumes, /dev, /Network, /net.

This is a starting denylist to extend, not a guarantee that every dangerous path is covered.

policy::approval::authorize() is the only way to construct an Approval — the type an executor is required to hold before acting. It is built with a private, unconstructible-outside-the-module marker type, so nothing outside approval.rs can fabricate one.

  • AutoSafe → always authorizes, no consent needed.
  • Protected → never authorizes, unconditionally, with no override parameter that can change that.
  • Ask → authorizes only if a matching UserConsent is supplied: the consent’s resource identity and fingerprint must match exactly. Consent granted for one resource instance does not carry over to a different instance at the same path, or to the same resource after its underlying fingerprint has changed.

There is no TTY prompt, but consent is no longer unreachable. Three surfaces supply a UserConsent today, and none of them is a prompt:

  • glomeris execute --confirm-ask --observed-fingerprint <token> — you pass back the exact fingerprint_token a prior explain --json printed on that same resource. Mismatched or missing, it is a usage error rather than a silent approval.
  • The menu-bar app’s Clean button, which builds precisely those two flags (CandidateDetailView.buildExecuteArguments) and nothing else. The GUI is the confirmation step; the consent still travels as a fingerprint.
  • glomeris autopilot enable --preauthorize-ask <kind>:<reason> — narrow advance consent for one Ask reason on one resource kind, recorded in the envelope. It is a flag on enable, not on run: the consent is written down before the run and autopilot run rejects the flag outright, so the run cannot grant itself anything the stored envelope does not already say. See Autopilot, and the limitation on that path in Known Limitations.

What has not changed is glomeris free’s recovery loop: it wires auto_approve_ask: false (src/main.rs), so inside that loop Ask candidates are still reported as declined/skipped and never executed. That is the loop’s own choice, not an absence of machinery.

Deletion-time TOCTOU revalidation

executor::execute() never trusts a previously computed PolicyDecision at face value. Immediately before mutating anything, it:

  1. Snapshots the resource’s live filesystem identity via symlink_metadata (never metadata, so a symlink is detected as a symlink, never resolved through) — captured before the slower correlation re-probe runs, so a symlink swap during that window can’t backdate the anchor.
  2. Re-collects evidence from scratch (re-probes size/mtime/fingerprint directly, re-runs the same EvidenceCollector).
  3. Aborts (ResourceIdentityChanged) if the fresh fingerprint doesn’t match the one the approval was granted against.
  4. Re-runs classify() on the fresh evidence.
  5. Aborts (PolicyClassDowngraded) if the class no longer matches what was approved.
  6. Aborts (PolicyReasonsWidened) if any new reason appears that wasn’t present at approval time — Ask is a heterogeneous bucket, so consent granted against RebuildCostHigh does not cover a freshly observed ResourceInActiveUse.
  7. Aborts (EvidenceDegraded) if completeness got worse since approval — a defensive, forward-looking guard.
  8. Only then builds the actual plan from the same fresh evidence just validated, and runs it. Immediately before any RunTool/DeletePath step actually mutates the filesystem, the resource’s identity is re-verified one more time against the early snapshot from step 1.

A plan with more than one step is refused outright rather than executed, because partial-deletion byte accounting has no way to report a correct total if a later step fails after an earlier one already succeeded. No registered action emits more than one step today.

AUTO_SAFE is end-to-end reachable through real execution

Detectors populate reclaimable_bytes at discovery time (HORO-992), and the deletion-time revalidation path (executor::build_fresh_evidence) reuses the same bounded recursive size estimate (HORO-1016) for reclaimable_bytes that it already used for logical_bytes (HORO-994) — it no longer hardcodes the field back to Unavailable. A real AutoSafe approval built from a detector’s evidence genuinely survives revalidation and executes for real, proven by tests/golden_chain_execute.rs. See Known Limitations for what’s still out of scope (Docker build cache’s own completeness gap).

Actions never receive raw commands

Every registered Action produces a typed ActionPlan/ActionStep from Evidence — never a caller-supplied string. ActionStep::RunTool invokes one of a closed set of binaries (ToolBinary::{Brew,Cargo,Npm,Pnpm,Yarn}) via argument-array Command::new(tool).args(args) calls — never a shell string. ActionPlan/ActionStep deliberately never derive Deserialize, so no external input (network payload, LLM text) can ever materialize one directly; the only way external input reaches execution is by selecting a pre-registered ActionId, whose real Action::plan implementation then decides the actual steps from Evidence it independently trusts. See BYOK LLM Planner for how this applies to the optional LLM path specifically.

An envelope narrows this model; it never widens it

glomeris autopilot adds the one thing the rest of this page does not describe: a grant that outlives the moment you typed it. It changes nothing above. An Autopilot candidate is classified by the same classify(), authorized by the same policy::approval::authorize, and revalidated by the same deletion-time TOCTOU check, in that order.

What the envelope adds is a filter in front of authorization, not an alternative to it. Its allowlist can only remove kinds from consideration; PROTECTED is refused whatever it says; UNKNOWN_INCOMPLETE is refused whatever it says; an ASK reason it has not been given by name is refused. The budgets — actions, bytes, wall clock, and an optional disk-pressure floor — bound a run’s total effect, so the worst case of an unattended run is a number you read before you granted it.

The clause this adds to the invariant is the last one: AI can recommend. Policy decides. Executor verifies. Filesystem reality wins — and the envelope bounds the outcome. See Autopilot.

Emergency Mode

glomeris emergency (src/emergency/mod.rs, macOS only) is a degraded recovery path that must produce a useful result even when SQLite/history/log writes fail, there is no network, there is no LLM provider configured, and no GUI is running. It is not glomeris free --target with a flag — its scope and limitations are deliberately narrower.

The menu-bar app does not expose this command, deliberately: emergency takes no arguments and acts machine-wide, which is not a thing to put behind a one-click control. It is CLI-only.

The invariant is unchanged

Emergency pressure is not permission to weaken the safety model. This module reuses policy::classify/policy::approval::authorize and executor::execute exactly as they are elsewhere — it does not define any separate, looser emergency-only rule set, and it never references anything network- or LLM-related (enforced by a test that scans the module for such references).

What it does, in order

  1. Frees its own disposable state first, unconditionally. Before anything else, it deletes the tool’s own best-effort pressure-history file (the same history.tsv the daemon appends to) — never gated on policy, because this file is not a developer resource under policy’s purview; it’s this tool’s own append-only log, and deleting it is safe by construction. A missing file is not an error: the function silently does nothing rather than fabricating work.
  2. Discovers candidates via the same bounded detector registry used elsewhere (DetectorRegistry::discover_all) — never the scanner’s slower, unbounded filesystem walk.
  3. Classifies and, for AutoSafe candidates only, attempts execution through the same classify/authorize/execute pipeline as normal operation.

AUTO_SAFE-only, no interactive Ask

A degraded, no-time, no-mechanism-for-consent path can only auto-execute AutoSafe candidates. Everything else (Ask, Protected) is refused and counted in EmergencyReport::denied_candidates — never escalated to interactive consent. This is a deliberate MVP scope decision documented in the module itself, not an oversight.

No LLM, no network, no required database

run_emergency never calls an LLM and never makes a network request — this is enforced by a dedicated test, not just a comment. Persistence-write failures (a plain write failure, or a failing directory creation) do not stop the run: run_emergency still completes and still frees the self-owned-state fixture, per its own fault-injection tests.

EmergencyReport

Every field is a plain, dependency-free primitive (u32/u64/String), so the report stays printable even if every other subsystem in the process has already failed:

  • actions_attempted / actions_succeeded
  • total_bytes_freed
  • denied_candidates — every candidate whose classify() result was not AutoSafe. Never incremented for an unrelated wiring gap (e.g. no registered action for an AutoSafe kind) — those go to errors instead.
  • errors — non-fatal error strings, bounded to at most 8 entries so an adversarial run can’t grow this field without bound.

Documented limitations (from the module’s own doc comments)

  • The preallocated emergency-reserve-file idea is deferred, not implemented. The originating ticket named it as an experimental idea to reject or defer without reliable measured evidence; this module does not build it.
  • Resolved (HORO-994): the candidate loop now frees real bytes via a real detector-produced candidate. executor::build_fresh_evidence’s deletion-time revalidation reuses the same bounded recursive size estimate (HORO-1016) for reclaimable_bytes that it already used for logical_bytes, so a real AutoSafe approval survives revalidation and executes. Both the self-owned-disposable-state step (step 1 above) and real detector-found candidates can now reclaim bytes.
  • Near-zero-real-disk-space testing was not performed. Driving a real machine’s free space to near zero to test this path was judged too destructive to be worth the risk; fault injection via fakes covers the equivalent failure modes instead.

BYOK LLM Planner

actions::llm (src/actions/llm.rs) is an optional, “bring your own key” LLM planner. As of HORO-1008, it is wired into the glomeris binary as an advisory, non-executing subcommand: glomeris llm-plan. This page documents both the library module and the CLI surface over it.

Since HORO-1308 the menu-bar app has an AI Plan card over the same subcommand. It is a thin client: it spawns glomeris llm-plan --json --progress-json, renders the report, and has no planner, no provider client and no execution path of its own. Everything on this page — what leaves your machine, what the validator drops, what the planner can and cannot do — applies unchanged to the GUI, because it is the same code doing the work. See Menu Bar App for how the card separates the model’s words from the machine’s.

Since HORO-1309 the provider itself can be configured from that app (Settings → AI Provider) instead of only from the shell, with the key in the login keychain rather than an exported variable — see Configuring this from the menu-bar app below. The GUI is still a thin client there too: it writes three values and launches a child process with them. It does not validate the URL, does not know which models exist, and never speaks to a provider itself.

What it is

An advisory ranking suggestion over evidence the crate already collected. LlmResourceView is an explicit, bounded projection of one Evidence record — deliberately not a Serialize derive on Evidence itself, so adding a field to Evidence later has no effect on what an LLM would see unless a human explicitly adds it here too. plan_with_llm is the entry point: it serializes a set of these views, calls a provider, and defensively validates the response.

What leaves your machine

Two strings: a fixed system prompt and a JSON array of LlmResourceView values. LlmRequestPayload is that request, and it is the whole outbound surface — nothing about the local machine is added downstream of it.

Each view identifies its resource by a positional wire alias — resource_1, resource_2, … — not by its real ResourceId. That matters because a ResourceId for a path-backed resource kind renders as an absolute path, and an absolute path under $HOME carries your OS account name and your directory layout. Since HORO-1298, no path, home directory or account name is transmitted; what the model gets is the resource’s kind, reclaimable size, age, regenerability, evidence completeness and the action ids offered for it, which is what a ranking judgement is actually made from.

The alias is positional rather than a hash of the resource. A hash would be stable across runs, which sounds better until you count the keyspace: a path like /Users/<name>/Library/Developer/Xcode/DerivedData has one unknown segment, so anyone holding a candidate account-name list can confirm it against the hash offline. A stable pseudonym is also, by construction, a handle for correlating your machine across requests. An index is neither, and costs nothing — the model is ranking resources it was just handed, not recognizing them from last week.

The wire-id → real-ResourceId table stays in the process’s memory and is never serialized. It is how a response’s resource_id regains meaning locally.

Checking this yourself

glomeris llm-plan --print-payload [--json]

Runs discovery, prints the exact request a live run would send plus the local alias table, and stops — before any provider is constructed. No API key is needed, and no network call is made, so the command cannot send the thing it is showing you. The two prompt sections are what leaves; the alias section is explicitly labelled as not sent, and is the only place the real absolute paths appear.

What it is not, and never will be

plan_with_llm’s output is a ranking suggestion only. It never calls policy::classify/authorize, and it must not: every surviving (ResourceId, ActionId, priority) tuple still has to go through the exact same policy classification any other candidate would, in the caller, before anything executes. A dedicated test (llm_plan_item_never_bypasses_policy) demonstrates this concretely: constructing evidence for SSH key material, getting an LLM response that “approves” deleting it, confirming the item survives plan_with_llm’s validation — and then showing policy::classify still returns Protected for it regardless.

glomeris llm-plan (the CLI layer built on top of this module) keeps the same guarantee: crate::cli::build_llm_plan_report never constructs a policy::Approval and never calls policy::approval::authorize or executor::execute. Nothing this subcommand prints is ever executed — there is no --execute/--yes flag, and there never will be one on this subcommand.

glomeris llm-plan usage

glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json]
glomeris llm-plan --print-payload [--project-root <path>]... [--json]
glomeris llm-plan --schema
  • Without --plan-file, calls a real OpenAI-compatible endpoint via actions::llm::provider_from_env.
  • --plan-file <path> reads the file’s raw bytes and treats them exactly as if they were the model’s raw response text, through the identical extract_plan/LlmPlan/plan_with_llm validation pipeline — useful for reproducing a scenario deterministically, with no network call and no API key.
  • --print-payload prints the outbound request without sending it, and returns before a provider is constructed — see “What leaves your machine” above. Requires no API key.
  • --json prints the LlmPlanReport (or, with --print-payload, the LlmPayloadReport) as JSON instead of the human-readable form.

Human-readable output always opens with:

LLM SUGGESTION — advisory only, nothing is executed by this command

See CLI Reference for the full flag/exit-code table.

glomeris llm-check — does the configuration actually work?

glomeris llm-check [--json]

The narrowest command in the CLI: no project roots, no detectors, no evidence, no policy, no actions. It sends two fixed prompts — "You are a connection test. Reply with the single word: ok." and "ok" — through the same LlmProvider::complete a plan uses, and reports whether the endpoint, the credential and the model name work.

Two properties are worth being explicit about, because a connection test that lacked either would be worse than none:

  • It proves the real path. A cheaper probe — a GET /models, a HEAD request, a DNS lookup — can pass while a POST /chat/completions with your model name fails. This sends the request a plan sends, minus the evidence.
  • It describes nothing about this machine. The two prompts are constants with no interpolation, and connection_test_prompts_describe_nothing_local pins that they contain no path separator, no home directory and no account name — an assertion that is only meaningful because they are constants. So the cost of testing a configuration is one trivial completion, not a disclosure.

The outcome is one of five tokens, and they are deliberately distinct because the fix for each is in a different place:

outcomeExitWhat it meansWhat to change
ok0The provider answered; its reply is in response_excerpt—
rejected1A non-2xx status — the provider answered and said noThe credential, or the base-URL path: read the path in error
unreachable1No HTTP response at all — DNS, TLS, refused, timeoutNetwork, host name, VPN
unusable_response1A 2xx that was empty, not JSON, or missing choices[0].message.contentThe model name, or a gateway not really speaking the OpenAI shape
misconfigured2validate_base_url refused before anything was sentThe base URL — see above

The 1/2 split carries a meaning the outcome token alone does not: on 1 a full report was printed first, so the diagnosis is on stdout; on 2 there is no report at all, and the single line on stderr prefixed glomeris llm-check: is the whole explanation. misconfigured is therefore never a token you will read out of a report — InvalidConfiguration can only come from constructing the provider, which happens before a report exists. It is in the vocabulary because it is the right name for the condition, and a caller branching on exit 2 applies it itself rather than inventing a sixth word. That is what lets the menu-bar app’s connection test branch without parsing prose, and it is also why a missing configuration is 2 rather than 1 — nothing was attempted, and the fix is entirely local.

actions::llm::llm_check_outcome is the one producer of those five tokens, and scripts/check-vocabulary-covers-cli-tokens.sh diffs the set of tokens the menu-bar app has wording for against the set that function can return, in CI — so the CLI and the GUI cannot drift into describing different sets of failures, and adding a sixth failure mode in Rust fails a check rather than silently reaching a user as a blank explanation.

Configuring this from the menu-bar app

Settings → AI Provider (HORO-1309) configures the same three values from the GUI, for the common case of someone who runs the app from Finder and never exported anything. A Finder-launched LSUIElement agent inherits no shell environment at all, so before this screen existed the GUI’s AI Plan card could only work for someone who had launched the app from a configured shell.

FieldStored in
Endpoint (API root)UserDefaults.standard — for a bundled app, the preference domain named by its bundle identifier
Modelthe same
API keythe login keychain, account llmApiKey, service = the app’s bundle identifier

Both namespaces are derived from the running bundle rather than written out as literals (HORO-1456), so a beta, renamed or diagnostic build gets its own stored settings and its own keychain items instead of the released app’s.

The precedence rule

Per field, what this app has configured wins over what it inherited. A field left empty here falls back to the inherited GLOMERIS_LLM_* variable, including falling back to it being absent. Nothing is ever removed — clearing a GUI field returns you to the environment rather than switching a working setup off.

Stated as a rule: the settings screen never lies. If the endpoint field shows a URL, that is the URL used. The alternative — environment-wins — was rejected precisely because it lets the screen display one endpoint while silently using another, with no way to correct it from the GUI.

Each field’s row says where its effective value came from (Configured here, From the environment, or Not set), so an inherited value is visible rather than mysterious. An empty-or-whitespace value counts as absent on both sides, by the same rule the Rust provider_from_parts uses, so an exported-but-empty GLOMERIS_LLM_MODEL= does not masquerade as configuration.

The CLI is unaffected either way (AC 3). glomeris in a terminal reads that terminal’s environment and knows nothing about this store. Nothing in this screen changes, shadows or requires anything for a shell user.

Where the key goes, and where it does not

Into the environment of the glomeris child process, and nowhere else. It is read from the keychain at the moment a command is spawned and not cached.

It is never a command-line argument (visible to ps, and refused by the CLI by flag name), never UserDefaults, never a file the app writes, never a log line, and never rendered — the field is a SecureField, there is no reveal control, and the typed text is cleared the moment it is handed to the keychain, whether the write succeeded or not. GlomerisLlmSettingsStore deliberately has no getter for the key; the only code that reads it is the one function that puts it straight into a child environment, so “show the user where their key came from” and “put the key on screen” cannot become the same operation. scripts/check-credential-store-uses-keychain.sh asserts that in CI over every line in the app that touches the key.

Removing it is explicit (AC 8): a Remove key button, worded as the destructive action it is, which deletes the keychain item. Afterwards the field falls back to the environment if one is exported, and reads Not set if not.

Note what that fallback means, because the precedence rule cuts the other way here. If GLOMERIS_LLM_API_KEY was exported into the environment this app was launched from, deleting the stored key does not stop Glomeris reaching a provider — it promotes the inherited key into use. Someone pressing a button labelled Remove key is more likely revoking access than tidying a field, so the app says so in that case rather than reporting a bare “Key removed”: it names the variable and tells you to unset it too. Unsetting it means relaunching the app from an environment without it — a running process’s environment is not editable from the settings screen.

What protects the stored key, and which programs can read it

Honest version, because an earlier one of these notes overstated it (HORO-1455).

The key is a generic-password item in your login keychain, which is a file keychain — ~/Library/Keychains/login.keychain-db. That means:

  • It is encrypted at rest, while the keychain is locked. Normally your login keychain is unlocked for your whole session, so for practical purposes the protection during that session is the access control list below, not encryption.
  • It is not synced to iCloud. True — a file keychain cannot sync at all.
  • It is included in a file-level backup of your home directory: Time Machine, a cloned disk, an rsync of ~. The item is still encrypted there and is useless without your keychain password, but it is not excluded from backups. The app sets the ThisDeviceOnly accessibility attribute, and that attribute is inert on a file keychain — it is honoured by macOS’s data-protection keychain, which needs an entitlement an unsigned build does not have. It is set so the value is already correct if Glomeris is ever signed and entitled.

Which programs can read it. The item carries an access control list naming the binaries allowed to decrypt it. Glomeris creates the item trusting exactly one program: the copy of the app that created it. Any other program asking for the value — including a different build of Glomeris, such as one you just upgraded to — does not silently get it. macOS asks you, naming the program.

You can see and edit that list yourself: open Keychain Access, find the item whose name matches the app’s bundle identifier, and look at Access Control. Nothing about this is only inspectable from inside Glomeris.

If you press “Always Allow” on that prompt, the program you allowed is added to the list permanently. Before HORO-1455 the list only ever grew: every build you upgraded past stayed on it, so over time the key became readable by every old Glomeris binary still on the disk.

How to reset the list. Save the key again — paste it into the API key field and save. Glomeris now removes the old item and creates a fresh one rather than updating in place, so the trusted list is rebuilt from one entry and every previously allowed program is dropped. (This is also why the list cannot be narrowed without re-saving: macOS does not permit an item’s access control list to be rewritten in place.) If you would rather start from nothing, Remove key deletes the item outright — read the note above first about what that promotes.

Two consequences worth stating plainly:

  • An upgrade will still prompt you once. Resetting the list is not the same as avoiding the prompt; a new build is not on the old item’s list, and macOS asks before letting it read or delete the item. What changed is that answering once no longer leaves a permanent grant behind.
  • If you decline that prompt, the save fails and your previously stored key is left exactly as it was. macOS refuses the deletion and the item stays put, so nothing is lost — the app reports that the save did not happen rather than quietly falling back to an in-place update, because that fallback is what let the list grow. Try again and allow it.
  • If the prompt is allowed but the re-add then fails, there is no stored key and the field falls back to the environment, as if you had removed it. That is the one regression this change accepts: previously a failed write left the old key in place. It is the safer direction — the cost is pasting the key again, against a key that stayed readable by binaries you had stopped trusting.

The privacy preview

The same screen offers a preview of the outbound payload, over glomeris llm-plan --print-payload --json with the app’s configured project roots — the same scope a real plan would use, so the preview is of your payload and not a generic example. It separates what leaves this Mac (the two prompts) from what stays on this Mac (the wire-alias table, which is where the absolute paths are).

Opening it sends nothing, and cannot: --print-payload returns before a provider is constructed (AC 7). The preview runs without the credential in its environment at all — only the connection test is given it — so there is no version of this screen in which looking at the payload transmits it.

LlmPlan schema — the concrete --plan-file example (HORO-1048)

glomeris llm-plan --schema prints an example, syntactically valid LlmPlan JSON document to stdout — the canonical way to get a starting point for a --plan-file fixture without reading src/actions/llm.rs’s LlmPlanItem struct directly. Its resource_id is in the ResourceId::to_string() form because a human writing a fixture by hand names resources the way glomeris detect prints them; a live model is handed wire aliases instead, and answers with those:

glomeris llm-plan --schema
{
  "items": [
    {
      "resource_id": "cargo_target_dir:/path/to/project/target",
      "action_id": "cargo.clean.target_dir",
      "priority": 1,
      "reason": "stale build artifacts, not modified in 30 days"
    }
  ]
}

Field meanings, all on LlmPlanItem:

FieldTypeMeaning
resource_idString, requiredMust match either a positional wire alias (resource_1, …) from the request the model was given, or — for a hand-written --plan-file fixture — a ResourceId::to_string() from the evidence set plan_with_llm was called with. Both are looked up in sets this process built from its own discovery, so neither can name a resource that was not found locally. An unmatched value drops just that item (dropped_unknown_resource), never the whole plan.
action_idString, requiredMust match a registered ActionRegistry action id — an unmatched value drops just that item (dropped_unknown_action).
priorityOption<u32>Informational ranking hint only.
reasonOption<String>Human-readable explanation. Never interpreted as an instruction, a path, or anything that reaches execution.

#[serde(deny_unknown_fields)] on LlmPlanItem means any extra field (e.g. a smuggled "command") fails deserialization of the whole document, not just that item — see “Safety properties” below.

This example is not a fixed, hand-maintained fixture: --schema’s output is proven, by a real subprocess round-trip test (tests/llm_plan_schema_round_trip.rs), to be accepted unchanged by glomeris llm-plan --plan-file <path> — i.e. it parses and validates through the exact pipeline above without a parse-error exit code. (The resource_id in the shipped example is deliberately one no real discovery run will ever produce, so a round trip against a real evidence set still drops it as an unknown resource — that is an expected, non-error validation outcome, not a parse failure.)

Configuration

Live mode (no --plan-file) reads three environment variables, all required, with no default base_url:

VariablePurpose
GLOMERIS_LLM_API_KEYBearer token sent as Authorization: Bearer <key>
GLOMERIS_LLM_BASE_URLAPI root of an OpenAI-compatible service — see below
GLOMERIS_LLM_MODELModel name sent in the request body

GLOMERIS_LLM_BASE_URL is the API root, not the host root

Glomeris appends /chat/completions to the base URL verbatim. It never inserts a /v1 segment for you, and never strips one. A single trailing slash is tolerated. So the base URL must be the path prefix your provider serves its API under — for most providers that includes /v1:

export GLOMERIS_LLM_BASE_URL="https://gateway.example.com/v1"
# request goes to https://gateway.example.com/v1/chat/completions

Setting the host root instead is the most common misconfiguration:

export GLOMERIS_LLM_BASE_URL="https://gateway.example.com"
# request goes to https://gateway.example.com/chat/completions  ← wrong path

Because Glomeris does not rewrite the path, a base URL that already ends in /v1 produces exactly one /v1 segment — there is no /v1/v1 failure mode.

What is refused locally, before anything is sent

actions::llm::validate_base_url runs when the provider is constructed, so a base URL that cannot possibly work fails with LlmError::InvalidConfiguration and exit 2 — no request, no round trip to blame. The refusals, each with the message naming what to change:

RefusedWhy
EmptyNothing to route to; there is deliberately no default base URL
Leading or trailing whitespaceA pasted URL routinely carries a trailing newline, and the resulting failure is a 404 nobody can explain by looking at the field
An interior spaceA URL cannot contain one
No scheme, or a scheme other than https:///http://Nothing else can be posted to
No host (https://, https:///v1)Same
A query string or a fragmentAppending /chat/completions to …/v1?key=… puts the path inside the query string, and the diagnostic path would then read /v1 and misdescribe why. Refusing also keeps a credential a user pasted into the URL out of the request Glomeris builds
The full endpoint URL (ends in /chat/completions)Would produce /v1/chat/completions/chat/completions — the other half of the API-root mistake above

The message never quotes the offending value, on purpose: of the three settings, the one a user most plausibly pastes into the wrong field is the credential.

What is not refused is the host-root form above. It is a perfectly valid URL that some providers really do serve their API at, so only the provider can say whether it works — which is what glomeris llm-check is for.

Do not try to identify this mistake from the status code: a gateway may answer an unrouted path with 404, but 403, 401 and even 400 are all things real gateways return instead. The path in the error message is the path that was actually requested, so it tells you directly which of the two forms you configured — and since a path of exactly /chat/completions can only have come from a base URL with no path at all, Glomeris says so itself rather than leaving you to notice:

provider returned HTTP 403 for openai:chat_completions POST /chat/completions:
{"error":"..."} — the configured address has no path, so this request went to
the host root; most OpenAI-compatible providers serve their API under /v1, so a
missing /v1 is the likeliest cause

(One line in reality; wrapped here to fit.)

The sentence comes after the provider’s own words, never instead of them, and it is a hint rather than a verdict: the host-root form is valid and some providers really do serve their API there, so this cannot be a refusal. A rejection at a configured API root — …/v1/chat/completions — gets no such sentence at all, because a 401 there is about the key.

A full configuration, using placeholders throughout:

export GLOMERIS_LLM_API_KEY="<token>"          # never commit; never pass in argv
export GLOMERIS_LLM_BASE_URL="https://gateway.example.com/v1"
export GLOMERIS_LLM_MODEL="example-model"
glomeris llm-plan --project-root ~/code/my-project

Prefer export in your shell (or a secret manager that exports into the process environment) over a plaintext .env file, and never pass the key on the command line — glomeris llm-plan rejects --api-key/--key/--token outright for that reason.

Diagnosing a provider failure

Provider failures are reported in the provider_error field (and on the text output’s provider error: line) as one readable, secret-free sentence. The classes are deliberately distinct, because each has its fix in a different place — glomeris llm-check reports the same distinction as its five-token outcome:

ClassMeaning
LLM provider configuration is invalid: …Refused locally; no request was attempted at all
provider unreachable: …No HTTP response at all — DNS, TLS, refused connection, timeout
provider returned HTTP <status> …The provider answered with a non-2xx status
provider response was unusable: …A 2xx response that was empty, not JSON, or missing choices[0].message.content

An HTTP failure names the status, the API style, the request path, the provider’s x-request-id when it sends one, and a bounded excerpt of the provider’s own error body:

provider returned HTTP 401 for openai:chat_completions POST /v1/chat/completions \
  (request_id=req-abc-123): {"error":{"code":"invalid_api_key"}}

What is deliberately not in that message: the Authorization header, the API key, the request payload, the scheme, the host, and the query string. The host is omitted because it may be private infrastructure and the query string because some gateways accept a credential there; the path alone is what diagnoses a base-URL mistake. The body excerpt is truncated to a bounded length and is scrubbed of the configured API key, so a provider that echoes the credential it just rejected cannot turn Glomeris’s diagnostics into the leak.

When that path is exactly /chat/completions, the message also names the host-root cause described above — the one reading of the path that needs no knowledge of the host to make.

actions::llm::provider_from_env is the only place std::env::var is called for these — the key is held just long enough to build the OpenAiCompatibleProvider and is never logged, printed, or returned any other way. A missing or empty value for any of the three returns LlmError::NotConfigured; glomeris llm-plan and glomeris llm-check report this by naming the three variable names, never a value, and exit 2.

A value that is present but cannot work returns LlmError::InvalidConfiguration instead, and also exits 2 (HORO-1309). Keeping the two apart matters: “you have not set this up” and “you set it up wrongly, here is what to change” have different fixes, and reporting the second as the first sent users to re-export three variables that were already set.

OpenAiCompatibleProvider is the one real provider implementation — any OpenAI-compatible /chat/completions endpoint (OpenAI itself, OpenRouter, a self-hosted gateway):

#![allow(unused)]
fn main() {
pub struct OpenAiCompatibleProvider {
    pub base_url: String,
    pub api_key: String,
    pub model: String,
}
}

The API key is never accepted as a CLI argument. glomeris llm-plan explicitly rejects --api-key/--key/--token with an error pointing at $GLOMERIS_LLM_API_KEY instead — a key on the command line would be visible to ps and land in shell history.

Never commit an API key to source control, a Jira ticket, a commit message, or a PR description. Treat it as you would any other credential.

Safety properties that are already verified

  • OpenAiCompatibleProvider deliberately does not derive Debug — a derived Debug would print api_key verbatim. A test confirms a real failed complete() call’s error output never contains the key.
  • LlmError’s variants never embed the key either — every variant’s Display output is asserted key-free by api_key_never_appears_in_display_output_of_any_variant, including the ProviderStatus body excerpt, which actively scrubs the configured key (redact_secret_removes_every_occurrence_and_collapses_whitespace) so a provider echoing the credential back cannot leak it through Glomeris.
  • The diagnostic endpoint path in a provider error is the path only: two tests (diagnostic_endpoint_path_drops_scheme_host_and_query, diagnostic_endpoint_path_never_contains_the_host) assert the scheme, host, and query string are all dropped.
  • The wire protocol itself is pinned against a loopback mock HTTP server rather than mocked at the trait: the appended path, trailing-slash tolerance, Authorization: Bearer scheme, model and both message roles, and the 401/403/404/empty-body/non-JSON/connection-closed failure modes each have a test that inspects the bytes actually sent.
  • No absolute path, home directory or account name is transmitted. Asserted at three levels: on the payload struct (payload_carries_no_path_or_account_name), on the CLI report (build_llm_payload_report_separates_outbound_prompts_from_local_aliases), and on the bytes of the real HTTP request, captured by a loopback listener standing in for the provider (tests/llm_plan_egress_privacy.rs) — headers included. That last test also pins the two halves that make the proof non-vacuous: the payload demonstrably does describe a path-backed resource, and the real path was available locally and was withheld deliberately.
  • Wire aliases are positional, so the same request shape carries byte-identical ids regardless of what the resources are (wire_ids_are_positional_and_independent_of_the_resource) — there is no stable identifier in the payload that could correlate a machine across requests.
  • An alias resolves only through the request’s own table, and a real-id string only through the evidence set: a model that invents either — an out-of-range resource_99, or a guessed path it was never shown — names nothing (out_of_range_wire_alias_is_dropped_and_counted, model_supplied_path_cannot_name_an_undiscovered_resource). No model-provided identifier gains execution authority; policy still classifies every surviving item.
  • The serialized view’s key set is pinned to its seven documented fields (serialized_view_exposes_exactly_the_documented_fields), so adding a field to LlmResourceView — the only way to widen what leaves — fails a test rather than passing silently.
  • LlmPlanItem uses #[serde(deny_unknown_fields)]: a model trying to smuggle an extra field (e.g. a "command") fails deserialization of the whole plan, not just that item — confirmed by unexpected_field_rejects_whole_plan.
  • An unknown resource_id or action_id in the model’s response drops just that one item (counted in dropped_unknown_resource/ dropped_unknown_action); it never fails or invalidates the rest of the plan.
  • LlmPlan/LlmPlanItem remain the only two #[derive(Deserialize)] types in the entire crate — LlmPlanItemReport/LlmPlanReport (the CLI report DTOs) are Serialize only. Everything a model response can produce is a String resolved against real, already-in-memory data, or a plain u32/Option<String> used only for display/ranking — never a path, never a shell fragment, never anything that reaches ActionStep construction directly.
  • The response parser handles a raw JSON object, a fenced ```json block, or a bare fenced block, and either produces a well-formed LlmPlan or nothing — never a partial parse.
  • A PROTECTED candidate is refused before its action is ever resolved: crate::cli::build_llm_plan_report checks PolicyClass::Protected before calling ActionRegistry::get/executor::dry_run for that item. tests/golden_llm_plan_protected_refusal.rs proves this end to end through the exact --plan-file input surface the CLI uses, closing the gap Known Limitations previously described as “verified at the code level only, not end-to-end through the CLI.”
  • The connection test’s two prompts are constants, and connection_test_prompts_describe_nothing_local asserts they contain no path separator, no home directory and no account name — so testing a configuration costs one trivial completion and discloses nothing.
  • The GUI’s provider settings are held to the same line by macos/GlomerisMenuBar/Tests/AiProviderPreferencesViewTests.swift, which asserts several things as absences, since they are code paths that must not exist: the key is bound to a SecureField and never read back or revealed; the typed text is cleared before the success-or-failure branch, so a failed keychain write cannot leave it in memory behind a visible error; neither command can run without the user pressing something (no .onAppear, no .task, no Timer, no .refreshable); exactly one call site is given the credential, and it is not the payload preview; and the preview’s outcome enum has no “not configured” case at all — a preview that could say that would be a preview that had tried to use a provider. Both llm-check reports are additionally pinned by golden fixtures on both sides (tests/dto_golden_fixtures.rs, DtoGoldenFixturesTests.swift), asserting no base_url, no api_key and no :// reaches a report a UI shows and a log keeps.
  • The GUI half of AI Plan is held to the same line by macos/GlomerisMenuBar/Tests/AiPlanSectionViewTests.swift: a row built from a confident, plausible recommendation to delete a credential carries the refusal and no willingness to act, and the card has no execution call site, no fingerprint token and no re-sorting of the model’s order. The repo-wide scripts/check-no-policy-label-branching.sh additionally proves nothing in the app — this card included — branches on a policy_label string to decide what may run.

Known limitations

  • Only one provider shape is implemented — no Anthropic-native, Azure-OpenAI-specific auth, or other provider shape.
  • No retry/backoff on transient network failures.
  • No streaming support — complete() returns the full response text at once.
  • No integration into any existing rule-only ranking loop (glomeris free) — deliberately out of scope for this ticket. A future ticket may wire it in as an optional enhancement, falling back to rule-only ranking whenever the provider call fails.
  • glomeris llm-plan has no --execute/--yes flag and never will on this subcommand — it is advisory-only by design, not an MVP gap.

Autopilot

Autopilot is a standing grant with limits, written down where you can read it.

Every other deleting command in Glomeris is something you type at the moment you want it: execute names one action on one resource, free runs a bounded recovery loop you asked for, emergency is the button you press when the disk is nearly full. Autopilot is the one that can act without you present — so the whole of its design is about bounding what “without you present” is allowed to mean.

The invariant the rest of this book states as

AI can recommend. Policy decides. Executor verifies. Filesystem reality wins.

gains one clause here:

…and the envelope bounds the outcome.

The envelope

An AutopilotEnvelope is the complete statement of what Autopilot may do. It is a small file of key = value lines at ~/Library/Application Support/Glomeris/autopilot.conf, and nothing else grants authority — not an environment variable, not the plan you hand run, and not the menu-bar app’s Autopilot tab, which grants by running autopilot enable and keeps no second copy of the answer. There is one grant and it is that file.

FieldDefaultHard ceilingWhat it bounds
enabledrevoked—Whether run may execute anything at all.
allowed kindsnone—Which ResourceKinds a candidate may be. An empty allowlist reaches nothing.
max actions325How many actions one run may perform.
max bytes5 GiB64 GiBThe byte total one run may reclaim.
max duration60s900sWall-clock budget for one run.
min pressurenone—A disk-pressure floor below which the run refuses.
pre-authorized ASKnone—Exactly which (kind, reason) pairs an ASK classification may proceed on.
respond to alertsoff—Whether a run may begin without being asked. The one row that is not a ceiling on outcomes — see Runs nobody asked for.

Two properties of that table matter more than the numbers in it:

The defaults grant nothing. AutopilotEnvelope::default() and ::revoked() are the same value, and a missing file reads as revoked rather than as an error. Autopilot being off is not a setting someone has to remember to choose; it is what the absence of a decision means.

A limit above its ceiling is refused, not clamped. Asking for 500 actions is a usage error (exit 2) and writes no envelope. Clamping would be the worse behaviour by a wide margin: you would have been told 500, be living under 25, and have no way to notice the difference.

Reading the envelope is read-only in the strict sense — glomeris autopilot with no arguments prints the grant and does not create the file it just read.

What the envelope cannot authorize

PROTECTED is unconditional. No flag on autopilot enable can reach it, and no ordering in a plan file can make a PROTECTED candidate executable. The allowlist is a narrowing filter over what policy already permits, never a widening one.

UNKNOWN_INCOMPLETE is likewise unreachable. It is the reporting label for an ASK decision whose evidence is incomplete, stale, or came from a probe that failed — so --preauthorize-ask refuses every evidence-quality reason. Pre-authorizing “I accept the rebuild cost” is a judgment a person can make in advance. “I accept that we do not know what this is” is not.

--preauthorize-ask takes one kind:reason pair at a time, e.g. node_modules:rebuild_cost_high. There is no wildcard, because the only purpose of a wildcard here would be to turn a narrow consent into a blanket one.

What a model may do, exactly

autopilot run --plan-file <path> reads an LlmPlan — the same schema BYOK LLM Planner documents, and produced by the same llm-plan command — and uses it to order the candidates this machine already discovered. That is its entire authority.

It cannot:

  • add a candidate (a resource ID the local scan did not produce is dropped);
  • choose an action (ValidatedPlanItem::action_id is deliberately ignored by the run; the action comes from the local registry, by kind);
  • supply a path, a command, or a command fragment;
  • raise any limit in the envelope;
  • change any policy class;
  • invent an action that is not registered.

The argument that this is airtight is structural rather than defensive: ordering is a permutation of a list, and reordering a list cannot add a member to it. The plan is consumed as a sort key over locally-derived candidates, so the set of things that can be deleted is fixed before the plan is read. A regression test hands the run a plan whose reason field is "; rm -rf /Users/dev/company-repo && echo pwned" and asserts, among other things, that the string never reaches the audit log — but the reason that test passes is that model_reason is never read by anything but the record writer.

There is deliberately no live-provider mode. autopilot run accepts a plan file and nothing else, so no deletion in this product waits on a network call. Asking a model is llm-plan’s job, and the two are separate commands so that the deleting one has no network path at all.

What one run actually does

For each candidate, in the order the plan (or the local ranking) put them:

  1. Admission gate — a pure function over the envelope, the candidate’s policy decision, the observed pressure and the run’s budget ledger. It returns either admission or a RefusalReason that names which bound stopped it: kind not allowed, policy class, unauthorized ASK reason, pressure floor, action budget, byte budget, time budget.
  2. Authorization — the normal policy::approval::authorize path, which is the only way an Approval can exist. A pre-authorized ASK supplies a real UserConsent matched to the resource and fingerprint; nothing forges one.
  3. Execution — the normal executor::execute, including deletion-time TOCTOU revalidation. An identity change, a class downgrade, widened reasons or degraded evidence aborts the action having deleted nothing. See Safety Model.
  4. Ledger — the attempt is charged exactly once, at max(expected, actual) bytes, so an action that reclaimed more than predicted cannot under-charge the budget.

The action budget stops the run; a byte-budget refusal does not. A candidate too large for the remaining byte budget is refused and the run continues to the next one, because the next one may well fit — whereas an exhausted action count is exhausted for everything.

A run that finds nothing it is allowed to do exits 0. “Allowed nothing” is a successful outcome, not an error.

The other consumer of the grant: free --autopilot

The envelope is not autopilot run’s private property. glomeris free --autopilot bounds the goal-driven recovery loop by the same stored grant, using the same admission gate above as one more narrowing stage between selection and authorization — so a goal-driven run can never reach further than a grant allows, and the grant remains the single place that authority is written down.

Two consequences worth stating, because they are what makes the second consumer safe rather than merely convenient:

  • The grant cannot widen anything. It has no power to add a candidate, choose an action, supply a path or raise a limit — only to withhold. Everything an --autopilot run does, the same command without the flag would also have done.
  • A stop the grant caused is reported as such, never as an exhausted disk: stop_reason: "envelope_refused" plus the refusal token. revoke therefore stops a free --autopilot run exactly as it stops autopilot run — both re-read the file, and neither has a cached copy.

See CLI Reference for the flag’s exit codes and its refusal alongside --dry-run.

Runs nobody asked for

Everything above bounds what a run may do. Whether a run may begin when nobody pressed anything — the standing answer to a disk-pressure alert — is a separate question, and the envelope answers it separately: autopilot enable --respond-to-alerts, off by default, exercised by free --autopilot --unattended.

It is two consents, not one. An enabled envelope does not authorize an unprompted run; --unattended without that permission exits 3 having attempted nothing. Reading “enabled” as consent to act unattended would widen every envelope already on disk — each written by somebody bounding a run they meant to start — into authority to delete while they were away. That is not a grant those people gave, so it is not one this reads out of their file.

It is the only field here that is not a ceiling on outcomes, which is why the table above says so rather than glossing it. It is still narrowing in the direction that matters: off, the default, means every run has a human behind it, and turning it on changes nothing about what a run may do once it starts. The budgets, the allowlist and the pre-authorizations all still apply, unchanged.

The check lives in Rust, twice over. The grant exposes starts_unprompted — in force and set — as one value, so that no client ever composes its own permission out of two fields; and free itself refuses an --unattended run the grant does not cover, so the check cannot be skipped by whatever launched the process. --unattended is the caller’s statement about itself, and demanding it is what makes a run nobody started distinguishable from one somebody typed.

revoke stops it with everything else, and keeps the setting on file — so revoking never reads as having silently cleared what you chose.

The audit trail

Every executed attempt appends one JSON line to ~/Library/Application Support/Glomeris/actions.jsonl, readable with glomeris history. Autopilot’s records carry two fields the interactive paths do not populate:

  • source — autopilot_auto_safe or autopilot_preauthorized_ask, so the authority an action ran under is recoverable from the log rather than inferred. Interactive records read execute, free or emergency.
  • model_rank — where the plan ranked that candidate, or null when no plan was involved. This is how you tell “the model suggested it and policy allowed it” from “policy allowed it and no model was consulted.” It is 1-based, and it is the same number the run’s own report printed as [AI rank N] — a log offset by one from the report it came from would be worse than no log.

“Where the plan ranked it” means its position in the plan’s items array, not its priority field. priority is advisory and its direction was never specified anywhere — nothing says whether 1 means most urgent or least — so ordering on it would mean inventing a semantics and then depending on it.

Refusals are not written to actions.jsonl — it is a log of what was done to the filesystem, and a refusal did nothing to the filesystem. They are in the AutopilotReport that run prints, which is where you look to find out why a run was quiet.

Revocation

glomeris autopilot revoke takes effect on the next run, and there is nothing to restart: every run re-reads the file. Revocation keeps the limits, so a later enable cannot come back carrying limits you never read.

enable replaces the previous envelope rather than merging into it, for the same reason. A grant should be one line you can read out loud, never the accumulated union of every enable you have ever typed.

Nothing here expires on its own. The envelope is a file: it survives quitting, restarting and logging out, and it stays in force until it is revoked. That is a deliberate omission rather than a missing feature — a grant that lapsed on a timer would mean the honest answer to “what may Autopilot do right now” depended on the clock, and show would have to be read differently depending on when you read it.

Granting it without a terminal

Everything above can also be read, granted and revoked in the menu-bar app, under Settings → Autopilot (Cmd+,). Reading a standing deletion grant only from --help and a config file would have meant that in practice most people who enabled it had never read it, so the GUI is part of the feature rather than a convenience on top of it.

It changes nothing about where authority lives. The tab renders autopilot show --json — the same envelope, the same ceilings, the same lists of what can never be allowlisted, pre-authorized or executed — and writes by running autopilot enable or autopilot revoke. It holds no policy logic, and it offers no choice that did not arrive from the CLI as data, which is the project-wide rule for that app and is checked in CI rather than trusted.

Two things it deliberately does not do: it cannot start a run (autopilot run has no --json for the same reason), and it cannot construct a grant the CLI would reject. See Menu Bar App for the screen itself, including why its form is pre-filled from the envelope in force and why its byte budget rounds up.

Known limitation: a pre-authorized ASK cannot complete today

A pre-authorized ASK candidate is admitted, authorized with real consent, and then always aborts inside deletion-time revalidation with PolicyClassDowngraded. Nothing is deleted.

The cause is upstream of Autopilot. ASK/RebuildCostHigh arises only from a per-instance Regenerability::NotRegenerable, while executor::build_fresh_evidence rebuilds regenerability from the resource kind’s static default — so the fresh classification lands on AutoSafe and the class comparison trips.

This is left as-is deliberately. The failure direction is the safe one (refuse, mutate nothing), and “make build_fresh_evidence carry a per-instance judgment” changes the TOCTOU anchor every deleting command shares. It is asserted by a test named for it — a_preauthorized_ask_still_aborts_at_deletion_time_revalidation — rather than left to be discovered, so whoever fixes it upstream starts from a failing test that points at the cause.

See Known Limitations for the rest.

Daemon Lifecycle

glomeris daemon <install|uninstall|status|run> manages an optional background poller via macOS’s launchd, implemented in src/platform/macos/launchd.rs.

User-level, not root

Everything here runs at the per-user level: the plist is written to ~/Library/LaunchAgents/com.glomeris.monitor.plist, and it is loaded with plain launchctl load/unload (gui/<uid> session semantics). Nothing writes to a system-level LaunchDaemons location, and no root/sudo access is required for any of these commands.

glomeris daemon install [--force]

Writes a plist that runs <program_path> daemon run every 60 seconds (StartInterval) by default, with RunAtLoad set, and stdout/stderr redirected to ~/Library/Logs/Glomeris/monitor.log/.err.log. Then runs launchctl load -w <plist_path>. If launchctl fails, the plist file is left in place (so status/a manual retry can still see it) — only the final launchctl result is treated as a hard failure.

What this owns, preserves, and refuses to clobber (HORO-1021): the plist at ~/Library/LaunchAgents/com.glomeris.monitor.plist is a wholly Glomeris-owned artifact — no other tool reads or writes it. Even so, install never silently discards a hand-edited value:

  • Idempotent. Re-running install against an already-up-to-date plist (matching a fresh default install) writes nothing.
  • Preserves a hand-edited StartInterval. If you’ve changed the poll interval by hand, re-running install keeps your value — only the program path is refreshed (the actual point of reinstalling after rebuilding/moving the binary).
  • Refuses any other unrecognized customization without --force. If the on-disk plist has been changed in a way this tool doesn’t manage (a hand-added key, a flipped RunAtLoad, etc.), install makes zero changes and exits with status 2, explaining the refusal. Pass --force to overwrite anyway — a backup of the current file is taken first, at <plist_path>.plist.bak.
  • Atomic, verified writes. Every real write goes to a temp file in the same directory, is renamed into place, and is read back and compared before install reports success — a crash or concurrent install mid-write can’t leave a corrupt/partial plist.
  • Aborts on concurrent modification. If the plist changes on disk between install’s read and its write (e.g. two daemon installs racing), the later write aborts with zero mutation of the live plist and exit status 2 rather than risk clobbering the concurrent change — rerun to retry. This is checked twice: once before anything is written, and again right before the live plist itself is overwritten (after the backup copy has been taken) — narrowing the window where a concurrent writer could land and be silently clobbered down to the instant between that second check and the write itself. That last sliver can’t be closed without holding a lock across the whole read-check-write section, which is out of scope here. If the second check is the one that catches a race, a .plist.bak backup of the pre-race content may already be on disk; that stray backup is harmless (never the concurrent writer’s data) and is not treated as a mutation of the live plist for the purposes of this guarantee.

glomeris daemon uninstall

Runs launchctl unload -w <plist_path> (best-effort — an already-unloaded agent reporting an error from launchctl is not treated as fatal here), backs up the current plist to <plist_path>.plist.bak, then removes the plist file if present. Idempotent — uninstalling an already-missing plist is not an error.

glomeris daemon status

Prints whether the plist file exists, its path, and whether launchctl list <label> currently reports the job loaded. This never errors: an unreachable launchctl (e.g. a non-macOS or sandboxed CI environment) is reported as loaded: false, not surfaced as a failure.

glomeris daemon run

Runs the polling loop in the foreground — this is the exact command the installed launchd agent invokes on its own schedule. It watches /, using ThresholdConfig::default() (see Pressure Model), the real macOS statvfs-backed FsStat, and the real osascript-backed notifier, appending confirmed pressure transitions to ~/Library/Application Support/Glomeris/history.tsv.

Known CI limitation

Only plist generation and path/file logic are unit tested in this repository’s CI. Actually asking launchd to load/run the agent needs a real macOS user session and is not exercised by cargo test.

Architecture

Module map

From src/lib.rs, the crate’s pub mod declarations:

#![allow(unused)]
fn main() {
pub mod actions;
pub mod detectors;
pub mod emergency;
pub mod evidence;
pub mod executor;
pub mod monitor;
pub mod platform;
pub mod policy;
pub mod scanner;
}

platform::macos is itself gated #[cfg(target_os = "macos")] — it does not exist in a build on any other OS. src/main.rs is a thin CLI shim over this library; the pure logic in monitor/scanner and their unit tests do not depend on being reachable from a binary at all.

ModuleResponsibility
monitorDisk-pressure state machine and the polling loop that drives it. Cheap, O(1) capacity checks only (statvfs) — never walks a directory tree. See Pressure Model.
scannerA generic, bounded, streaming filesystem walker that reports the top-K largest entries by size. Unrelated to the detectors below — it has no notion of resource kind, regenerability, or policy.
detectorsBounded, per-tool probes (Cargo, Homebrew, Node, Xcode, Docker) of known roots, with a bounded recursive size estimate under each root — not a full filesystem walk. Produces discovery-stage Evidence with the four correlation fields always Unavailable(NotAttempted).
evidenceThe Evidence/ProbeOutcome/Completeness domain model (evidence::model, evidence::probe), plus the correlate submodule that fills in the four correlation fields via lsof/git/pgrep. See Evidence Model.
policyThe single deterministic classify()/authorize() pair that decides AUTO_SAFE/ASK/PROTECTED. See Safety Model.
actionsTyped, pre-registered cleanup actions (ActionRegistry) that turn Evidence into a closed ActionPlan/ActionStep — never a caller-supplied string. Also home to the optional BYOK LLM planner (actions::llm).
executorDry-run, real execution with deletion-time TOCTOU revalidation, and the bounded closed-loop recovery orchestration (executor::recovery_loop, glomeris free).
emergencyThe degraded, AUTO_SAFE-only, no-LLM/no-network recovery path (glomeris emergency).
platform::macosThe only place real OS-level I/O happens: statvfs (disk stats), osascript (notifications), launchd (daemon lifecycle).

Data flow

Observe (monitor)
    -> statvfs, classify pressure, debounce, notify/persist (best-effort)

Discover (scanner | detectors)
    -> scanner: generic top-K-by-size walk (no resource semantics)
    -> detectors: bounded per-tool probes -> discovery-stage Evidence
       (correlation fields still Unavailable(NotAttempted))

Correlate (evidence::correlate)
    -> lsof / git / pgrep, per-field timeout and failure handling
    -> fills open_by_process, process_cwd_match, git_state, tool_liveness

Decide (policy::classify / policy::approval::authorize)
    -> AUTO_SAFE | ASK | PROTECTED, fail-closed ordering
    -> Approval is the only thing an executor may act on

Plan (actions::Action::plan)
    -> typed ActionPlan/ActionStep from Evidence
    -> optional: actions::llm ranks candidates, never authorizes them

Execute & revalidate (executor::execute)
    -> re-collect evidence, re-run classify, abort on any disagreement
    -> re-verify filesystem identity immediately before mutating
    -> re-measure actual reclaimed bytes after the fact

executor::recovery_loop (the orchestration behind glomeris free) and emergency are both thin callers that wire the above stages together, one candidate at a time, without reimplementing or loosening any of them — neither introduces its own execution or authorization primitive.

Why the scanner and the detectors are separate

The scanner (scanner::walker) is a generic, dependency-free (std-only), non-recursive directory walker with no concept of what it finds — it reports size and depth only. The detectors are the opposite: narrow, tool-aware probes of specific known locations (e.g. ~/Library/Developer/Xcode/DerivedData, brew --cache’s reported path, <project>/target) that produce policy-relevant Evidence. glomeris scan exercises the former; glomeris detect, glomeris free, and glomeris emergency all exercise the latter.

Troubleshooting

External tools Glomeris shells out to

ToolUsed byPurposeIf absent
dockerdetectors::dockerdocker system df --format '{{json .}}' to report build/image cache sizeToolAbsent — normal, expected, not an error
brewdetectors::homebrewbrew --cache to find Homebrew’s cache pathToolAbsent — normal, expected
lsofevidence::correlate::open_files, ::processFind processes with a resource open, or with it as their cwdThat correlation field becomes Unavailable(ToolAbsent) for the affected resource; other fields are unaffected
gitevidence::correlate::gitDetermine repo root, dirty/untracked state, worktree-nessgit_state becomes Unavailable(ToolAbsent) for the affected resource
pgrepevidence::correlate::tool_livenessCheck whether Xcode.app or Docker’s backend daemon is runningtool_liveness becomes Unavailable(ToolAbsent) for the affected resource
launchctlplatform::macos::launchdLoad/unload/query the background daemondaemon status reports loaded: false; install/uninstall report an error but the plist file itself is still written/removed
osascriptplatform::macos::notifyPost a macOS notification on a confirmed pressure transitionNotification failure is captured as notify_error on the poll outcome; the monitor loop keeps running

A missing tool is treated as a normal, expected state everywhere in this codebase, not an error — every detector’s own doc comment says so explicitly (e.g. “not every detector’s tool is installed on every machine”). The one nuance: for Docker specifically, a running-but-unreachable daemon (e.g. Docker Desktop not started) is also folded into ToolAbsent, not a separate Failed state.

Cargo, Node, and Xcode detectors never shell out to cargo/node/npm/ xcodebuild at all — they only check the filesystem (known_project_roots for a target/node_modules directory, or ~/Library/Developer/Xcode/DerivedData). Their ToolAbsent really means “expected resource not present” (no project roots configured, or the directory doesn’t exist), not “binary missing from PATH.”

“Why does glomeris detect show tool_absent for something I have installed?”

For Cargo/Node, tool_absent also appears if no known_project_roots were configured for the DiscoveryContext used — these two detectors never search the filesystem on their own; they only check specific roots handed to them. Check how the caller (CLI/daemon) constructed the DiscoveryContext.

“Why did a probe come back Unavailable(Failed) instead of ToolAbsent?”

ProbeReason::Failed means the tool ran but something about its output or exit status wasn’t a recognized “nothing found” or “tool absent” shape — for example lsof exiting non-zero with stderr content, or brew --cache succeeding but reporting a path this process can’t canonicalize. This is deliberately never coerced into a safe default; treat it the same as “we don’t know,” not “nothing to clean up.”

“Part of the scan failed — where do I see which detector?”

A failed detector and one whose tool is absent both contribute zero candidates, so a count alone cannot tell them apart while they mean opposite things: “we don’t know what is there” versus “there is nothing there.” Every surface therefore reports the outcome and not just the count.

SurfaceWhere a failure appears
glomeris detectWith no candidates at all, the line reads no candidates discovered by the detectors that succeeded — N failed, so this is not a clean bill of health rather than the bare no candidates discovered
glomeris detect --jsonA detectors array — one entry per detector, in registration order, with status (found/tool_absent/failed), candidates_found, and a reason present only on failed — plus a derived discovery_complete
glomeris detect --progress-jsonEach detector_finished event carries outcome, and reason when it failed
glomeris free, glomeris emergencyA discovery incomplete: N detector(s) failed block naming each one; if the run stopped at SafeExhausted, it also says in so many words that this is not a finding that nothing safe is left
Menu-bar app“Nothing found where Glomeris could look” in place of the all-clear, or “This list may be incomplete” below a non-empty list — either way naming the checks that did not finish

An absent tool appears as tool_absent and does not make discovery_complete false. That is a real answer, not a missing one, and flagging it would make the caveat permanent on any machine without Docker — which is the fastest way to teach people to ignore it.

“The menu-bar app says the glomeris CLI was not found, but I installed it”

The app lists the locations it searched in the same message. If your binary is not in one of them, that is the mismatch — move or symlink it into /opt/homebrew/bin or /usr/local/bin, or launch the app from a shell whose PATH contains its directory. A GUI app launched from Finder gets only PATH=/usr/bin:/bin:/usr/sbin:/sbin, so a PATH that works in your terminal is not visible to the app. See Menu Bar App for the full resolution order.

You do not need to restart the app after installing the CLI: it re-resolves on every invocation, so the next poll picks it up.

macOS permissions

Glomeris’s own I/O runs as the invoking user; it does not request or use Full Disk Access, and nothing in this codebase currently prompts for or checks any macOS privacy permission (TCC). If a probe or detector needs to read a location gated by such a permission on your system, expect a PermissionDenied-flavored Failed/Unavailable outcome rather than a silent empty result — check the affected detector/probe’s specific error message.

glomeris free/glomeris emergency doesn’t actually delete anything

This is expected today, not a bug you need to work around — see Known Limitations for exactly why, and Safety Model for the revalidation logic responsible.

Non-macOS platforms

daemon, emergency, and free print an error to stderr and exit 1 on any OS other than macOS, because platform::macos (real statvfs, osascript, launchd) does not exist in that build at all. scan and detect are not gated this way and should work cross-platform, though the project’s CI only exercises macos-14.

Security & Privacy

No telemetry, no cloud requirement

Glomeris has no backend, no account system, and no telemetry. Every command documented in this book runs entirely locally. The only network access anywhere in the codebase is to whatever BYOK endpoint you configure yourself, from exactly two commands — glomeris llm-plan (the planner) and glomeris llm-check (the connection test, which sends two fixed words and nothing about this machine). See BYOK LLM Planner. Nothing else in the crate makes a network request.

No raw filesystem inventory, and no filesystem paths, sent to any LLM

As of HORO-1008, glomeris llm-plan is wired into the binary (see BYOK LLM Planner). No filesystem data is sent anywhere unless you explicitly run that subcommand without --plan-file and have all three provider settings configured — every other command in this book either makes no network request at all, or (glomeris llm-check) sends two fixed words that describe nothing local. Since HORO-1309 those three settings may come from the menu-bar app’s Settings rather than from GLOMERIS_LLM_* in a shell, which changes where they are stored and nothing about what is sent: the GUI’s AI Plan card spawns the same subcommand with the same bounded payload. Even then, what is sent is bounded and explicit: LlmResourceView is a hand-maintained projection of one Evidence record containing a resource’s kind, size estimate, age, regenerability, completeness, and the action ids offered for it — never raw file contents, never a directory listing, never anything beyond that fixed set of fields. Adding a field to Evidence later has no effect on what a model sees unless a human explicitly adds it to LlmResourceView too, and adding a field to LlmResourceView itself fails a test that pins its serialized key set.

Nor are resource paths sent. Until HORO-1298 each view identified its resource by its real ResourceId, which for the five path-backed resource kinds renders as an absolute path — and an absolute path under $HOME discloses the OS account name and the machine’s directory layout. Note that --project-root never bounded this: it scopes the cargo and node detectors only, while the Xcode detector is $HOME-bounded and the Homebrew and Docker detectors shell out and are bounded by neither. Each view now carries a positional wire alias — resource_1, resource_2, … — and the table mapping an alias back to a real ResourceId has no Serialize derive and stays in the process’s memory. The aliases are positional rather than hashed on purpose: a hashed path would be both brute-forceable (one unknown segment in /Users/<name>/Library/...) and a stable handle for correlating your machine across requests.

Run glomeris llm-plan --print-payload to see the exact request a live run would send, including the local alias table, without sending it and without configuring a credential. Three tests enforce the property, the strongest of them (tests/llm_plan_egress_privacy.rs) by capturing the real HTTP request with a loopback listener and asserting the transmitted bytes contain no path separator at all.

BYOK secret handling

  • OpenAiCompatibleProvider deliberately does not derive Debug, so an accidental {:?}-log of the provider value cannot leak the API key.
  • Verified by test: neither LlmError’s Debug output nor a real failed complete() call’s error output contains the key.
  • glomeris llm-plan and glomeris llm-check read the key only from GLOMERIS_LLM_API_KEY via actions::llm::provider_from_env — std::env::var for this value is called nowhere else in the crate. The key is never accepted as a CLI flag: --api-key/--key/--token are explicitly rejected with an error pointing at the environment variable instead, so a key never appears in ps output or shell history. The rejection message names the flag only, never the value beside it, so --api-key=<secret> does not get echoed back. See the BYOK page for the full configuration contract.
  • In the menu-bar app (HORO-1309) the key lives in the login keychain, not in UserDefaults, not in a file the app writes, and not in a log line. It is read at the moment a glomeris child process is spawned, placed in that child’s environment, and not cached. GlomerisLlmSettingsStore deliberately exposes no getter for it — the only code that can read the value is the one function that builds a child environment — and the settings screen binds it to a SecureField with no reveal control and no read-back. Removing it is an explicit destructive control, not a side effect of clearing a field. scripts/check-credential-store-uses-keychain.sh enforces all of that in CI over every line in the app that touches the key, so a future shortcut — stashing it in a plist “just for now”, printing it in a debug view — fails a check rather than shipping.

Subprocess invocation

Every external tool Glomeris shells out to — docker, brew, lsof, git, pgrep, launchctl, cargo, npm/pnpm/yarn, brew (as an action) — is invoked via std::process::Command::new(program).args(args) with a Vec<String>/&[&str] argument array. There is no sh -c anywhere in this codebase, and no place where untrusted content is interpolated into a shell command string.

The one place a string is built and passed to an external tool is platform::macos::notify’s osascript invocation, which constructs an AppleScript source snippet via format!. That string’s only two substituted values are the notification title and body produced by monitor::notifier::notification_text, which only ever emits fixed template text derived from the closed PressureState enum — never external or attacker-controlled input — and the whole script is still passed to osascript as a single argument in an argument array, never through sh -c.

No arbitrary LLM-composed commands reach execution

ActionPlan/ActionStep deliberately never implement Deserialize, so no externally-sourced input (a parsed LLM response, a network payload) can ever construct one directly. The only way external input reaches execution is by selecting a pre-registered ActionId string, which a real Action::plan implementation then interprets against Evidence this process already collected and trusts independently. See Safety Model for how this interacts with policy classification.

Bounded, printable failure reporting

EmergencyReport caps the number of retained error strings at 8, and no type in this crate’s execution/reporting path derives Serialize in a way that would let an unbounded or adversarial input grow a report without limit.

Known Limitations

This page collects specific, verified limitations pulled from the actual code and from the PRs that introduced each piece — not a generic disclaimer. This is an experimental MVP; treat every claim elsewhere in this book as scoped by what’s on this page.

Resolved: real detector-produced candidates now complete real deletions

Fixed in HORO-994. Detectors populate Evidence::reclaimable_bytes at discovery time (HORO-992) — Xcode/Cargo/Node/Homebrew via a real size estimate, Docker via docker system df’s own reported figure — and executor::execute()’s deletion-time TOCTOU revalidation (executor::build_fresh_evidence) now reuses the same size-estimate computation for reclaimable_bytes that it already used for logical_bytes, rather than hardcoding it back to Unavailable. A golden end-to-end integration test (tests/golden_chain_execute.rs) proves the full chain — real detector → AutoSafe classification → Approval → execute() → real deletion → real re-measured freed bytes — against a disposable fixture. glomeris free --target and glomeris emergency can both now actually free real bytes, not just report a dry-run plan.

Resolved (HORO-1016): the size estimate was non-recursive and badly under-counted nested trees

Before this fix, detectors::shallow_logical_bytes (the shared size probe behind every detector’s logical_bytes/reclaimable_bytes and behind executor::build_fresh_evidence’s revalidation) summed only a directory’s immediate entries — for a subdirectory entry it counted the directory inode’s own size, never its contents. On a real machine this reported a 4.8 GB Xcode DerivedData tree as 34.9 KB, a 1.0 GB Homebrew cache as 1.3 MB, and a 16 GB Cargo target/ directory as 4.6 KB. Fixed: detectors::estimate_logical_bytes walks the full subtree via an explicit stack (never real recursion, so an arbitrarily deep tree cannot overflow the stack), bounded by a shared 200,000-entry / 750ms budget (detectors::size_estimate_budget) used identically at discovery time and at deletion-time revalidation. Budget exhaustion always yields a truthful partial-sum lower bound — never Unavailable — so a truncated walk can never look like a probe failure to Evidence::completeness() or trip a spurious abort in executor::execute’s TOCTOU revalidation; when the walk does stop early, a provenance note is attached via Evidence::push_source (advisory only, never a policy input). tests/reclaimable_bytes_reaches_auto_safe.rs and the new executor::tests::estimate_matches_total_size_best_effort_on_the_same_tree lock the estimate’s byte semantics to executor::total_size_best_effort’s existing convention (files and symlinks counted by their own size, directories contribute 0). On a real developer machine the 750ms deadline truncates the Xcode DerivedData walk at roughly 42,000 entries, so that resource’s reported bytes are normally a lower bound, by design — no extrapolation is attempted.

Docker build cache can never reach Completeness::Complete

DockerBuildCache/DockerImageCache resources use ResourceLocator::Tool (a tool-native id, not a filesystem path) because Docker’s build cache has no single canonical path. DefaultEvidenceCollector can only run tool_liveness for a Tool-locator resource — open_by_process, process_cwd_match, and git_state always come back Unavailable(NotAttempted) for it, and required_evidence() still requires all three for Complete. AutoSafe for Docker build cache is out of reach until a future ticket gives this resource kind a real path-based or tool-native correlation strategy.

Docker image cache is unconditionally Protected; Docker has no registered cleanup action

Detectors do not yet distinguish a persistent, named Docker volume from disposable image cache, so every DockerImageCache resource is classified Protected unconditionally (see Safety Model). Separately, Docker (both build cache and image cache) has no registered Action at all — it stays detect-only, per an accepted design cut. Neither of these is accidental; both are documented design decisions, not bugs.

Resolved (HORO-957): Homebrew’s cleanup action is registered but never actually executes

HomebrewCleanupCache’s ActionStep::RunTool has scoped_path: None because there is no narrower, safe way to ask brew to clean only one thing — real brew cleanup -s has no path argument to scope to. HORO-957’s independent golden-scenario evaluation found that, before this fix, that meant executor::execute() ran this step with zero identity/TOCTOU guard, and confirmed the real Homebrew cache classifies AutoSafe on a real developer machine today — a live risk that emergency/free --target could silently trigger a real, irreversible brew cleanup -s. Fixed: execute() now refuses any RunTool step with scoped_path: None unconditionally, fail-closed, before ever spawning the tool. dry_run/ clean --dry-run still render this action’s plan; real execution stays permanently refused unless a scoped equivalent becomes available upstream. HORO-1358 carried that refusal upstream into reporting: detect --json and explain --json now report this candidate as executable: false with no offered action, because reporting puts the action’s own plan to the same pre-mutation structural rule execute() applies. HORO-1360 moved that rule into one shared predicate (actionability) and pointed two more surfaces at it: the BYOK LLM prompt no longer names homebrew.cleanup.cache among a resource’s offered action ids, and autopilot reports such a candidate as ineligible without spending one of its bounded attempts, rather than charging an attempt and reporting an execution failure that was certain in advance. clean --dry-run, free and emergency still resolve actions independently and do not read the predicate, so they can still nominate homebrew.cleanup.cache — where it remains refused at execution. That remainder is HORO-1359. See HORO-1005 for a related, non-blocking follow-up (scoped_path isn’t yet structurally tied to what a RunTool step’s args actually mutate — not currently exploitable, since this was the only unscoped action and it’s now refused outright).

Resolved (HORO-957): Cargo target/ and node_modules are now discoverable

Before this fix, DiscoveryContext::known_project_roots defaulted to an empty list and was never populated by any real CLI code path — so the two most canonical “developer storage hotspots” named in the epic were structurally undiscoverable no matter what was actually on disk. Fixed: a repeatable --project-root <path> flag is now wired into detect/ explain/clean/free (deliberately not emergency, which takes no arguments by design).

glomeris free still declines every ASK candidate

This section used to say Ask had no handling at all, because no interactive prompt existed. Narrower than that now: glomeris execute --confirm-ask --observed-fingerprint <token>, the menu-bar app’s Clean button, and Autopilot’s --preauthorize-ask all supply a real UserConsent — see Safety Model.

What remains is specific to one command. glomeris free’s recovery loop wires RecoveryConfig::auto_approve_ask: false (src/main.rs), so every Ask-classified candidate inside that loop is reported as declined/skipped and never executed, however much of the target it would have reclaimed. There is still no TTY prompt anywhere in this codebase; a free run cannot ask you mid-loop, so it does not ask at all. To act on an Ask candidate, use explain --json to read its fingerprint_token and then execute.

A pre-authorized ASK under Autopilot cannot complete (HORO-1310)

Autopilot’s --preauthorize-ask <kind>:<reason> grants narrow advance consent for one ASK reason on one kind. The gate admits such a candidate, policy::approval::authorize issues a real Approval for it, and then executor::execute’s deletion-time revalidation always aborts it with AbortReason::PolicyClassDowngraded. Nothing is deleted.

The cause is upstream of Autopilot and shared by every deleting command. Ask/RebuildCostHigh arises only from a per-instance Regenerability::NotRegenerable, while executor::build_fresh_evidence rebuilds regenerability from the resource kind’s static default — so the fresh classification lands on AutoSafe and the class comparison trips.

Left as-is deliberately: the failure direction is the safe one (refuse, mutate nothing), and changing build_fresh_evidence changes the TOCTOU anchor execute, free, emergency and autopilot run all depend on. It is pinned by src/autopilot/run.rs’s a_preauthorized_ask_still_aborts_at_deletion_time_revalidation, so the fix starts from a failing test that names the cause. See Autopilot.

Correlation depends on lsof/git/pgrep being present and stable

Runtime correlation is subprocess-based. CI validates this against macos-14 only, where all three tools are confirmed present; behavior on other macOS versions is not separately verified. tool_liveness for non-daemon tools (Cargo, Npm, Pnpm, Yarn, Homebrew) is structurally always Unavailable(ToolNotRunning) — there is no persistent process whose presence would mean “this tool is active” for them, which caps their Evidence::completeness() at Partial/Confidence::Medium, never Complete/High. This is a deliberate fail-closed consequence, not an anomaly to “fix.”

launchd real-scheduling behavior is not exercised by cargo test

Only plist generation and path/file logic are unit tested. Actually asking launchd to load and run the agent on a schedule needs a real macOS user session and is not covered by CI.

The protected-path matcher is conservative, not exhaustive

See Safety Model — the denylist covers a specific, named set of categories (credential material, git internals, infra state, system paths, unsafe mounts) and is explicitly documented in its own module comment as a starting point to extend, not a completeness guarantee.

Resolved (HORO-1008): the BYOK LLM planner now has a CLI surface, and PROTECTED refusal is proven end-to-end

glomeris llm-plan [--project-root <path>]... [--plan-file <path>] [--json] is an advisory, non-executing subcommand wired on top of actions::llm — see BYOK LLM Planner and CLI Reference. It never constructs a policy::Approval and never calls policy::approval::authorize or executor::execute.

--plan-file <path> feeds a fixture response through the exact same extract_plan/LlmPlan/plan_with_llm pipeline the live provider uses, with no network call. tests/golden_llm_plan_protected_refusal.rs uses this to prove, through the real CLI-facing crate::cli::build_llm_plan_report function, that a plan request to delete SSH key material is refused — closing the gap the previous version of this section described: HORO-943’s golden acceptance scenario step 6 (“prove a protected resource cannot be deleted even if an LLM plan requests it”) was previously verified at the code level only (llm_plan_item_never_bypasses_policy), never through an actual CLI input surface.

Remaining, deliberate scope cuts for this subcommand specifically:

  • Advisory-only, by design — no --execute/--yes flag exists or is planned for this subcommand.
  • No retry/backoff, and no streaming — mirrors actions::llm’s own existing limitations (see BYOK LLM Planner).
  • Only one provider shape (OpenAiCompatibleProvider, any OpenAI-compatible /chat/completions endpoint) — no Anthropic-native or Azure-OpenAI-specific auth.
  • Not wired into glomeris free’s recovery loop — that remains a future ticket’s optional enhancement, per actions::llm’s own module docs.

Separately, and independently of the LLM planner: a previous version of this section claimed that no live detector could ever emit a resource whose path matches a policy::protected pattern, and that PolicyClass::Protected was therefore enforced at the code level but not reachable by an evaluator driving only the shipped product’s detectors. The first half of that is wrong, and HORO-1313’s release-gate pass disproved it on a disposable fixture.

What a protected pattern matches is a path component, not an installed location, so anything a detector can discover underneath one is protected. A real node_modules tree created under a path containing an .ssh component was found by the live Node detector, classified Protected with reason protected_credential_material and executable: false, and then refused by the real executor twice — once by a plain execute, once with a valid --confirm-ask and matching --observed-fingerprint — exiting 3 both times with nothing deleted.

So the accurate statement is narrower, and the safety conclusion is stronger rather than weaker. The conventional locations of the tool caches Glomeris knows about do not normally sit under a protected path, which is why Protected is uncommon in day-to-day use; a project root that does sit under one reaches it through ordinary discovery, and the refusal holds when it happens. --project-root is the usual way to get there, deliberately or by accident.

tests/golden_llm_plan_protected_refusal.rs still reaches Protected through a hand-built Evidence fixture, and that remains the right shape for a hermetic test — it does not depend on a tree existing on the machine running CI.

Resolved (HORO-957): prebuilt release artifacts

cargo-dist packaging produces macOS artifacts for aarch64-apple-darwin and x86_64-apple-darwin, published as GitHub Release assets with checksums. Building from source remains fully supported.

Resolved (HORO-1305 through HORO-1309): there is a GUI

This page used to say a SwiftUI app was not part of this MVP. There is one: a LSUIElement menu-bar app, documented in Menu Bar App.

What has not changed is where authority lives. The app is a thin client over the same CLI — it shells out to glomeris and renders what comes back. It classifies nothing, decides nothing, and holds no policy logic, which a CI guard (scripts/check-no-policy-label-branching.sh) enforces mechanically rather than by convention. Every screen maps back to a named CLI invocation; that table is at the end of the Menu Bar App page.

What the app offers no way to invoke is running something without being asked: glomeris emergency has no button, and neither does glomeris autopilot run. Both still show up in the app’s history when run from a terminal, which is the point of a shared audit trail.

Autopilot’s grant is a different question, and this page used to get it wrong too: it said the envelope could only be granted or revoked on the command line. That is no longer true. Settings → Autopilot reads autopilot show --json and writes through autopilot enable/autopilot revoke, so the one authorization in this product that lets something delete without asking again can be read and withdrawn by someone who never opens a terminal — see Menu Bar App. The distinction the app keeps is between authorizing and acting, not between the terminal and the GUI.

What holds this book to the code

Three mechanical checks, because the drift this page is about was found by reading rather than by CI:

  • tests/help_golden.rs pins every rendered help surface byte for byte against committed fixtures, so a command’s own help text cannot change silently.
  • scripts/check-docs-cover-cli-commands.sh requires CLI Reference to have a section for every command in src/cli/help.rs’s single COMMANDS table. A new subcommand now fails CI until it is documented.
  • scripts/check-vocabulary-covers-cli-tokens.sh compares the CLI’s JSON tokens against the menu-bar app’s wording for them, in both directions.

None of that can catch prose that goes stale, which is what the rest of this page is for. Where this book and src/main.rs disagree, the source is right and the disagreement is a bug on this page.