mcp-salt: letting AI agents operate a SaltStack fleet
mcp-salt is a Model Context Protocol server. It lets an AI agent such as Claude Code work on a fleet that is managed by SaltStack. The agent can read state, run commands and apply changes, all through Salt’s own API and as a named identity with its own grants. There is no SSH and no shared admin account.
It was built for a homelab that is deliberately run like production (about-this-homelab), as a testbed for patterns meant for much larger networks. The design therefore aims at the general case rather than the smallest one that would work here.
Design goals
- Auditable. Every action is a Salt job, with a job id, attributed to the identity that made it, and stored. Wiki changes link back to those job ids.
- Least privilege, enforced where it can’t be bypassed. Salt’s external authentication (eauth) on the master decides what each identity may do. mcp-salt itself only advises.
- Preview before change. Anything that mutates the fleet is two-phase. You get a dry run and a digest first, and the commit happens only when that exact digest is confirmed.
- Secrets never pass through the model. Credentials the agent needs are held as opaque handles in memory, or written straight to disk. They are never put in its context.
- Humans stay the merge gate. An agent proposes changes to code as merge requests from its own fork. A person merges them.
- Not tied to one model. Any MCP client works, and the launcher can point the agent at a locally served model instead of a hosted one. Nothing in the tool surface assumes a particular model.
Architecture
AI agent (MCP client)
| stdio, JSON-RPC
mcp-salt (one process per session; holds the session's salt-api token)
| HTTPS + mutual TLS client certificate
TLS gate (nginx: client-cert verify, source allowlist, request hygiene)
|
salt-api -- salt-master (eauth via LDAP; master-side "ai_*" runners)
| \-- job cache (PostgreSQL, redacted, append-only role)
minions (execution modules) audit table (append-only)
- Two gates in front of the API. The gate rejects any request without a client certificate from the dedicated AI-ops CA. A valid certificate only gets you to the login. The identity still needs its LDAP password to do anything. The certificate is treated as authoritative for who the client is, and its CN must be the eauth identity.
- Startup preflight. On launch mcp-salt logs in once. A bad credential fails the launch
with a clear message, instead of failing on the agent’s first call. The login response
carries the identity’s resolved grants, and mcp-salt renders them into a per-tool
“granted / not granted” map in the model’s initial context, so the model doesn’t learn its
limits by trial and error. When a capability is missing it can say which grant would be
needed rather than looking for a way around it.
salt_whoamireturns the same map on demand. - The degrade engine. Salt is asynchronous underneath, and a “synchronous” API call just
holds the wait inside an HTTP connection. If that connection dies, the job id goes with
it. So every slow call is dispatched asynchronously (a job id in under a second) and then
awaited for a sync budget of about 50 s. The wait uses the event stream with
per-minion accounting, so a fan-out job counts as done only when every targeted
minion has answered, not just the first. A short call feels synchronous. A long one comes
back as a handle,
{status: running, jid, returned, pending}, never as a transport error. The job keeps running and can be picked up withsalt_job_waitorsalt_job_lookup. - Response-size budget. A shared helper keeps every reply within a character budget (8000 by default, adjustable per call). JSON is trimmed by halving its largest lists and stays valid JSON. Text is cut at a line boundary. Every cut is labelled with what was shown, what the total was and how to narrow the query. A truncation is never silent. Wiki reads can be paged by table of contents, section or offset.
- Plugins. The core supplies the Salt tools. Plugins add tool families, found through
Python entry points, so a new family needs no fork of mcp-salt. A plugin can also supply:
- its own section of the server instructions;
- grant rows for the startup map;
- live health checks for
mcp-salt --preflight.
- Credential handling on the workstation. The salt password is prompted for each session by a launcher and lives only in the process environment. Each session binds its own private Unix socket for its credential helper, so concurrent sessions can’t take each other’s credentials.
Tool families
Fleet inspection and execution (core)
- Read-only tools:
salt_test_ping,salt_grains,salt_file_read,salt_file_findandsalt_state_show_sls. - Generic dispatchers:
salt_execution_runandsalt_runner_run, plus async variants. These reach any execution module or runner the identity is granted. Salt has hundreds of modules, so one curated tool per module wouldn’t scale, and eauth already decides which of them a caller may use. - Commands:
salt_cmd_runruns a shell command. eauth scopes it by host class, so most tiers can’t reach core hosts with it at all (see the security model). - Jobs:
salt_job_waitandsalt_job_lookuppick up anything the degrade engine handed back as a handle.
Changes: two-phase by construction
salt_state_apply. Phase 1 runs the state withtest=Trueand returns the would-change set plus a digest. Phase 2 has to present that digest. The runner re-runs the dry run, recomputes the digest and refuses if anything moved between preview and commit. Only then does it apply for real. There is no way to skip the preview.salt_edit. This is an exact-string replacement that has to match exactly once (zero or several matches are refused). The write is atomic and keeps the file’s mode and owner.
The LLM-maintained wiki
The wiki is a git repository of Markdown pages, rendered with Quartz. This page lives in it. Agents read it and maintain it as the fleet’s institutional memory. The rule is that what was figured out once gets written down here, so it stays figured out.
- Read:
list,readandsearch. Search is moving from grep to hybrid retrieval (vector, full-text and literal, through the RAG index below), which falls back to grep automatically and says so. Reads can also be paged by section. - Write:
write,editandapply.applylands a set of related page changes as one git commit. All three are two-phase: the preview returns a diff, lint warnings and a digest, and the commit needs the digest. All three use compare-and-swap: the writer passes the hash of the body it read, and a stale hash refuses the write, so the agent re-reads before it overwrites anyone. - The “Be Prepared” invariant. Lint and the snapshot happen in the same operation as the write. Every write is a git commit, and git history serves as the wiki’s snapshot store. A change set is all-or-nothing, so no commit is ever a half-applied, inconsistent restore point.
- Lint: checks frontmatter, dead links and orphans. It runs on every write, and on the whole wiki on demand.
- Reads from a replica. Reads are served from a replica on the master, so they add no job-cache rows and still work when the wiki host is down. Writes go to the authoritative repository, where compare-and-swap makes the replica’s few seconds of lag safe.
- Agent coordination. Each identity keeps a small intent page, saying what it’s working on and what it’s touching. A hook loads it into context every turn. Another hook reminds the agent to check other agents’ intent pages before a write that could overlap.
GitLab: the fork-and-merge-request workflow
- The plugin mints short-lived personal access tokens for the GitLab user whose name matches the eauth identity. Tokens last an hour by default and a day at most.
- The token never reaches the model. mcp-salt keeps it in process memory and hands back an opaque handle. A git credential helper fetches it over a private local socket when git needs it, so it never appears in a remote URL, on a command line or on disk.
- AI identities are external, non-admin GitLab users with Reporter access. They can read
and fork anything, and they can’t push to a canonical branch. Work happens on the
identity’s own fork, and the agent opens a merge request with
salt_gitlab_open_mr. A human merges it.
mTLS client certificates
- First certificate. A new client redeems a single-use, short-lived join token that an operator mints. The client generates its keys locally and sends a CSR. The CA imposes the subject and extensions itself, ignores what the CSR asked for, and returns only the certificate. Private keys never move.
- Renewal. A live client renews itself over its own mTLS session. The renewed key and
certificate are written straight to disk (
salt_client_cert_issue) and never pass through the model. - The lease. The right to self-renew is a lease that only successful logins keep alive. An abandoned or stolen machine can’t keep minting certificates for ever.
- Re-enrollment. Mapped identities, including non-interactive service identities, are re-enrolled monthly, which rotates both keys.
The external brain
Agents waste tokens rediscovering the same infrastructure every session, by reading long prose and running ad-hoc commands, and nobody can check where the “facts” they end up with came from. The external brain is a model-agnostic substrate that sits behind a small agent API. The model does recognition; a logic engine does the multi-hop calculation and returns proofs. Its main metric is tokens per correct inference.
| Verb (tool) | Engine | What it answers |
|---|---|---|
pack | renderer | ”What is X, what runs where”: a compact, budgeted context pack around some seeds. With no seeds it returns the cached default pack, a map of the fleet of roughly 1.5k tokens. |
need | abductive reasoner (SWI-Prolog) | “What breaks if X fails”, “does A affect B”, “how does A reach B”. |
facts | fact store | Exact values, including past ones (as_of, superseded history). |
search, history, why | wiki RAG | The section that explains something, how a page evolved, and why it says what it says. |
collect | collectors over Salt | Refresh stale or missing facts. This is two-phase like every other write. |
registry, status | — | Which collector produces which predicate; health and freshness. |
The posit (hypotheses on a scratch branch) and act (parameterised, idempotent states
taking a proof) verbs are designed but not built.
-
Facts. A fact is subject–predicate–object, with an
as_ofand a validity interval. Its provenance is one of five kinds:- live: observed by a collector, keyed to the Salt job that saw it;
- declared: taken from configuration;
- prose: a wiki page at a given commit;
- derived: produced by a rule from premises;
- scratch: a session’s hypotheses.
The store is PostgreSQL with Apache AGE for the graph and pgvector for embeddings. Nothing is deleted. A fact seen again is confirmed, and one that disappears from a complete scope is superseded. The writer role has no DELETE.
-
Unknown is not false. Only a harvesting collector can declare a scope complete (
#closed). Outside a closed scope, a missing fact is UNKNOWN.needanswersproof,partial,unknown(listing the missing leaves and the collector that could supply each one) orno, and it saysnoonly when every failing leaf is inside a closed scope. -
Collectors are Salt execution modules listed in a registry, with targets, predicates, TTLs and triggers. Each run is one Salt job, and that job id becomes the provenance of every fact it yields. They collect structure, not secrets: names, states, mount points, listeners and dependency edges, never a process’s environment or command line. The authentication collector records who may authenticate to what, and by which method, never the credential itself. A redaction guard keeps every secret seen while parsing in memory and scans the serialized return for each one before it leaves the minion. A single hit withholds all of that run’s facts and reports only a count. Unit tests full of realistic fake secrets pin this behaviour.
-
Wiki RAG. The wiki’s entire git history is indexed append-only, by commit and by section, for hybrid search: vector, full-text and literal, fused with reciprocal-rank fusion, with optional reranking. Embeddings come from a small locally served model. That index enables:
as_ofqueries against the wiki at a past point in time;historyfor a page or topic;why, which finds the commit that introduced a statement and returns its recorded rationale. Wiki commits carry the Salt job id that made them, so that rationale joins the audit trail.
On an internal question set, recall@5 rose from about 0.2 with grep to about 0.6.
-
Default pack and fold handles. Packs are rendered deterministically and cached after every harvest. Detail that doesn’t fit the budget is folded into counts with handles (
+pack:x,+detail:x/pred,+collect:c) thatunfoldexpands on demand. Nothing is dropped silently. -
Event-driven ingest. The rule is “events trigger, ranges carry, reconcile verifies.”
- A source fires a small Salt event once a change is readable, without blocking the write and without secrets.
- The reactor checks the sender and starts the consumer fire-and-forget.
- The consumer ignores the payload and processes from its own watermark to the head, so a lost or duplicated event heals on the next one.
- A scheduled reconcile proves the events work.
A wiki commit becomes searchable in a few seconds.
On an internal set of operations questions, need answered far more of them correctly than
reading the wiki prose did, using a small fraction of the tokens.
Salt-side components
mcp-salt deliberately contains little logic. The interesting code runs on the master, so
the same safety rules apply to every caller: an agent through mcp-salt, a human at
salt-run, a timer or the reactor.
| Component | What it does |
|---|---|
ai_safe runner | The guarded write surface: two-phase state_apply and edit, wiki read/write/apply/lint, GitLab token minting, client-certificate issuance. |
ai_exec | Target-scoped execution. Execution modules can be scoped by host in eauth, but runners run on the master and have no target to scope. So the containment for runner-mediated work lives inside the runner: caller check, target check, a secret-path deny list, audit. |
ai_ldap, ai_enroll, ai_lease | Directory operations limited to ai-* accounts, certificate enrollment, and the self-renewal lease. |
| storage control plane | A Proxmox storage plugin calls a runner over salt-api instead of SSH-ing to the storage host. Pool, target and host come from trusted master-side configuration, never from the caller. |
ai_common | One shared library for every runner. See below. |
ai_safe_jobcache returner | The master job cache: PostgreSQL first, local cache as fallback, with redaction before any backend. |
brain, brain_rag runners + modules | Collectors, fact store, reasoner and wiki RAG for the external brain. |
ai_common came out of an architecture review. The runners had grown up as
copy-and-paste siblings, with six caller lookups implementing four different policies, three
pillar readers and six audit writers in three formats. The library replaces that with:
- One caller policy. Each caller is classified into a tier:
ai,human,root,service,reactororunknown. An unidentified caller or the reactor is never allowed to mutate. Reads by unidentified callers are allowed but logged. The daemon’s own run-as user is never mistaken for the caller. - A master-side pillar read that compiles in-process, so no minion job is involved. It refuses a compile with render errors, so a runner never carries on with keys silently missing.
- The digest kernel, described under the security model.
- One audit row format with two sinks, described under the security model.
- A helper convention for the hop to a minion. Argv is a fixed absolute path plus fixed words, so data never rides on the command line. The request goes base64-encoded on stdin, and the reply is the last line of stdout as JSON. Because the request is itself job space, the library refuses any secret-named field that isn’t encrypted, and anything that looks like a private key or a token.
- A caller-encrypted envelope (below), with no cleartext fallback.
- A Postgres connect helper with tight timeouts and a circuit breaker shared across processes, so an unreachable database costs one short wait rather than one per call.
- The protected-entities registry and the level ladder (below).
Events and triggers. Periodic work runs on systemd timers on the master. Events go
through the reactor. Each mutation domain is meant to announce its commits as
salt/ai/<domain>/commit events, so the brain can re-collect exactly what changed.
Security model
Identities
-
Every human gets an AI identity,
ai-<identity>, deliberately separate from their own login. It holds only the grants its AI work needs. Its audit trail is distinct from the human’s. A narrowly scoped credential is acceptable to cache in places where a person’s primary credential would not be. -
Who and where. Grants are tiers of LDAP groups (FreeIPA): read-only or read-write, for tech, dev and admin scopes. Host classes are FreeIPA hostgroups that the master resolves into Salt nodegroups. Group membership is declared in git and converged by a Salt state. A service principal with a narrowly scoped IPA role does the converging; it can only manage
ai-*groups and can’t delete anything. -
Interactive and autonomous identities are separate classes. A human-paired session works with a person approving each tool call. Autonomous identities get narrower, path-scoped grants. Their wiki writes, for example, are confined to their own namespace.
-
Levels. On top of the tiers there is a generic ladder, which each domain adopts:
Level Meaning L0 Reads only. L1 AI-owned and reversible. L2 Non-core, non-protected, not declared elsewhere. L3 Core, edge or destructive. Needs an explicit acknaming the objects, and a second group (two keys).declared Owned by a Salt converge. Refused at every level, with a pointer to the merge request to make instead. Each domain declares its maximum level, whether callers must opt in, and its rate limits. An identity that hasn’t opted in to a domain is L0 there. Unknown levels or domains fail closed.
eauth is the authority; the tools only advise
The MCP tool surface is a view: it exposes a convenient subset and adds safety wrappers. Enforcement belongs to salt-master’s eauth, on every call, by identity, target and function. A compromised mcp-salt, or a compromised gate, is still bounded by what the identity is granted. The grant map shown to the model is advisory, a courtesy that the master enforces anyway. On top of all that, the MCP client asks a human to approve each write or exec call unless that person has explicitly turned approvals off.
Core hosts and protected entities
- Core hosts are the hosts an AI could use to change its own constraints or hide its tracks: the salt master, the git server, the wiki host, the identity server and the public edge. Every tier except one explicit core-write tier targets “all hosts except core”. Admin tiers reach core only through runners that carry their own guards.
- The core set can’t silently empty. “Everything except core” turns into “everything” if the core set is ever empty, so the nodegroup also ORs in a literal floor of the core hosts. A test keeps that floor in step with the declared list, and the resolver keeps its last good answer rather than shrink to nothing. Core can grow through the directory alone, but shrinking it takes a merge request.
- One protected-entities registry lists core hosts, public edge names, protected addresses and per-host secret paths. Every guard is meant to render from it, rather than each keeping its own list that slowly drifts.
Writes go through runners
The agent never gets a credential for a backend. Directory changes, database writes, wiki commits, token minting and certificate signing all happen inside master-side runners, using credentials only the master holds. The rule is: no writer gets a database credential. Authentication to backends uses host certificates generated in place, wherever possible, so for those there is no password to leak at all.
Preview, then confirm, with caller-bound digests
digest = sha256(version, domain, tier:caller, level, normalized changes, hash(before-state), minute)
The confirm recomputes the digest against the current caller and state. It is refused if:
- a different identity, or the same name in a different tier, tries to confirm;
- the level changed;
- the changes or the before-state moved;
- more than about 15 minutes have passed;
- the digest is from the future.
So a preview can’t be replayed by someone else or a day later. The action kernel being built
on top of this adds an explicit ack for destructive changes, and refuses a whole batch at
preview if any item in it is refused.
Secrets never enter job space
Job space is everything in a Salt job’s arguments, its return and a state’s change diff. All of it is published on the event bus and stored in the job cache, readable by any authenticated caller with the right grants for the cache’s retention period. Pillar is not job space: it is compiled per minion and delivered point to point. The rules follow from that asymmetry.
-
Fetch secrets in-process on the master, never through a minion job. Fetching a secret with a remote
pillar.itemorconfig.getbroadcasts it. -
Never return a secret in the clear. When a runner has to hand back something it just generated (a token, or a certificate’s private key), the caller sends a one-time public key. The runner returns an envelope: a data key wrapped with RSA-OAEP, around AES-256-GCM ciphertext. Only the calling process can open it, and it throws the private half away after the call. The bus and the cache carry only ciphertext. The format is hybrid because RSA-OAEP alone can’t wrap a whole PEM key.
-
Redact at the job cache. A returner scrubs named secret fields from known functions before any storage backend sees them. It is the backstop if a caller forgets the envelope.
-
Deliver secrets by pattern, chosen by where the secret can be generated.
- CSR: the key is generated where it is used and never moves. This is the preferred pattern.
- Plain pillar scoped to one host, for a single static secret.
- ext_pillar, for per-minion secrets minted elsewhere, from an explicit allowlist in git. It fails closed by returning nothing, never by raising.
In each case the consuming state suppresses its diff, so the secret doesn’t go straight back into job space.
-
No secrets on argv. Command lines show up in process listings and in job arguments. Secrets travel on stdin, in files with tight modes, or through the credential-helper socket.
-
Handles, not values, in the model’s context. Tokens live in the MCP process’s memory and die with it. Certificate keys go straight to disk.
Audit everywhere
- The job cache is in PostgreSQL. The master’s database role can insert and read but cannot delete or alter schema, so neither the master nor anything using its identity can erase history. Retention is enforced in the database. Authentication is by the master’s host certificate; no password exists. The local cache stays on as a fallback, so a database outage degrades the tooling rather than breaking it.
- The audit table. The shared library writes one JSON audit row in a common
format for every mutation: caller, tier, function, level, target, digest, result and a short summary.
Fields with secret-like names are redacted. A
startedrow is written before the effect and the outcome row after it. Each row goes first to a local log file, the durable copy, and then to an append-only table that triggers protect from update, delete and truncate for every role, the database superuser included. Because that table sits beside the job cache, audit rows join job results by job id. - Attribution. eauth records which identity made each job, even when a group membership was what authorised it.
- Commit trailers. Every wiki commit carries the Salt job id of the write that made it. So “why is this like this?” can be answered by machine: from the commit, to its rationale, to the audited job.
Git is the review gate
Declared configuration (states, pillar, group membership, the core host floor) lives in git. AI identities change it only through merge requests from their own forks, reviewed and merged by a human. Anything declared is refused by the imperative runners, which point at the merge request to make instead. The imperative change would be reverted by the next converge anyway.
Lessons that became principles
These came from real near-misses found while building this system. They are written as general rules.
- A dry-run flag that isn’t honoured is a write. A
test=Truethat one code path quietly ignores, or a misspelt keyword argument that gets silently dropped, turns a rehearsal into the real thing. Check that the dry run is one before you trust it. And build the refusals so they hold even when it isn’t. - Fetching a secret can broadcast it. In a pub/sub control plane, a remote read is a publication. Read secrets where they already are.
- What a runner returns is published and kept, not just handed to the caller. Assume every return is readable by every authorised caller for the whole retention window.
- Secrets leak through process listings and broad config reads. Argv, generic “get config” calls, and library calls that log their whole configuration on error are all leak paths. Ask for the one key you need, in-process.
- Longer retention raises the stakes. Moving a cache from hours to weeks multiplies the cost of every redaction miss. Tighten redaction before lengthening retention.
- Fail closed on partial data. A configuration compile with render errors, or an empty answer from a group resolver, must stop the operation, not let it go ahead with defaults.
- Group changes don’t reach live sessions. eauth tokens carry the groups they were issued with. A revocation takes effect for an existing session only when its token is replaced.
- Running the CLI as root can defeat retention. A job-cache entry written by root may be one the daemon can’t delete. Run tools as the service identity.
- Searching for a leak can create the hit. On a bus that publishes commands, a search for a secret pattern publishes the pattern. Build the search string at runtime.
- A hostname that answers isn’t evidence it’s the right one. When a host contradicts the documentation, fix the drift at the source rather than trusting whichever you read last.
- Collect structure, not secrets. Collectors record that a credential, rule or account exists, and its shape, never its value. A guard redacts anything that slips through.
Status
- Live:
- fleet tools and the degrade engine;
- two-phase state apply and edit;
- wiki read, write, apply and lint;
- the GitLab fork and merge-request workflow;
- mTLS enforcement and enrollment;
- the PostgreSQL job cache with redaction;
- the shared runner library, with its caller policy and append-only audit table;
- the external brain’s runners.
- Rolling out:
- the brain tools in mcp-salt (the agent-facing
salt_brain_*surface, default packs), hybrid wiki search and paged reads; - migrating the older runners’ digests, audit writers and pillar reads onto the shared library.
- the brain tools in mcp-salt (the agent-facing
- Next:
- one generic two-phase action kernel (classify, digest, ack, snapshot, rate-limit, audit, commit event), shared by upcoming directory, DNS and secrets-vault domains;
- commit events that feed postcondition facts back into the brain;
- a typed in-memory secret store for more kinds of credential.