@agent-core

Security

agent-core's security posture: the shell sandbox, URI-policy gating, env sanitization, data isolation, and the MCP boundary.

agent-core enforces the platform's deny-by-default posture at the points where a model's output turns into real action: the shell, tool dispatch, package data, and the MCP boundary. This page covers the agent-core-specific mechanisms; the platform-wide model is described in the enterprise security model.

The shell sandbox

execute action="shell" passes two independent guards before anything spawns.

1. Command analysis. The command is tokenized into an AST and inspected structurally — no regex matching. Always rejected: destructive system binaries (mkfs, shutdown, sudo, user/group management, partitioning, firewall tools), eval used as a command, fork bombs, dangerous rm -rf forms, and any reference to the platform app zone or credential files. The parse is bash-faithful, so escapes and quoting cannot hide a name: \sudo, s\udo and "\mkfs" are resolved exactly as bash resolves them before the check runs.

A nested bash -c "…" is re-analysed, not rejected — and by both guards. The literal script inside it goes through the same checks, so a blocked binary there is caught, and its paths are checked by the URI-policy gate below exactly as if the command had been written flat: bash -c "…" grants nothing that the same command without the wrapper would not. Nesting itself is legitimate and stays allowed, up to four levels — deeper than that the command is refused rather than merely left uninspected, since a limit that stopped looking would be a way around the checks rather than a bound on them. The recognition is not tied to the exact spelling: bash reads a short-flag cluster as a set, so -lc, -cl, -ec and -xc are all the same flag and are analysed the same way.

A wrapper prefix is resolved to the command it actually runs, in both guards — timeout 5 rm <path>, env A=1 rm <path> and busybox rm <path> are judged exactly as rm <path> is, and a blocked binary cannot hide behind one. find is read the same way: -delete and an -exec that runs a write-class command mark the directories being walked as writes — while find <tree> -exec cp {} <dest> still reads the tree it walks, so backing up a read-only directory keeps working. The enumeration deliberately under-approximates — a wrapper nobody listed keeps its ordinary classification rather than being newly denied, a command built at runtime from a variable stays undecidable by definition, and a wrapper that relocates the base (env -C, unshare -w, chroot) moves the meaning of the paths after it without this layer following, so those are checked against the session's own directory rather than the relocated one. The kernel sandbox below is what confines all of these.

This layer is defence in depth, not the boundary — the OS-level sandbox is. The one exception is the operator's confinement: unconfined host mode: there no kernel sandbox is applied at all, so these static checks and the URI-policy gate below are the only path enforcement a command meets. That mode is an operator opt-in on a single-operator machine, never something a role can select.

2. The URI-policy gate. The shell is a full enforcement layer of the URI-policy system — the same evaluator that gates connectors, routes, and the UI:

  • The working directory needs exec. The resolved cwd must fall under a registered source whose persisted policy grants source-level exec to the caller — the same check the terminal applies to its tabs.
  • Path arguments need read or write. The command is parsed with a bash-faithful parser, so comments end at the line (a write on the next line is still seen) and heredoc bodies are not mistaken for commands. Every absolute path token is mapped back into its owning source's URI space and evaluated: read for reads, write for mutating commands (rm, mkdir, tee, chmod, copy/move destinations, sed -i, redirect targets, and the destination of an archive extractor such as tar -x … -C, unzip -d, or cpio -i). Relative tokens are resolved against the cwd and checked the same way, reads included — closing directory-escape forms like cat ../../../etc/passwd and a plain cat admin/audit/x alike. The check cannot tell a filename from a search word, so a word that names a protected path next to the cwd is refused with a reason saying it was read relative to the cwd. A command that cannot be parsed is denied.
  • Uncovered paths deny by default. The narrow allowances are /dev/null-style streams, /tmp reads and cwd, and the active skill's ${SKILL_DIR} subtree (read and cwd only — writing into a skill directory still denies). /tmp writes are confined to the caller's private per-conversation scratch directory (/tmp/nrs-…, created 0700 before the shell runs, with TMPDIR pointed at it) — shared /tmp is no longer writable across conversations.
  • Fail-closed. If the gate or its source index cannot be built, the call is a structured deny — never a passthrough. Every denial is audited with the command, the failing path, and the missing capability.
  • The container plane has an entry feature: exec.container. A shell opens in three steps, the same on every plane — the tool (execute requires core.execute), the door (the plane's entry feature), then the regulator (the source's own exec policy on the working directory). exec.container is granted wherever core.execute is, so a packaged skill's shell scripts need both; a caller holding neither it nor the bypass below is refused before any path is evaluated.
  • exec.unconfined bypasses the gate. The feature carries no grant below the admin tier — admin holds it through its enumerated default grant and owner through the '*' wildcard, and it is grantable to any custom role. Holders can target any container path; for everyone else the source policies are the boundary. It opens the container plane by itself, so a holder needs no separate exec.container — and it grants nothing on the host or on a remote source. This power is intentional and unlimited — the safeguard for it is the tool-activity audit, not a path restriction: every shell call a holder makes is recorded.

Active-skill substitutions (${SKILL_DIR} and friends) are expanded before the token scan, so skill-script paths are visible to the gate rather than hidden inside variables.

The command analysis is defense-in-depth, not the boundary. A path that is constructed at runtime — inside a shell variable, a $(...) substitution, or an interpreter (node -e "…writeFileSync…") — cannot be seen by any static parser. Those are confined at the operating-system layer instead.

The OS-level sandbox

For every agent shell except holders of the unlimited exec.unconfined feature, the command is spawned inside a Landlock kernel sandbox (Linux's built-in filesystem LSM). A small helper turns the sandbox on and then runs the command: the kernel refuses to open any file outside the caller's allowed roots, no matter how the path was written — closing the runtime-path gap that no static analysis can.

  • The allowed roots are exactly what the caller's policy already grants (writable stays writable, read-only stays read-only), plus a read-only system baseline so standard tools still run, plus the private per-conversation scratch directory. A source whose rules are the same everywhere is one root; a source whose rules separate users, roles or agents is granted path by path below its root. A directory those rules split carries no grant of its own, so a shell can cd through it and use what it is granted below it, but cannot list it, or create, remove or rename entries directly in it. A file granted on its own there can be rewritten in place, but not replaced through a temporary file (as sed -i and many editors save). The filesystem tools are unaffected.
  • exec.unconfined holders are not sandboxed — the same intentional, audited bypass as the URI-policy gate.
  • The sandbox fails closed: if the host kernel lacks Landlock (ABI 3, ~Linux 6.2+), a lower-trust shell is denied rather than run unconfined; owners are unaffected. Availability is governed by the host kernel, not the container, so on-prem deployments should provision a Landlock-capable host. The same applies to the helper itself: a confined spawn whose helper cannot be found is refused with a clear reason, never quietly run without a sandbox.
  • Socket isolation applies here too. Landlock governs the filesystem and only the filesystem — it decides what a process may open, not what it may connect to. Confined shells and confined terminal sessions are therefore also denied the ability to create a Unix-domain socket, on the same terms described for host access below: anonymous ("abstract") sockets included, ordinary networking, DNS and process IPC unaffected. Programs that depend on a local Unix socket — a syslog daemon, an SSH agent, an X11 or D-Bus client — will not work inside a confined shell.
  • The terminal is not a second, weaker path. A source-scoped terminal session is spawned through the same helper with the same flags as an agent shell, and running a command in an existing session runs it inside that session's sandbox. The one deliberate exception is the explicitly named Container Root destination, which is gated on a separate feature that no role below administrator holds by default.

Host access

By default every shell — agent or human — runs inside the application container. Neuralis can additionally reach the host machine, so an agent can drive tooling that only exists there, but this is a separate plane with its own gates. It ships off, and it has no per-role bypass — not even for an owner; the one relaxation is a mode the operator writes into the broker's own ceiling file (floor 4 below).

The host plane is reached through a small broker: a service that runs on the host itself, one per installation, started by the operator. The container never mounts the host filesystem. A mount would be passive and permanent — once / is bind-mounted, every process in the container sees it forever, regardless of features. The broker is active and gated: nothing is reachable unless all four floors below hold. Its default transport is a Unix socket with no port at all; where that cannot work (Docker Desktop), a loopback-only TCP fallback exists. Binding a routable address is a hard startup error in either mode.

Five floors, all required

  1. Operator provisioning. The service unit, a secret file, and the deployment's read-only bind-mount of the broker runtime directory. No in-app role can create any of the three. This is deliberately not an in-app setting: platform configuration is editable by project administrators, so an in-app flag would not be a floor at all.

  2. Authentication. Every privileged request carries a shared secret read from an operator-owned file and compared in constant time. GET /healthz is the only unauthenticated route and exposes only bounded capability/count health. The container verifies actual readiness through authenticated GET /readyz; neither route reveals a path, the secret, or the ceiling.

  3. The operator ceiling. The operator writes a file listing which host paths the broker may touch. Every command resolves to effective = requested ∩ ceiling, and the intersection never widens: asking for a directory that contains a ceiling entry narrows down to the ceiling entry. An empty — or missing, or malformed — ceiling denies everything, because "not configured" must never mean "unlimited". Neuralis computes what it wants from the host source's URI policy, and that computation remains the policy and audit truth; the ceiling is the trust boundary, because the in-app half is exactly what a project owner controls.

  4. Filesystem confinement — with one operator relaxation. Every host command and host PTY is spawned inside the same Landlock sandbox described above, with the clamped roots — unless the operator's ceiling file declares confinement: "unconfined", a mode meant for a machine one person owns and runs: every host spawn then runs bare as the operator's own account, with that account's login environment, reaching everything installed for it. That is root-equivalent on the host, as the operator; it loads only beside trustedSingleOperator: true in the same file, every result reports unconfined, and it must never be used on a shared deployment. If the sandbox helper is missing — or present but unable to prove both of its layers in its self-test (a copy older than the socket filter) — the request is denied — in both modes, because the helper is required in both: outside the operator's unconfined mode there is never a bare spawn on this plane, and unlike the container plane there is no bypass feature here, not even for an owner. In that mode the ceiling file, the broker's service unit, its secret and the sandbox helper are all writable by the agent's own process — so the allowlist and the request channel become documentation rather than enforcement, and switching back to sandboxed is only trustworthy after you re-verify those artifacts (upgrade proves the helper; the ceiling and the unit by eye). The operator refreshes the helper with pnpm neuralis:host-broker upgrade.

    Landlock governs the filesystem, and only the filesystem. It decides what a process may open; it does not decide what a process may connect to. A confined child cannot open a host control socket as a file, but the kernel right that would stop it from connecting to one does not exist in any currently deployable kernel. That is why floor 5 exists rather than being folded into this one.

  5. Socket isolation. The confined child is additionally denied the ability to create a Unix-domain socket at all, which is what actually closes the gap floor 4 leaves open — including anonymous ("abstract") sockets, which no path-based rule can name. The broker self-checks this at startup by running the real spawn path against a fixed list of host control endpoints, and refuses to serve if any of them is still reachable. The check reports endpoint classes, never paths.

    A syscall filter is only as good as the numbering it covers. On 64-bit Intel hardware the kernel accepts a second numbering for the same calls, and a filter written against the ordinary numbers alone can be walked straight around by using the alternate ones — which is exactly what a startup check written in a high-level language cannot see, because such a language can only issue the ordinary form. The filter therefore rejects the whole alternate range outright, and the sandbox helper proves it about itself: it forks a child, installs the real filter there, and asserts both the denials and the capabilities that must survive (ordinary networking and process IPC). That self-proof runs when the image is built — a helper whose filter does not enforce fails the build — and again at startup on both the container boot probe and the broker readiness probe.

    Two consequences worth stating plainly. Programs that rely on a Unix socket — a local syslog, an SSH agent, an X11 or D-Bus client — will not work inside a confined host command; ordinary networking, DNS and process IPC are unaffected. And the strongest form of this floor is not something the broker can apply to itself: a process cannot shed the group memberships and login session it inherited, so a deployment that exposes this plane beyond its own operator should run the broker under a dedicated service account with no groups and no session.

The Host Plane kill-switch in platform configuration is exactly that: turning it on is necessary but never sufficient, while turning it off disables the plane immediately regardless of provisioning.

Features and routing

  • exec.host lets execute target the host; terminal.native opens exact host-source terminal tabs. Neither is granted below the admin tier (admin through its enumerated default grant, owner through '*'), and holding one is still insufficient without the provisioned broker.
  • terminal.supervise controls project-wide managed sessions. It has no manager/member default; supervising host sessions additionally requires terminal.native.
  • Attaching or editing a host-plane source additionally requires brain-core's drive.mount.host together with drive.mount.privileged. On the host plane every root is privileged by construction — a subdirectory is gated exactly like / — and editing is gated like attaching, because a permission widening is an escalation.
  • The plane is chosen by the owning source's connector, never by a name. A command's working directory decides where it runs; an ordinary in-container source called host-notes stays in the container. A bare absolute path resolves against the container first and only falls through to the host when no container source owns it, so a deployment with no host sources attached behaves exactly as before.
  • exec.unconfined bypasses the container gate and grants nothing here.
  • Background (detached) shells run on the host plane as broker-detached runs when the operator ceiling names a lifetime limit (maxDetachedLifetimeMs); otherwise the denial says so. They survive an application rebuild — the broker keeps the process and the platform re-attaches at boot — and end with the broker (lost) or at the lifetime limit (killed).
  • A host source is never a container terminal tab. Each host source is an exact broker-backed destination behind terminal.native; no host path is passed through container cwd translation and there is no aggregate host-root tab.

Host filesystem connector

A host source uses the common filesystem connector contract over broker RPC: list/read/write/delete/move/mkdir/rmdir plus bounded walk/count/delta/content scans. Brain-core still enforces the caller's source scope and uri-policy first; the broker independently clamps reads and mutations against the operator's read/write ceiling.

Reads go through the broker under the same operator ceiling as commands, with two additions: a path is resolved to its real location before it is checked, so a symbolic link cannot lead out of an allowed directory, and non-regular files are refused outright. The ceiling's executable directories — the interpreters and toolchains a command needs in order to start — are not readable through this path. Being reachable so a program can run is not the same grant as being readable, and conflating them would publish far more of the host than the operator agreed to.

A host source defaults to user scope: it belongs to whoever attached it. Project scope is an explicit opt-in, because file reading is granted to members and viewers by default — the host-plane features gate attaching and executing, not reading a source that is already attached.

Whole-home access is trusted-operator mode

If both the host source and operator ceiling grant the entire home directory read/write, the process can also reach credentials and user-owned control files inside it — including the default broker secret and ceiling. Landlock grants roots and cannot subtract a sensitive child directory. Keep ordinary roles narrow; strong isolation with whole-home visibility requires a separate OS identity and root-owned broker control files.

Platform support

Linux hosts, and WSL2 with an in-distro Docker engine, work with the default Unix socket. Docker Desktop requires the loopback TCP fallback. macOS has no host plane today: the sandbox helper is required in both confinement modes and it is a Linux kernel feature, so a macOS host resolves to "confinement unavailable" and every host command is denied.

Host sources participate in ordinary connector reads, scans, and indexing, so their Markdown and files are available through brain-core. Direct discovery and loading of executable source packages stays local-only: Neuralis never imports host JavaScript/WASM across the broker boundary. A future implementation needs an explicit, controlled staging/trust step rather than treating a host path like an in-container package directory.

Environment sanitization

The child process environment is rebuilt, not inherited. Exact-name secrets (database URLs, auth secrets) and any variable matching secret-like name patterns — *SECRET*, *TOKEN*, *PASSWORD*, *API_KEY*, *CREDENTIAL*, provider prefixes like OPENAI_* / ANTHROPIC_* / GEMINI_*, and more — are stripped from both the base environment and caller-supplied overrides.

The only way through the filter is the skill credential allow-list: when an active skill declares credentials: in its frontmatter, exactly those keys, resolved from the credential store at the caller's scope, are injected on the override side. The base environment is still stripped, so a host-level secret with the same name cannot piggyback on a skill's declaration.

One identity for skill calls

Skill scripts call back into the platform with the per-stream session ticket — an opaque token minted from the live stream's verified SessionContext and re-minted on every shell call so multi-turn skill sessions keep working. Routes derive the project and agent from the verified ticket and gate on the caller's real role and features. There is no service account, no synthetic owner, and no way for a script to claim an identity the stream does not have.

Per-agent data isolation

agent-core's manifest URI policies partition its data:// tree per agent:

  • An agent's own subtree ($self) is read-write; foreign agent subtrees are read-only, and their conversations/**, workflows/**, plans/**, and memory/** are not readable at all by default — owners and admins override.
  • A conversation belongs to the user who started it. Even under the agent everyone shares, another member can neither list, read nor continue it — not through the chat history, not through the filesystem tools — while owners and admins can. The rule names the requester (users: ["$owner"], see URI policies), so an owner who wants a conversation shared adds a grant in the source's policy editor. The composer's upload staging folder stays open to every member. A member's own container shell is held by the same rule at the OS layer: its sandbox is derived path by path from the source's policy, so a computed path reaches another user's conversation no more than a literal one does.
  • Each agent's identity files (SOUL.md, USER.md, HEARTBEAT.md under identity/) live inside that same $self boundary: an agent evolves its own identity through filesystem tools but cannot rewrite another agent's.
  • Workflow runs/** directories are engine-written — read-only for every caller.
  • The runtime-owned conversation files — the transcript (messages.jsonl), its metadata, the active SUMMARY.md projection, and the immutable summary revision archive — are write-protected by an immutable floor: no filesystem tool or route can write them, not even the owning agent or an owner/admin, so only the internal service that manages them can. Ordinary conversation artifacts (e.g. a plan.md) stay writable — the floor is scoped to the raw files.

Because most of these rules ship as manifest baselines and union into the persisted source policy, tightening or loosening them is a policy edit, not a code change — except the immutable floor above, which is deliberately not editable through the policy surface.

Summary privacy and the vector index

A conversation summary is stored as an immutable revision; the visible SUMMARY.md is a regenerable projection of the currently active revision. The revision archive is excluded from the vector index and from sync, so historical summaries are never embedded or searchable, while the active SUMMARY.md projection remains indexable like any other conversation file — and inherits the same per-path read policy, so a read-denied member never sees it.

Resync authorization

Re-provisioning an agent from its team template is a mutation, so the coarse core.agents route feature is only a discoverability gate. The real decision is a capability floor evaluated fail-closed: ordinary file/config resync needs agent-update capability, and an identity reset needs full agent-management capability — never a hard-coded role name. See agents.

Audit trail

Destructive and mutating conversation operations leave a trail. Deleting a conversation emits a conversation.delete audit event, and updating its title or metadata emits a conversation.update event — each recorded with the caller's userId, the conversation id, and the project/agent scope — so a deletion through the UI, a direct API call, or a skill is always accountable.

Tool activity is audited by two built-in host hooks that observe every agent turn:

  • tool.shell.exec — every successful execute action="shell" call, with the caller's identity, the (capped) command, the working directory, and whether the command itself failed. This is what makes an exec.unconfined holder's unlimited shell access accountable even though it bypasses the uri-policy gate — including the case where an agent collaterally deletes a conversation directory with a raw rm, which the route-level conversation.delete event never sees.
  • tool.failed — every thrown tool failure across all tools, with the tool name, a redacted input summary, and the error.

The terminal writes to the same trail — terminal.session on PTY attach/detach and terminal.blocked when the input guard rejects a typed command, keyed by user, project, agent, and session. Those records live in the platform app zone, not in a package data directory the audited agent could reach.

Both are observe-only and fire-and-forget: they can never block, alter, or slow a tool call, and an audit-write failure never breaks the agent. They do not log credentials — the shell hook records only the command and cwd, never the caller-supplied environment. The trail is the same app-zone audit.jsonl the admin Logs surface reads.

The MCP boundary

The hosted MCP server applies the same identity discipline to external clients:

  • Callers authenticate with a per-agent API key or an OAuth 2.1 JWT; both resolve to a fixed (userId, projectId, agentId) scope.
  • Header-based scope override is rejected — x-user-id, x-project-id, and x-agent-id headers from external callers are ignored as identity sources.
  • The system sentinel is rejected before dispatch — a request whose user, project or agent id is the reserved __system__ value never reaches a handler, so no external caller can claim background-system identity. The skill principal __skill__ is deliberately not part of that set: it is an audit identity, not a system one, and it is refused on its own path instead — an activated skill cannot mint a session ticket from another skill's identity.
  • Tool results are size-capped at 20 KB on content and structured content, so a tool cannot flood the wire or the model.

The MCP Apps sandbox

Interactive HTML from a connected MCP server (the ui:// MCP Apps surface) is untrusted external content and never touches the app origin:

  • The app runs inside the spec's isolated-origin sandbox — a throwaway browser origin on a second published port that serves nothing but a static relay shell. Session cookies do reach that port (cookies ignore ports), so a structural middleware fence 404s every other path there; the template HTML itself travels from the authenticated app origin to the shell over postMessage, never over the sandbox origin's HTTP surface. The isolated sandbox origin is mandatory and fail-closed: without it the card refuses to render ("sandbox unavailable") rather than falling back to a weaker mode.
  • The template's Content-Security-Policy is injected server-side from the resource's declared domains over a deny-by-default baseline (connect-src 'none'), and the template fetch has its own size cap.
  • Rendering requires the dedicated mcp.apps feature — no grant below the admin tier: the seeded admin role holds it through its enumerated default grant, owner through the '*' wildcard; grantable to any custom role — re-checked at the template route and at the app bridge.
  • Every app operation is bound to a short-lived, opaque View lease. The server mints the lease only after verifying the caller's fresh read access and that the app's own tool-result marker anchors the exact server and resource. The lease token lives only in host memory — never in a URL, query string, log line, or the sandboxed view's JavaScript — and the server derives the user/project/agent/conversation and server from the lease, not from repeated request fields. Each operation re-runs a fresh access (and, for approvals, feature) check: the lease is an ownership handle, never an authorization cache. A 60-second heartbeat renews it; when access is lost the lease is revoked on the next call and the view tears down.
  • An app-initiated tool call may target only the server that owns the app, and passes the exact model-call policy gate — hard deny rules, then the conversation's guard profile. When the guard asks, the approval is a standard pending interaction whose decision is durable in one write — there is no parked "approved" state. Approving runs the tool once server-side, from the originally persisted input (a replacement input is rejected), and returns the result; approving additionally re-checks the mcp.apps feature at decision time, so losing the feature between render and approval fails closed. A view can never approve itself, and one browser tab can never resolve another tab's pending call. A lost lease (page reload, expiry) closes the pending call by server-side expiry rather than replaying it.
  • External links from an app open only after an explicit human confirmation, and an app message can only prefill the composer — it never auto-sends.

Tool dispatch gates

Every tool call — native or package-contributed — passes the policy engine first, in a fixed order, and each of these five arms can deny outright: the global tool blocklist, the calling package's trust-tier capability ceiling, its capability gate, the untrusted-package tool allowlist, and the feature gates from the tool's declared requires.

Two of those arms are evaluated across every loaded package rather than against the caller's own rules, so only a first-party package may contribute to them: a tool blocklist and a per-tool timeout declared by an installed or dropped package are dropped at load. The feature gates are cross-package in the same way, and are bound to the package that DECLARED the tool — so a package cannot gate a tool it does not own, including by declaring one under a name that already exists. Everything a package declares about its own tools — allowlist, capability gate, trust tier, input validation — applies at every trust level, unchanged. Only after all five pass does the conversation's guard profile decide whether the call runs or waits for approval — the guard can never grant anything the arms above refused — and the per-tool timeout (bounded by the package's trust tier) is resolved last. The guard engine itself never reads the call's arguments. The one declared refinement is presence-based and can only ever ADD an approval prompt: a tool may list root input key names in x-neuralis.askOnKeys, and a call that sends one of them asks in the asking profiles while a call that sends none stays quiet — read only from a declaration the host itself stamped, so a remote MCP server can never narrow the asking about its own tool. A non-object argument payload asks. The runtime then validates the input against the tool's JSON schema and runs the PreToolUse hooks in priority order before dispatch; a hook that REPLACES the input has the same posture re-evaluated on what it substituted, so it cannot introduce a key the guard was never asked about. Tools the caller's features do not cover are absent from the model's catalog entirely, and a hallucinated call to one fails with a generic not-available error. Approval replays re-validate against the originally persisted input, so an approver cannot be tricked into approving different arguments than the ones reviewed.

What a package file may do to the prompt

A package or a synced source contributes markdown that is assembled into the system prompt inside labelled containers — one for each injected file, one for the package overview, one for an agent's identity. Those containers are built by concatenation, so the platform treats every value it did not author as data rather than structure.

Two rules make that real. A body may not close the container it sits in: any occurrence of a container's own tag inside injected text — a rule file's body, a package or source name, a tool description, a file label, an agent handle — is visibly folded before assembly, so a file cannot end its own block and have the rest of it read as prompt-level instructions, and cannot forge a second block of the same kind. The fold is tag-specific by design: text that spells out some other container's tag stays visible, sitting inside its own correctly-closed block, where the model reads it as the contributed content it is. Every injected value is bounded: attributes are escaped, and names, paths and descriptions are capped, so a contribution cannot spend an unbounded share of the context window either. Identifiers the model is meant to pass back — a tool name, a command name, an agent's delegate handle — are the one exception to capping: they are never truncated, because half an identifier is not a shorter identifier, it is a wrong one. Above a far higher ceiling such a value is dropped and the drop is stated, never silently cut.

The rules apply to every contribution regardless of trust tier, including first-party ones. They are a floor under the gates rather than a replacement for them: feature visibility, the caller's own package toggles, scope filtering and the per-file size cap all still decide whether a file is injected at all.

On this page