Sources and connectors
Source kinds, the URI vocabulary, and how files resolve across environments.
A source is a mounted data location: a project's data folder, an extra
host directory, the vector memory, a machine sandbox. Each source has a slug,
and that slug is a URI scheme — every file in the platform is addressed as
<source>://<path>. A connector is the driver behind a source kind; the
runtime contract is described in
the connector port. This page covers the
vocabulary brain-core ships and how source configuration behaves in practice.
Govern external data before connecting it
Sources can hold an organisation's shared corpus or one user's private data. Before copying or syncing, establish the dataset owner and approved location, intended audience and source scope, sensitivity and retention, credential scope, copy permission, and provenance/refresh record. If copying is not authorised, do not create or index the source. Public availability does not by itself permit collecting sensitive personal data, and external text is data rather than trusted instruction.
Corpus content belongs only behind a configured source/connector with explicit policy and sync lifecycle. A writable project-data folder is not automatically a governed corpus. Agent identity and durable memory must not contain dataset rows or shadow copies; memory can retain a compact canonical source-URI pointer with owner/scope, provenance, permitted use, refresh date and governing retention/policy reference. Packages add skills, deterministic tools and bounded workflows around those source URIs instead of embedding the data in package prose.
The coordinate four-tuple
Every file artifact lives at the coordinate
(uri, connector, os_uri?, container_dir?):
uri—<source>://<path>, the stable identity used by tools, routes, the UI, and the vector index.connector— the connector instance bound to that source slug at boot.os_uri— optional full host-OS path, produced byconnector.resolveOsUri(uri). When the platform runs inside a container and no host mirror path is configured, the connector throws a structuredOsUriUnresolvableErrorinstead of fabricating a path — the UI and the vector index persistos_uri: undefinedrather than a stale guess.container_dir— optional container-side path counterpart, used when a shell or skill script needs to execute against the file.
The host-versus-container split is why the local connector's config has two
fields: root (the path the server actually reads, required) and hostRoot
(the host-side mirror when the server runs in Docker, optional). The source
listing shows these paths only to someone who could attach that source or can
read its root, and the last sync's error text only to someone who may run syncs;
everyone else sees the connector kind and a generic "the last sync reported
errors".
URI vocabulary
| Scheme | Connector kind | What it addresses |
|---|---|---|
data:// | local | The project's data zone — the default working area for agents and uploads |
brain:// | brain | Vector-only memory backed by the index, agent-scoped by default |
packages:// | local | The project's installed package directories |
app:// | local | The platform app zone — a privileged mount (drive.mount.privileged); agent writes are denied |
machine:// | webtop | A machine sandbox, contributed by machine-core |
| custom slugs | any | Every additional source you attach — the slug you choose becomes the scheme (repo://, notes-alex://) |
Connector kinds shipped by brain-core
local (LocalDiskConnector) is the full-capability on-disk source:
list, read, stat, write, delete, move, mkdir, rmdir,
scanDelta, walk, count, peekCount, scanContent, subscribeEvents,
resolveOsUri — every capability in the port except exec and its
execDetached companion. stat answers "is this still there" without
transferring the file. It blocks
symlink traversal and .. escape and refuses explicit access to the platform
app zone. Code execution is NOT a connector capability — shell/skill scripts run
through the execute tool's shellRunner, gated per-path by the uri-policy
exec permission, not by the connector. Sync excludes keep secrets, lockfiles,
.env files, key material, dotfile credential stores, browser profiles, and
build artifacts out of the index — the credential shapes on a runtime floor that
also covers sources attached earlier (for what gets indexed next — it does not
remove what was embedded before it shipped), the rest as an editable per-source
default.
brain (BrainInternalConnector) is the Qdrant-backed internal memory.
It is discoverable like any other kind, but reads and writes are serviced by
the filesystem service against the vector index — there is no on-disk tree
behind it, and no os_uri. Viewer roles are read-only on brain:// by
default.
Other packages contribute further kinds through their manifests — for
example the webtop kind from machine-core (addressed as machine://). The
registry, capability checks, and the trust gate are the same for every kind.
Source scope and naming
A source belongs to a project, a user, or an agent. Each connector
kind declares which scopes it allows (allowedScopes) and the default;
project-scoped sources are visible to every project member, while user- and
agent-scoped sources require ownership or a scoped-observation feature grant.
Source slugs must be unique because they are URI schemes. The Files UI
suggests scope-qualified keys when you attach a source: a project source named
after its folder (repo), a user source qualified by the user (repo-alex),
an agent source qualified by the agent (repo-nova). Never reuse one slug
across scopes — the scheme is the identity.
Declarative sources and package attribution
Packages don't just register connector kinds — they can also declare the
concrete source instances they own, right in their manifest. A declaration
names the slug, the connector kind, the root, the scope, and a required
one-line description, and it marks the source as either auto (created in
every project at initialization) or discoverable (surfaced for on-demand
attachment). This is how the project data zone, the package directory, and the
vector memory exist in a fresh project without anyone attaching them by hand,
and how the platform app zone is offered as a privileged add — gated by
drive.mount.privileged, which no role below the admin tier holds by default.
Because the instance is declared, every source carries a contract-backed
owning package — its origin_package. The platform resolves it by matching
a live source against the declarations, and the match is deliberately
unspoofable: a rooted on-disk source must match BOTH the declared slug AND the
resolved root, so a user mount that merely borrows a well-known slug (say,
naming a mount app while pointing it somewhere harmless) is never mistaken
for the package-owned source. Rootless sandbox kinds (vector memory, machine
sandboxes) match by kind alone, since their slugs are per-user.
origin_package appears on the sources listing and on each source's
runtime-stack row, where the agent sees it as a [pkg: …] tag — so the
Sources panel can group sources by the package that owns them, and an agent
can tell a first-party source from one a teammate attached. A source's
description follows the declaration too: an explicit per-source description
wins, otherwise the agent falls back to the declared description, then the
connector kind's default — so a runtime-stack row is never blank.
Enabled toggle versus URI policy
Two independent controls answer two different questions:
- URI policy answers which URIs may this caller touch. Per-source path
rules are evaluated against the caller's role and agent at four enforcement
layers — route, UI, tool (including the sandboxed shell), and the vector
index — see URI policies. Connectors are
deliberately not one of them: a connector is the raw transport for a URI
scheme, and every gate sits one layer above it. A
read: falsepath is hidden on both axes: its content is refused (read/diff/show), and its name is dropped from the file tree and search — the listing and search agree, so a denied path never leaks its existence. One authoring note about globs: a rule onfoo/**hides everything insidefoo, but not the folder nodefooitself — to hide the folder's name too, add a rule onfoo(or**/foo). - The enabled flag answers is this source live at all.
SourceConfig.enabledistrue,false, or'auto'(defaulttrue). A disabled source returns a structured error (code: 'source_disabled', withenable_sourceas the suggested next step) from everyfs_*tool and route, regardless of what the policy would allow.
'auto' exists for machine sources only: the effective state derives from a
live probe of whether the machine sandbox is currently running, so a stopped
machine reads as disabled without anyone flipping a switch. The toggle is
exposed as POST /sources/enabled with { source, scope, enabled } and as a
switch in the Files UI sources panel.
- Connector liveness answers can the connector reach its backing store right now. This is a display-only signal, separate from both the enabled flag and URI policy: when a source's connector is disconnected — a local mount that has gone missing, a sandbox container that isn't running — its row dims (icon and name) in the file tree and the sources panel, and the agent's runtime view marks it offline with a reason. Already-synced content stays listed and readable; the full colour returns once the connector reconnects.
Manifest baselines and the persisted config
When a source is first attached, the filesystem layer unions every loaded
package's matching manifest uriPolicies baselines into the source's
persisted configuration JSON. From then on the persisted config is
authoritative: owners and admins edit rules through the sources panel or the
policy routes, and the manifest baseline is only re-seeded when a source is
added or replaced.
Immutable floors are the exception, on purpose. A restrict-only floor rule
from a trusted manifest — or from a source connector's own declaration — is
re-asserted on every boot, so a floor added by a newer package version reaches
sources that already exist, and one that went missing comes back. Ordinary rules
never do; if you edited a path rule away, it stays away — and a no-op policy
edit no longer disturbs the repair bookkeeping, so a rule you deliberately
removed is not resurrected by a later restart.
A floor also decides indexing, not only reading: what cannot be read cannot
be embedded, so a path a floor denies read on is never newly embedded into
the vector store. That is one declaration covering both axes — nothing needs to
be added to a sync-exclude list, and nothing editable can undo it. The merge semantics live on
URI policies; the conversational editing
workflow is the manage-uri-policy skill on
brain-core skills.
Each source also carries an optional one-line description, editable from the
sources panel and the source routes. Agents see it in their runtime-stack
source table, so a good description directly improves how an agent picks the
right source — see runtime stack.
Discoverable sources
The "add a source" picker offers a union of two suggestion families. The first
is declared discoverable sources: a package surfaces an on-demand source —
for example the privileged platform app zone — and the declared label and
description prefill the add form. These suggestions carry their owning
origin_package. The second is host-mount suggestions: directories the
operator has bound into the runtime (the Neuralis tree, the projects zone, the
whole host, or any named mount). Those depend on a runtime mount rather than a
package, so they have no owning package. Either way, the access controls that
gate seeing a high-blast-radius suggestion and attaching a source rooted in
a protected zone remain authoritative — a declaration can tighten visibility,
never loosen it.
A suggestion may also carry defaults for what happens after the attach: a sync include allowlist, a sync trigger, and uri-policy path rules. Path rules offered this way are ordinary editable rules, never immutable floors — only a trusted package manifest can author one of those, and a create request carrying a floor marker is rejected. The distinction that matters when reading a suggestion: an include narrows what is indexed, while what an agent can read is decided by the source root and the path rules. A wide root with a narrow include is still a wide root. Where a zone offers both a narrow and a wide proposal, the narrow one is listed first.
Deleting a source
DELETE /sources/:slug first cancels any running or queued sync job for the
source (so an in-flight sync can never write entries for a configuration that
no longer exists), then removes the source-config JSON and the in-memory
connector binding. The underlying folder is never touched — brain-core
owns the wiring, not your content. Vector index points that belonged to the
source are left in place by default; pass ?purge=1 (the Files UI exposes
this as a Purge vector index entries checkbox, off by default) to remove
them. The purge runs as a background job with a single bounded index filter —
the delete call returns immediately with the job reference instead of
blocking on index size. Leftover orphan points can always be cleaned up
later through the vector health repair tools — see
memory and sync.
Sources are config, files are content
Attaching, disabling, or deleting a source never modifies the files behind it. Source operations edit configuration and index state only — content mutations go exclusively through the write tools and routes, where policies and approvals apply.