Memory and sync
Vector memory, embedding, and the sync pipeline.
brain-core keeps a vector index of your sources so agents can search by
meaning, not just by filename. The index is multi-tenant — every point
carries the owning projectId, userId, and (for memory entries) agentId
— and is backed by Qdrant in production, with an in-memory substitute for
tests and degraded operation. This page explains how content becomes
searchable, when sync runs, and how to keep the index healthy.
The embedding pipeline
Embeddings are produced by a declarative model registry: every supported model is an explicit record with its provider, protocol, default dimension, supported dimensions, and credential id. Supported providers are Ollama (local), OpenAI, Gemini, Qwen, and Voyage, plus a deterministic offline embedder for tests. Two carried models are multimodal — Gemini Embedding 2 (images, PDF pages, audio and video) and Voyage's multimodal model (images and PDF pages) — while every other carried model is text-only. Models that tell queries and documents apart are told which one each call is.
Embedding has its own credentials, separate from the chat keys:
embedding.openai, embedding.gemini, embedding.qwen and embedding.voyage,
kept in the credential store — packages never
read embedding secrets from the process environment. They are read at the
platform (global) scope only: indexing is background work with no user in
context, and one collection serves every project, so a member's own value for
the same id is never consulted. Because they are not LLM keys, the
per-credential use limits count them, and the platform meters embedding spend
in dollars per call: embeddingDailyUsdCap stops embedding for the rest of the
UTC day once today's priced spend reaches it (sync pauses with a resumable
error, semantic search falls back to full-text; 0, the default, means no cap).
Any server that speaks the OpenAI /v1/embeddings protocol can be added as an
embedding endpoint — from the admin Vector section or during setup: its
URL, its models with the dimension each returns (and optionally a price), its
authentication, and its address class. A keyed endpoint's key is the
credential embedding.endpoint.<id>, never part of the entry, and its models
appear in the model list as endpoint:<id>:<model>. An endpoint whose model is
the active one, the target, or a running rebuild's target cannot be removed,
disabled or renamed (409 endpoint_in_use, nothing written) — switch the index
away from it first. A public endpoint is
address-checked and IP-pinned on every call, exactly like the cloud providers;
private is the operator's explicit statement that the address is inside
their own network. The two cloud base-URL overrides (openaiBaseUrl,
qwenEmbeddingBaseUrl) accept public addresses only — a proxy or gateway
inside your network is added as a private embedding endpoint instead.
Many models expose MRL (Matryoshka) selectable dimensions: their
supportedDimensions list (or range) lets you pick a smaller vector — trading
a little quality for less storage and faster search.
Three properties make the pipeline safe to operate:
- One active model, recorded. The model that built the collection is
recorded with it and stays the active model: every write and every
search embeds with it, and every vector written is stamped with the model that made
it, so vectors from two models never mix in one index. The configured
embeddingModelId/embeddingDimensionis only the target. Changing it starts nothing: the admin Vector section reports the switch as pending — how many points sit on the active model, the target, and on request a cost and time estimate computed from a sample (a range, not a quote) — and waits. - A model switch is an explicit rebuild. An owner starts it from the
Vector section (or the
rebuild-generationcollection action below). The target model embeds eligible points again from stored text, or from source bytes when inline content is unavailable or the input is media, into a new collection while the old one keeps serving. Retained recovery content, includingbrain://content and its version history, is preserved. An already recovery-only record whose input is unavailable stays preserved without vectors in the new collection, including during a source outage; verified input restores its searchable embedding. Unmarked missing input still pauses the rebuild. Writes and deletes that happen meanwhile are applied to both, and when the new collection is complete the two are swapped in one step; the dimension may change freely. A rebuild can be cancelled (the new collection is dropped, the active one is untouched), and one interrupted by a restart waits, paused, until it is resumed. The previous collection is dropped after the swap unlessvectorKeepPreviousGenerationkeeps it as a rollback copy (at double the storage). While it runs, storage is doubled and every new write is embedded twice. - No silent fallback. A missing credential raises
EMBEDDING_CREDENTIAL_MISSINGnaming the credential id and its scope, and an unsupported dimension raisesUNSUPPORTED_EMBEDDING_DIMENSION. The platform never swaps in a different provider or another key behind your back. Nothing is embedded without the active model's key; semantic search then answers with full-text results and awarningthat names the missing credential, rather than failing the call.
pnpm reset:vector (and the admin Vector section's confirmed reset) is not
the way to change models: it drops the collections and the model record, and
with them every brain:// file and its version history, which live only in the
vector store. Connector-backed sources come back on their next sync, embedded
again at provider cost. Keep it as a last resort.
Each indexed file produces a parent document point plus content chunks, so semantic search can rerank chunk hits against document hits. The parent keeps only the file's first 200k indexed characters; the chunks hold the rest, which is where every search mode reaches it. Chunks are skipped by content hash when unchanged, which keeps re-syncs cheap.
That skip is also why the per-source controls are three different things.
Rescan changes (sources/reindex) walks the whole source again, but a file
whose content did not change keeps its vectors. Forget sync state
(sources/reset) only clears what the last sync saw and cancels a running job,
so the next sync compares every file — it deletes no file and no vector.
Re-embed… is the one that spends: it reads every file in its scope — a
whole source, a folder or one file — and embeds it again with the active model,
changed or not. It shows a sample-based cost and time estimate first and starts
only on confirmation; it is offered in the Sources panel, on the Files tree's
context menu, and to agents through the manage-sources skill, and it needs
drive.sync like every sync control. Re-embed is not offered on brain://,
whose vectors are the content itself.
Media. With a multimodal active model, sync also embeds media from the
file's own bytes: images, PDFs (one embedding per rendered page, up to
brainSyncMaxPdfPages, default 20), and — only where you switch them on —
video and audio. Each source chooses its kinds in its sync settings
(sync.media {images, pdf, video, audio}; images and PDFs default on, video
and audio off), and a file larger than brainSyncMaxMediaMb (default 8 MB) is
skipped. A kind is embedded only when the ACTIVE model supports it, so on a
text-only model nothing is sent and the Sources panel says how many media
files were not embedded and why. Structurally invalid PNGs are skipped locally before provider work and counted as invalid when the media lane is invoked; existing indexed and recovery content is preserved. This is a structural preflight, not a pixel-decoder guarantee. PDF pages are rendered from local-disk
sources; embedding endpoints (the OpenAI /v1/embeddings protocol has no media
input) embed text only. A media hit in search carries its kind and, for a PDF,
its page.
When content is indexed
Writes that go through brain-core (fs_write, uploads) update
the vector index inline — the file is searchable as soon as the call
returns. The background sync exists for everything else: files changed on
disk by shells, skills, external editors, or other processes.
Inline writes honor the same exclude/include filters as sync: a file an
agent creates under an excluded path (say inside node_modules) is written to
disk and tracked for review, but it is not embedded — exactly as the sync
walk would skip it. So what is searchable is decided by one consistent rule, no
matter how a file arrives. Excluded files are still fully readable and editable
through the connector; they simply stay out of the meaning index.
Each source declares its sync trigger in its configuration:
| Trigger | Behavior |
|---|---|
manual | Sync only when explicitly triggered (UI button, route, or skill) |
auto | Event-driven: syncs on the first run, when the filter configuration changes, and on a periodic full-reconcile cadence (reconcileEvery, default 24h) — never a full tree walk on every tick |
interval:<N><s|m|h|d> | Sync at most once per interval (for example interval:15m) |
The legacy on_change trigger (which re-walked the whole tree every
scheduler tick) was replaced by auto; existing configurations migrate
automatically when read.
auto sources are watched at the operating-system level: any change —
an agent write, a shell command, a git checkout, an external editor —
marks the touched files dirty, and after a short debounce a targeted
sync reconciles exactly those files (read, re-embed if changed, tombstone
if deleted, move detection for renames). Excluded trees are ignored at
the watcher level too, so they cost no OS watch resources. If the
watcher ever loses events, the source falls back to a full reconcile
rather than trusting the stream.
Filters are two lists. Exclude patterns (gitignore syntax) always apply
on top of the built-in defaults. An optional include allowlist narrows
the source further: when non-empty, only files matching an include pattern
sync — and an exclude match still wins, so an excluded path can never be
re-included. Either edit changes the configuration hash and forces a
reconcile pass on the next run, and every full sync removes the index
search data of files the filters no longer cover. Pattern/media exclusions remove obsolete parent/chunk search data, including modern parentless chunks in the same project and source. Chunk cleanup has its own count and refreshes Files without inflating removed-file counts; legacy chunks without source ancestry need explicit health maintenance. Recovery-bearing parents — pending/deleted records, saved baselines and sole-copy brain:// content — retain their complete payload with recovery_only:true, no embedding stamp and no dense or BM25 vectors. Retrieval, browsing and review remain policy-gated; search, enrichment and source-package discovery exclude retained recovery points. Re-inclusion produces a real new embedding before clearing the marker. Enabled/auto inactivity never purges an index, and an immutable READ floor alone never authorizes cleanup. Generation rebuilding classifies current source policy before provider work and again at its config-coordinated swap: unknown provenance or unavailable/invalid configuration pauses without swapping. The dollar estimate samples only eligible content and exposes protected/excluded/unresolved counts; it remains a range, not a hard spending bound. A new project's
data source starts with the platform's numbers-only machine-written files
(usage counters) already excluded; an existing project's filters are never
rewritten. Sources also offers these defaults when enabling sync on its own data source with no sync block; Save applies the editable proposal. Custom sources and explicit-empty configured lists keep their choices.
Every source walk — for indexing and for the displayed file count alike —
stops at an effective scan cap: a per-source Max scan files override,
falling back to a platform-wide default (200 000), so a very large source
never walks unbounded. The displayed connector count applies the same
exclude set as the sync, so it counts the files actually eligible for
indexing rather than the whole tree. When a walk reaches the cap the count
is shown as N+ (the source genuinely holds more than the cap allows) —
raise the cap to index and count beyond it.
A platform-wide auto-sync mode (admin setting) controls the background
scheduler: on (default), off (manual sync only), or packages-only —
in which only sources pinned always active keep auto-syncing
(package-content sources are pinned by default). Manual sync is available
in every mode.
A background scheduler polls source configurations every 60 seconds (tunable
via BRAIN_SYNC_POLL_INTERVAL_MS) and enqueues due jobs on a shared sync
queue. The queue is the single execution path for every sync-shaped
operation — scheduled runs, the UI Sync button, reindex, and delete-time
purges all go through it. It runs a bounded number of jobs at a time
(default two), rotates fairly across projects so one large source can never
starve another project's sync, deduplicates repeat requests per source, and
gives every job its own cancellation handle: triggering a sync returns a job
reference immediately (HTTP 202), and cancelling stops the run within
moments — mid-walk or mid-embedding — instead of waiting for it to finish. A
cancelled run is not an error; everything indexed before the cancel stays
valid and the next run converges. The runner walks the
connector's changed-files-since-checkpoint delta (the checkpoint token from
the previous successful run is persisted and passed back, so nothing slips
through the gap between scan and completion), embeds new and changed
files, removes deleted points, and detects renames by exact content-hash
match: a moved file keeps its index entry and every chunk embedding, and only
its file-level vector is refreshed when the name changed. An append that
leaves the first 200k characters as they were re-embeds only the new tail.
Default excludes keep
node_modules, package-manager content stores (.pnpm-store, .yarn/cache),
build output, lockfiles, .env files, and key material out of the index, and per-file caps (5 MB, 200k indexed characters) bound the
embedding cost. Append-shaped files (.jsonl, .ndjson — conversation
transcripts, event logs) are treated differently, because they only ever grow
at the end: they have their own, higher size limit (brainSyncMaxAppendFileMb,
20 MB, never above the read limit), they are never cut at the per-file chunk cap, and an append embeds only
the new tail — the platform checks that the stored file is an exact prefix of
the new one before it trusts that, and re-reads every chunk hash otherwise.
Some of those patterns are more than a default. A runtime sync floor is
unioned under every source's own exclude list on every check, so it also covers
sources that were attached before a pattern shipped and an include rule can
never bring a floored path back. Like every exclude pattern, it also clears
what was already embedded under it on the next full sync. It holds the dotfile credential stores and
browser profiles — .ssh, .gnupg, .aws, .azure, .kube, .docker,
.netrc, .git-credentials, .npmrc, .pypirc, the login keyrings, the NSS
certificate and private-key store in both of its locations (.pki and
.local/share/pki), the gcloud / gh / rclone config trees, and the
Chromium, Chrome and Firefox
profiles — written as path shapes rather than as one connector's paths, because
a home directory synced through a local source and a sandbox /config synced
through a virtual desktop are the same exposure.
What cannot be read cannot be embedded. Beyond that pattern list, indexing
consults the source's immutable
uri-policy floors directly: a path whose
floor denies read is never newly embedded, on every source, with no exclude
entry to add and none that can be edited away. The two inputs are different
kinds of thing on purpose — the exclude list is a preference you own and can
change at any time, while a floor is a security declaration that belongs to
the package owning the connector. That is why there is no package-declared
"sync-exclude" contract: deriving the index rule from the read floor covers
every floored path automatically.
Two boundaries worth knowing. A per-role denial does not keep a file out of the index — somebody can read it, so it stays searchable and each caller's own results are filtered for them; this is why an owner still finds conversation content a member cannot. Pattern cleanup preserves recovery-bearing history while clearing search vectors; an immutable READ floor hides content and blocks new indexing without alone purging old records.
Conversation summaries are a deliberate part of these defaults. A conversation's
active SUMMARY.md is indexed like any other file, so an agent can search
what it learned by meaning. The immutable summary revision archive behind it,
however, is on the same floor — historical revisions are never embedded or
searchable, on new and pre-existing projects alike.
After each successful run the source stores an absolute index snapshot (indexed, unindexed, and total file counts) that the UI and the agent's runtime stack read cheaply — no live directory walk is needed to display counts.
Contribution categories at sync time
The indexer assigns every file a category payload field. Beyond the
runtime-data categories (conversations, usage logs), files whose path matches
the shared contribution layout get a first-class contribution category — the
same vocabulary packages use:
| Path pattern | Category |
|---|---|
SKILL.md (any directory) + commands/*.md (legacy single-file form) | skill |
instructions/**/*.md | instruction |
AGENTS.md, CLAUDE.md, CLAUDE.local.md, GEMINI.md, .github/copilot-instructions.md | instruction (identity form) |
rules/**/*.md, .cursor/rules/*.mdc, .cursorrules, .windsurfrules | rule |
agents/**/*.md, *.agent.md | agent |
docs/**/*.md (READMEs and templates excluded) | docs |
package.json, neuralis.package.json, known plugin manifests | package |
.mcp.json, hooks.json | connector / hook (indexed only) |
These patterns also match inside well-known tool directories (.claude,
.cursor, .codex, .gemini, .github, and friends). A markdown keeps its
contribution category only when it is admitted: the filename-identity
forms (SKILL.md, commands/*.md, the identity basenames, .cursorrules,
.mdc) always are; a markdown under agents/, rules/, instructions/ or
docs/ and a *.agent.md need frontmatter id or name, otherwise the
file is indexed as a plain file. A free-form category a caller sets on a
file (note, config) is kept as given; a contribution category is always
derived from the path and the content, never taken from a request, and a
stale one is re-derived on the next full sync. Admitted markdown
contribution files additionally carry a whitelisted frontmatter capture
(id, name, title, description, required features, credential ids), manifests carry a
usability flag when their content declares a recognized package shape, and
files under an agent's private agent-core/<agentId>/ subtree carry an
owner marker. Unchanged files pick these fields up through a payload-only
refresh on the next sync pass — no re-embedding. This category layer is the
discovery substrate for
source packages — skills, rules,
instructions, agents, and docs loaded from synced sources.
Search semantics
fs_search has three modes, chosen by search_mode:
semantic— meaning-based vector search over the index. Requires a warm index; results carry relevance scores. On Qdrant the query is hybrid: the vector ranking is fused with a keyword leg over the same content, ranked by BM25. A fresh install creates its collection with BM25 ranking from the first file (each indexed write also runs one BM25 encoding on the vector server; this needs Qdrant 1.15.2 or later); an older collection gets it when an owner enables text ranking (below). A file not yet covered by it still matches through the plain keyword filter. A search whose scope — your files in a project, or one folder of them — holds at mostvectorExactSearchMaxPointsindexed points (20 000 by default) compares the query with every one of them instead of walking the approximate index, so a small folder finds everything it holds; each search counts its scope first to decide.filter— metadata and keyword queries: category, tags, status, and creation-date ranges, plusqueryas a tokenized, case-insensitive keyword over indexed content — the file's chunks included, one result per file — (results ranked by term frequency) andglobas a filename keyword match. No embedding involved.grep— literal text search (ripgrep, git grep, or grep — whichever is available) over live connector content. Bypasses the index entirely, so it always reflects the current state of disk.
Across all three modes, fs_search enforces the per-URI read policy on
every result, not just the drive.search feature gate — a read:false
path (another agent's private memory subtree, for example) is silently absent
from results, and a fully-denied folder_uri returns a 403, giving search the
same path protection as fs_read and fs_write. This makes grep the
uri-policy-read-gated counterpart to shell grep: a caller holding
drive.search plus read policy on a path can grep it without any
execute / shell access.
Reads and listings are connector-first: fs_read and fs_list take live
content and tree structure from the connector even when the index is cold,
then reconcile index presence per listed file. Presence states distinguish
vector-only entries (brain), files where index and connector agree
(synced), on-disk files not yet indexed (unindexed), and indexed files no
longer present on the connector (stale_indexed).
brain:// memory queries are isolated per agent inside a project: an agent
searches its own memory unless sibling inclusion is explicitly requested, and
cross-agent observation requires the corresponding feature grant.
Vector health and degraded mode
If Qdrant is unreachable when the package initializes, brain-core wires an
in-memory index instead and the platform stays up: connector-first reads,
listings, and grep search keep working, while semantic search and
sync operate against the in-memory substitute. On Qdrant-backed deployments a
background watch re-probes connectivity every 30 seconds, logs loss and
recovery transitions, and reports { backend, qdrantReachable, lastCheckedAt }
into the host's health endpoint. The live backend is not hot-swapped — after
Qdrant recovers, a restart re-enables Qdrant-backed search.
The vector health service is the project-scoped janitor behind the Files UI health panel. It separates cheap from expensive work so the panel never slows the rest of the UI:
| Operation | Cost | Effect |
|---|---|---|
| Stats | Cheap (exact counts, no scan) | Total / document / chunk / deleted counts, per source — computed from indexed fields, so they stay instant at millions of points |
| Scan for problems | Expensive (walks every point) | Orphan-chunk and duplicate-document detection — explicit, on-demand only |
| Garbage collect | Destructive | Drops records marked deleted that are older than the retention window, except those a pending change still needs for its revert |
| Repair orphans | Destructive | Drops chunks whose parent document is missing or flagged deleted |
| Find duplicates | Read-only | Report of documents sharing a content hash |
The stats load when you open the panel; the integrity scan runs only when you ask. Garbage collection and repair are destructive and support a dry-run flag; the duplicates report never removes anything.
Garbage collection also runs on its own, once a day by default
(vectorGcIntervalHours, 0 turns the schedule off), for every project as
explicit system work — never under a user's identity, and whether or not
automatic sync is on — and every run is logged with who started it. The
schedule's clock starts when the server starts, so the first automatic run
comes one interval after a restart, never at boot; running it by hand is always
immediate. Deleted records a still-unresolved pending change refers to
are kept, so reverting that change keeps working. Garbage collection and
orphan repair never run twice at once for the same project: a second request
while one is running answers 409 action_in_progress, and an automatic run
that meets a manual one skips that project and tries again on the next tick.
The index also watches its own writes. Every write is counted per file over a
rolling hour, within a bounded list per project, and a file rewritten
vectorWriteHotUriThreshold times or more in an hour is logged as a hot file — the
signature of an automated writer that rewrites instead of appending. Members
who may change files (drive.write) can list the hottest files through
GET health/hot-uris; the list is filtered for the caller exactly like a file
listing, so it names only files that caller may read.
One vector collection serves every project, so the work on the collection
itself belongs to the owner, in a maintenance window: health/collection
requires the platform feature platform.vector.maintain, which no role holds
by default. GET reports what the server has — missing or differently
configured indexes, quantization, text-ranking coverage — and POST runs one
action: rebuild the payload indexes with their declared settings (only the
fields vector searches filter on get extra graph structure, which keeps an index
rebuild cheap), apply the segment settings (vectorDefaultSegmentNumber,
vectorMaxSegmentSizeKb, vectorMaxIndexingThreads), apply
vectorQuantizationMode, or add the BM25 text-ranking vector to a collection
created before it was the default and backfill it in bounded, resumable batches. Each of these
rebuilds index structures on the server; none runs at startup. The same surface
carries the model switch: GET also reports the active model, the target, the
point counts, a running rebuild's progress, the last estimate (dropped once the
target model changes) and today's
embedding spend, and the actions rebuild-estimate (a sample-based estimate,
changes nothing), rebuild-generation (starts the rebuild, or resumes a paused
one) and rebuild-cancel run it — the admin Vector section is the same surface
with buttons. Startup creates
indexes that are missing, rebuilds a keyword-search index only when its
tokenizer settings are wrong, and otherwise just checks and logs what differs. Startup also compares the Qdrant server version
with the client the platform ships and warns when they are more than one minor
version apart — the admin Vector tab shows the same.
Qdrant does not log its own background optimization, so the platform reports
it: every 30 seconds it reads the server's list of optimizer runs and writes one
vector.optimizer.completed line per finished run (kind, points, segments and
the duration the server measured), plus a warning for a run still going after
ten minutes — in the container log as well as the vector log. The Qdrant
the setup runs — the generated container or the downloaded binary — also has
its anonymous usage telemetry turned off.
The pending-change journal
Every revertible filesystem change an agent makes is recorded in a durable, append-only journal (with a baseline copy of the prior content alongside it). This journal is the source of truth behind the Changes tab: listing, approving, and reverting a change all consult it directly and survive a restart, including files that were never indexed. Listing can fall back to the journal during an index outage; review mutations refuse backend errors rather than treating them as missing state.
Listing pending changes reads two sources in parallel — the durable journal
and a query for files awaiting review in the vector index — and merges them,
preferring the journal. If the vector index is unreachable, the list falls back
to the journal alone rather than failing. Review uses a fresh POST /pending {action:"preview",id} receipt, separating current connector bytes from the recorded snapshot and saved baseline. The preview reports missing/unavailable/binary/too-large content and missing/stored/stale/recovery-only index states within pendingPreviewMaxBytes (default 262144). Approve (POST /approve) and Revert/Dismiss (POST /pending) carry expectedFingerprint; states requiring acknowledgement additionally need acknowledgeCurrentState:true after review. Freshness, tenant, owner, source scope and URI READ/WRITE gates run again at apply. Missing state is never fabricated, backend errors never become absence, and acknowledgement grants no permission. Approve/Dismiss preserve the stored snapshot and index state, including a genuinely absent index; recovery-only metadata uses explicit vector-clearing retention. Revert uses the saved baseline, refuses missing baselines and occupied destinations, and gates both move coordinates. Legacy unsafe pending /write revert:true requests refuse; ordinary non-pending one-level undo remains available.
The journal is never vectorized: it lives in an excluded path, so sync skips it, and it is readable only by owners and admins. Each entry is re-checked against per-file path policy and per-user ownership when it is listed, so a user only ever sees their own changes (unless their role carries the scoped-observation feature).
Source delete and vector purge
Deleting a source removes its configuration and connector binding but leaves its vector points in place by default — they become orphans whose parent files are no longer synced. Pass the purge option on delete to remove the source's points in the same operation, or run the orphan repair afterwards. The full source-deletion contract is on sources and connectors.
Memory entries are files
brain:// is a source like any other: agents write memory with fs_write
create, refine it with an edit, and retrieve it with fs_search in semantic
mode. There is no separate memory API to learn — the URI vocabulary covers
durable notes, insights, and working state.