Skip to main content

Memory or context?

Orca has two kinds of text that reach an agent before it starts, told apart by one question: who wrote it. Memory is what the agent wrote about itself. Context is what you wrote about the work: the site, the repo, the brand, attached by hand to a session or an automation. They are separate stores on purpose. A context carries the authority of a human having written it, and is labelled that way in the prompt; a memory carries only the authority of an agent’s own earlier reading. Merging them would lose the distinction the agent uses to weigh one against the other. An automation has a third thing, and it is memory rather than context: one short note, rewritten by its automation agent after every firing so the next run does not start blind. It belongs to the schedule, dies with it, and is read and corrected through GET|PUT /api/automations/{id}/memory; an empty body clears it. There is no dashboard UI for it yet: it is API-only today. The note records state, never rules. The automation agent describes what the last run found and where it stopped; it never writes a condition the next run must satisfy, and a run that ended by asking a question is recorded as a run that stopped short, because nobody answers a scheduled run. The note reaches the agent as a context block, so it opens with a preamble saying it is a scribe’s observation and not operator input, and that the agent’s own instructions take precedence. The Memory Bank entries below belong to the agents the schedule runs, and outlive it.

What is the Memory Bank?

The Memory Bank is a profile-scoped store of long-lived knowledge that survives across every session of an agent. It captures preferences, facts, behaviors, and short-lived context the agent has decided to remember, and surfaces the most relevant entries back into the prompt on every run. Two delivery channels:
  • Automatic injection: every run for a profile that opts in to @memory gets a prefixed --- CONTEXT FROM MEMORY --- block built from the core tier plus query-relevant entries.
  • On-demand: agents can call memory_save, memory_recall, memory_list, or memory_delete at any point during a run.
The bank is a port of the Brain Dump relevance model. No embeddings: keyword overlap plus a small set of domain-specific clusters bridges the lexical gap between query and memory.

Storage Layout

Memories are persisted to the dedicated internal bucket (INTERNAL_S3_BUCKET, for example orca-internal), under a prefix in that bucket:
<memoryId> is mem-<8hex>, matching Orca’s other ID conventions (run-, sess-, agent-, pool-). Memory is a platform artefact and is kept separate from the tenant/artifact-tools bucket (S3_BUCKET) that save_artifact/get_artifact read and write. Only in single-bucket dev/CI setups, where INTERNAL_S3_BUCKET is unset, does the bank fall back to the primary bucket. The in-memory index is the runtime source of truth; the bucket is the persistence tier hydrated at startup via LoadAll. A successful save does not return until the object lands in storage, so a crash never loses an explicitly saved memory. When neither bucket is configured the bank still operates, degraded to in-memory only.

Memory Schema

RawInput is capped at 4 KiB. Larger inputs are rejected with a clear error.

Categories

The core tier is preference + fact entries with confidence >= 0.8, ordered by confidence then access count. Cached per profile for 5 minutes to keep the run hot path off S3.

Relevance Scoring

QueryWithRelevance ranks active memories by a composite score:

Maturity

Until the bank has accumulated query history, recency and topic dominate. The maturity factor is min(1, totalAccesses / 50): once a profile has crossed ~50 cumulative recalls, the usage weight reaches its full 0.30, and the slack from the early-bank period is redistributed back to recency and topic.

Semantic clusters

Eight built-in vocabulary groups bridge cases where a query and a memory share a topic but no literal words. A “diet” query still matches a “peanut allergy” memory through the food cluster. Clusters: food, health, work, tech, travel, entertainment, communication, shopping. A non-zero contribution requires the cluster to appear on both sides.

Staleness

Each memory carries a stalenessScore in [0, 1] recomputed on every save and access:
Category weights tilt the decay: preference=0.5, fact=0.7, behavior=1.0, context=1.5, general=1.0. The score is surfaced for inspection but does not currently gate retrieval (Brain Dump parity).

Prompt Injection

Profiles that opt in to the @memory capability receive an automatic context block prepended to every run’s prompt:
Two soft caps shape the block: After the block is rendered, the IDs that made it in are passed to MarkAccessed fire-and-forget: accessCount and lastAccessedAt update without blocking the run. Profiles without @memory (or any of the four memory_* tools by name) never receive injection. Same opt-in shape as @artifacts.

LLM Processor

When an agent saves a memory with only rawInput (no pre-structured processedContent), the bank consults a fast LLM to extract processedContent, summary, category, and confidence. The processor is a thin wrapper around Anthropic’s Messages API with a 5-second deadline. Configuration: When the key is unset or the API call fails, the extraction step is skipped and the save proceeds with whatever the caller supplied; confidence defaults to 1.0 in that case. Only when the LLM responds but the response cannot be parsed does the bank fall back to a deterministic preview: processedContent = rawInput, summary = first 50 chars, category = general, confidence = 0.5.

Memory Tools

@memory is opt-in (not in @default). Every tool requires session context and operates against the session’s profile.
Selecting @memory (or naming any memory_* tool individually) both registers the tools and enables prompt injection.

REST API

All endpoints return 503 Service Unavailable when the bank is not wired. Reads call a 2-second TTL RefreshIfStale so the dashboard sees runner-side saves promptly without paying for an S3 round-trip on every request.

Per-profile

Global

See the Conductor API reference for request and response shapes.

Configuration

The bank prefers the same internal-bucket contract used by other platform artefacts, see Storage and Files. When INTERNAL_S3_BUCKET is set it boots in fully persistent mode; otherwise it falls back to S3_BUCKET for single-bucket dev/CI setups, and if neither is set it runs in-memory only and logs a warning at startup. Both the runner and the conductor build their own bank from these variables. The runner’s bank is the source of truth on writes; the conductor’s bank is read-through synced every 2 seconds for dashboard reads.

Limits and Caveats

  • Per-profile scope only. Memories are not shared across profiles or pods, by design. Use pods and the shared filesystem for cross-agent state.
  • Listing cap. LoadAll walks at most 1000 entries per profile (the underlying ListObjectsV2 page cap). Banks larger than that need pagination plumbed through ListInput.
  • No retention enforcement. Staleness is computed but not acted on. Operators delete cold memories by hand or via the dashboard.
  • MarkAccessed is best-effort. Access count and timestamp updates persist asynchronously; a runner crash within a few seconds of a recall may lose them.
  • Source inferred is reserved for v2. All v1 saves carry source: "explicit"; future post-run extraction will populate inferred.
Last modified on September 6, 2026