Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions apps/memos-local-plugin/adapters/hermes/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,36 @@ keepalive, reconnect generation, and host callback dispatch. All algorithm
logic (L1/L2/L3, skills, retrieval, feedback, decision repair) remains in the
shared TypeScript core.

## Desktop and custom Unix installations

Use the backend source directory and Python environment used by the desktop
app, which may differ from the CLI installation. The Unix installer validates
`hermes_cli` and `plugins.memory.load_memory_provider` before stopping Hermes
or deploying the plugin. It checks literal Python/Bash launchers and the default
backend's `venv` and `.venv` directories.

If auto-detection fails, use the updated `install.sh` with explicit paths:

```bash
HERMES_INSTALL_DIR="/actual/path/to/hermes-agent" \
HERMES_PYTHON="/actual/path/to/hermes-agent/venv/bin/python" \
bash install.sh --agent hermes
```

`HERMES_INSTALL_DIR` is the backend source directory containing `hermes_cli`
and `plugins/memory`, not simply the `.app` bundle. `HERMES_PYTHON` must be the
interpreter that runs that backend. Explicit paths fail with diagnostic output
instead of falling back to another installation. Paths containing spaces work
when quoted. For a non-default data/config directory, also set `HERMES_HOME`;
it defaults to `~/.hermes`. The MemOS package/data location remains
`~/.hermes/memos-plugin` for compatibility with existing installations.

If the chosen interpreter cannot import the host's memory provider API, repair
or update that Hermes environment. Creating an empty `plugins/memory` directory
does not supply the missing API. Restart the desktop app after installation.
Desktop distributions with unrecognized layouts require explicit paths; these
options do not imply that every desktop release has been tested.

## Protocol surface

The adapter calls the following methods on the bridge:
Expand Down
8 changes: 5 additions & 3 deletions apps/memos-local-plugin/core/config/defaults.ts
Original file line number Diff line number Diff line change
Expand Up @@ -222,6 +222,8 @@ export const DEFAULT_CONFIG: ResolvedConfig = {
// an early-life install can still cluster into a world model;
// strict 0.6 starved L3 in real usage.
clusterMinSimilarity: 0.3,
maxPoliciesPerCluster: 20,
maxPromptChars: 32_000,
policyCharCap: 800,
traceCharCap: 500,
traceEvidencePerPolicy: 1,
Expand Down Expand Up @@ -249,9 +251,9 @@ export const DEFAULT_CONFIG: ResolvedConfig = {
// real usage; 1 lets the candidate→active transition happen
// immediately on first successful invocation.
candidateTrials: 1,
// Lowered from 6 hours → 0: no cooldown, skills can re-evolve
// as soon as new evidence arrives.
cooldownMs: 0,
// Verification failures are retried after six hours by default;
// operators may set this to 0 when immediate re-evaluation is desired.
cooldownMs: 6 * 60 * 60 * 1000,
traceCharCap: 500,
evidenceLimit: 6,
useLlm: true,
Expand Down
4 changes: 4 additions & 0 deletions apps/memos-local-plugin/core/config/schema.ts
Original file line number Diff line number Diff line change
Expand Up @@ -326,6 +326,10 @@ const AlgorithmSchema = Type.Object({
* are ignored (policies too disparate to share a world model).
*/
clusterMinSimilarity: NumberInRange(0.6, 0, 1),
/** Maximum policies included in one L3 abstraction prompt. */
maxPoliciesPerCluster: NumberInRange(20, 1, 100),
/** Hard total character cap for one L3 abstraction prompt. */
maxPromptChars: NumberInRange(32_000, 4_000, 128_000),
/** Chars of L2 body handed to `l3.abstraction`. */
policyCharCap: NumberInRange(800, 200, 4_000),
/** Chars of trace body handed per evidence trace. */
Expand Down
39 changes: 27 additions & 12 deletions apps/memos-local-plugin/core/llm/client.ts
Original file line number Diff line number Diff line change
Expand Up @@ -282,6 +282,7 @@ export function createLlmClientWithProvider(
maxTokens: opts?.maxTokens ?? config.maxTokens ?? DEFAULT_MAX_TOKENS,
jsonMode,
stop: opts?.stop,
op: opts?.op,
};
}

Expand Down Expand Up @@ -355,37 +356,33 @@ export function createLlmClientWithProvider(
notifyError: true,
});
} catch (hostErr) {
const normalizedHostErr = normalizeError(hostErr, ERROR_CODES.LLM_UNAVAILABLE, "host fallback failed");
failures++;
const failAt = markFail(hostErr);
const failAt = markFail(normalizedHostErr);
facadeLog.error("host.fallback_failed", {
primary: summarizeErr(err),
host: summarizeErr(hostErr),
host: summarizeErr(normalizedHostErr),
});
// Primary AND host bridge both failed. Trip on a terminal
// primary error (the one the operator typically needs to fix
// — host bridge failures are usually transient stdio issues).
if (breakerIsTerminal(err)) breakerTrip(err);
notifyOnError(hostErr);
notifyOnError(normalizedHostErr);
notifyStatus({
status: "error",
provider: provider.name,
model: config.model,
message: summarizeErrMessage(hostErr),
code: hostErr instanceof MemosError ? hostErr.code : undefined,
...extractRetryDiagnostics(hostErr instanceof MemosError ? hostErr.details : undefined),
message: summarizeErrMessage(normalizedHostErr),
code: normalizedHostErr.code,
...extractRetryDiagnostics(normalizedHostErr.details),
at: failAt,
durationMs: Date.now() - startedAt,
fallbackProvider: "host",
op,
episodeId: opts?.episodeId,
phase: opts?.phase,
});
throw hostErr instanceof MemosError
? hostErr
: new MemosError(
ERROR_CODES.LLM_UNAVAILABLE,
`host fallback failed: ${(hostErr as Error).message ?? String(hostErr)}`,
);
throw normalizedHostErr;
}
}
failures++;
Expand Down Expand Up @@ -840,3 +837,21 @@ function summarizeErrMessage(e: unknown): string {
if (e instanceof Error) return e.message;
return String(e);
}

function normalizeError(
err: unknown,
fallbackCode: (typeof ERROR_CODES)[keyof typeof ERROR_CODES],
prefix: string,
): MemosError {
if (err instanceof MemosError) return err;
if (err instanceof Error) return new MemosError(fallbackCode, `${prefix}: ${err.message}`);
if (typeof err === "object" && err !== null) {
const record = err as { code?: unknown; message?: unknown; data?: unknown };
const code = typeof record.code === "string" ? record.code : fallbackCode;
const message = typeof record.message === "string" ? record.message : String(record.data ?? err);
return new MemosError(code as (typeof ERROR_CODES)[keyof typeof ERROR_CODES], `${prefix}: ${message}`, {
bridgeError: err as Record<string, unknown>,
});
}
return new MemosError(fallbackCode, `${prefix}: ${String(err)}`);
}
20 changes: 16 additions & 4 deletions apps/memos-local-plugin/core/llm/prompts/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -50,8 +50,13 @@ export function languageSteeringLine(lang: PromptLanguage): string {
* Heuristic:
* - Count CJK Unified Ideographs (U+4E00..U+9FFF) as `zh`.
* - Count ASCII letters A-Z/a-z as `en`.
* - If CJK accounts for more than `zhRatioThreshold` of counted
* CJK+ASCII signal, pick `zh`.
* - Treat Japanese kana as an explicit non-Chinese signal. This keeps
* Japanese prompts from being mistaken for Chinese just because they
* contain a few shared Han characters.
* - CJK characters carry a small weight because technical identifiers
* (package names, commands, file paths) can contribute many ASCII
* characters inside an otherwise Chinese sentence. If weighted CJK
* accounts for more than `zhRatioThreshold` of the signal, pick `zh`.
* - Otherwise pick `en`.
*
* This intentionally treats Japanese / Korean prompts with filenames,
Expand All @@ -66,18 +71,25 @@ export function detectDominantLanguage(
samples: ReadonlyArray<string | null | undefined>,
opts: { zhRatioThreshold?: number } = {},
): PromptLanguage {
const zhRatioThreshold = opts.zhRatioThreshold ?? 0.7;
const zhRatioThreshold = opts.zhRatioThreshold ?? 0.6;
let zh = 0;
let en = 0;
let kana = 0;
for (const s of samples) {
if (!s) continue;
for (let i = 0; i < s.length; i++) {
const code = s.charCodeAt(i);
if (code >= 0x4e00 && code <= 0x9fff) zh++;
else if (
(code >= 0x3040 && code <= 0x30ff) ||
(code >= 0x31f0 && code <= 0x31ff)
) kana++;
else if ((code >= 0x41 && code <= 0x5a) || (code >= 0x61 && code <= 0x7a)) en++;
}
}
if (kana > 0) return "en";
const total = zh + en;
if (total === 0) return "en";
return zh / total > zhRatioThreshold ? "zh" : "en";
const weightedZh = zh * 4;
return weightedZh / (weightedZh + en) > zhRatioThreshold ? "zh" : "en";
}
8 changes: 8 additions & 0 deletions apps/memos-local-plugin/core/llm/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -257,6 +257,14 @@ export interface ProviderCallInput {
maxTokens: number;
jsonMode: boolean;
stop?: string[];
/**
* Logical call site (e.g. `capture.summarize`, `retrieval.filter`,
* `skill.evolve`). Forwarded from `LlmCallOptions.op` so providers
* can apply per-op behavior (request-body tweaks, routing overrides,
* reasoning kill-switches, per-op budget caps). Optional — providers
* must not assume it is set.
*/
op?: string;
}

/** What providers return — pre-facade post-processing. */
Expand Down
66 changes: 66 additions & 0 deletions apps/memos-local-plugin/core/memory/l2/induce.ts
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ import type {
EmbeddingVector,
EpisodeId,
PolicyId,
PolicyMetadata,
PolicyRow,
TraceId,
TraceRow,
Expand Down Expand Up @@ -156,6 +157,7 @@ export function buildPolicyRow(args: {
inducedBy: string; // prompt id + version
now?: number;
id?: PolicyId;
sourceSignature?: string;
}): PolicyRow {
const now = args.now ?? Date.now();
const vec = centroid(args.evidenceTraces.map((t) => t.vecSummary ?? t.vecAction ?? null));
Expand All @@ -170,16 +172,80 @@ export function buildPolicyRow(args: {
gain: 0,
status: "candidate",
sourceEpisodeIds: Array.from(new Set(args.episodeIds)),
sourceTraceIds: Array.from(new Set(args.evidenceTraces.map((trace) => trace.id))),
inducedBy: args.inducedBy,
// Fresh policy starts without learned guidance — populated by the
// decision-repair pipeline as user feedback / failure bursts arrive.
decisionGuidance: { preference: [], antiPattern: [] },
vec: vec as EmbeddingVector | null,
createdAt: now,
updatedAt: now,
metadata: derivePolicyMetadata(args.evidenceTraces, args.sourceSignature),
};
}

function derivePolicyMetadata(
traces: readonly TraceRow[],
sourceSignature?: string,
): PolicyMetadata {
const domainTags = uniqueStrings(traces.flatMap((t) => t.tags ?? []));
const toolNames = uniqueStrings(
traces.flatMap((t) => (t.toolCalls ?? []).map((c) => c.name ?? "")),
);
const errorCodes = uniqueStrings(
traces.flatMap((t) => {
const text = [
t.agentText,
t.reflection ?? "",
...(t.toolCalls ?? []).map((c) =>
typeof c.output === "string" ? c.output : "",
),
].join(" ");
return Array.from(
text.matchAll(/\b[A-Z][A-Z0-9]{2,}_[A-Z0-9_]+\b/g),
(m) => m[0],
);
}),
);
let zh = 0;
let en = 0;
for (const t of traces) {
for (const s of [t.userText, t.agentText, t.reflection ?? ""]) {
for (const ch of s) {
const code = ch.charCodeAt(0);
if (code >= 0x4e00 && code <= 0x9fff) zh++;
else if (
(code >= 0x41 && code <= 0x5a) ||
(code >= 0x61 && code <= 0x7a)
) en++;
}
}
}
const total = zh + en;
const language =
total === 0
? "unknown"
: zh / total >= 0.7
? "zh"
: en / total >= 0.7
? "en"
: "mixed";
return {
version: 1,
language,
domainTags,
toolNames,
errorCodes,
...(sourceSignature ? { sourceSignature } : {}),
};
}

function uniqueStrings(values: readonly string[]): string[] {
return Array.from(
new Set(values.map((v) => v.trim().toLowerCase()).filter(Boolean)),
).slice(0, 32);
}

// ─── helpers ────────────────────────────────────────────────────────────────

function packTraces(
Expand Down
2 changes: 2 additions & 0 deletions apps/memos-local-plugin/core/memory/l2/l2.ts
Original file line number Diff line number Diff line change
Expand Up @@ -253,6 +253,7 @@ export async function runL2(
evidenceTraces: traces,
inducedBy: `${L2_INDUCTION_PROMPT.id}.v${L2_INDUCTION_PROMPT.version}`,
now: input.now ?? Date.now(),
sourceSignature: bucket.signature,
});
const owner = ownerFromTraces(traces);
policy.ownerAgentKind = owner.ownerAgentKind;
Expand Down Expand Up @@ -632,6 +633,7 @@ function mergePolicyEvidence(existing: PolicyRow, incoming: PolicyRow, now: numb
...incoming.sourceEpisodeIds,
]),
vec: existing.vec ?? incoming.vec,
metadata: existing.metadata ?? incoming.metadata,
updatedAt: now as PolicyRow["updatedAt"],
};
}
Expand Down
38 changes: 25 additions & 13 deletions apps/memos-local-plugin/core/memory/l3/ALGORITHMS.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,13 +50,15 @@ rust|cargo|go|java|maven|gradle|typescript|javascript
```

Matches are ordered by the first-hit position so the sweep is
deterministic, then we pick the top two distinct tokens. We deliberately
**do not** embed free-form LLM tags here — domain keys must be cheap
and stable enough to hash.
deterministic, then we pick the top two distinct tokens. Newly induced L2
rows also persist trace-derived metadata (language, tags, tool names, error
codes, and the source signature); that structured metadata is preferred over
this legacy prose heuristic. We deliberately **do not** embed free-form LLM
tags here — domain keys must be cheap and stable enough to hash.

Policies with no recognised domain keyword fall into the bucket
`__generic|` and are still candidates for clustering by vector
similarity.
Policies with no recognised domain keyword fall into the generic bucket and
are only clustered when at least two vector-bearing policies pass the cosine
gate. This prevents unrelated no-label policies from becoming an L3 prompt.

---

Expand Down Expand Up @@ -126,16 +128,20 @@ strict, high-gain clusters surface first.

## 4. Evidence packing

Per cluster we assemble a prompt payload:
Per cluster we assemble one or more prompt payloads. Policies are sorted
deterministically and split into batches of at most
`maxPoliciesPerCluster`; the limit bounds prompt size and never discards
cluster members. Batch drafts are then unioned into one draft before the
single merge/create decision.

```
{
primary_tag: string,
domain_tags: string[],
avg_gain: number,
avg_support: number,
policies: PolicyPrompt[], // up to |cluster|, each capped
evidence: TracePrompt[] // at most traceEvidencePerPolicy × |cluster|
policies: PolicyPrompt[], // up to maxPoliciesPerCluster, each capped
evidence: TracePrompt[] // at most traceEvidencePerPolicy × batch size
}
```

Expand All @@ -144,8 +150,8 @@ Per cluster we assemble a prompt payload:
* For each policy we fetch the most recent non-redacted supporting
trace (by `episodeId`) and include up to `traceCharCap` characters of
`userText + reflection`. Evidence is **read-only**, never mutated.
* Total token budget is bounded by `policyCharCap × |cluster| +
traceCharCap × evidencePerPolicy × |cluster|`, which is deterministic
* Per-call token budget is bounded by `policyCharCap × batchSize +
traceCharCap × evidencePerPolicy × batchSize`, which is deterministic
and easy to debug.

---
Expand Down Expand Up @@ -250,9 +256,15 @@ if (now − kv.get(key)) < cooldownDays × 86_400_000:

## 9. Failure policy

* Storage error → propagate. Partial state remains; next run sees the
same eligible policies and re-drives.
* Storage error → warn for the affected cluster and retain retry state.
Other clusters continue; a later run re-drives the failed cluster.
* LLM error → single-cluster skip, reason logged. No cooldown update.

Failed LLM drafts use a persisted retry key scoped by cluster membership. The
retry delays are 5 minutes, 30 minutes, 2 hours, then 6 hours (capped), so a
repeated provider failure cannot consume one LLM call per episode. A successful
world-model insert/update clears the retry key and only then records the normal
cooldown. Storage failures keep the retry state and do not record success.
Other clusters continue.
* Invalid draft (missing `environment/inference/constraints`) →
treated as LLM error.
Expand Down
Loading
Loading