Troubleshooting¶
Symptom-first. Each entry: what you see, why, what to do. The Debug Tools are the companion page.
Startup¶
Server exits immediately.
The four required secrets are validated at boot: LLM_PROVIDER/LLM_API_KEY (or the Azure pair), MEMORY_OS_MONGODB_PASSWORD, ARANGODB_PASSWORD. The log names the missing one.
connection refused to ArangoDB during init-ltm.
ArangoDB boots slowly. Wait ten seconds, re-run; the command is idempotent.
Server starts but /health says degraded.
Redis or MongoDB is unreachable. The /api/v1/health endpoint reports per-dependency status with error strings.
Nothing is being remembered¶
Interactions store fine but no chains ever form. In order of likelihood:
- Workers disabled: check
ENABLE_MEMORY_WORKERS=true. - Not enough user messages: chain-break checks need two user messages to compare; a single exchange forms nothing until more conversation arrives or the 20-event force flush hits.
- Same topic throughout: no break, no chain (yet). The force flush (
STM_MAX_EVENTS_BEFORE_FLUSH, default 20) eventually partitions long single-topic sessions. - Quality gate rejection: small talk is filtered by design. Check with
verifydb; lowerMTM_QUALITY_MIN_SCOREonly if real content is being dropped. - LLM failing: run
test-gemini; watchllm_fallback_calls_total.
Chains form but nothing reaches the knowledge graph.
- Promotion threshold too high:
.env.exampleshipsMEMORY_OS_MTM_HEAT_THRESHOLD=0.8(conservative); the pipeline default is 0.3. Verify which your deployment uses. - Chains cooled before the promoter ran: check heat values in
verifydboutput against the decay math. - Extraction schema failures: check
memos_extractor_schema_failures_totaland server logs. - LTM disabled or ArangoDB unreachable: confirm
init-ltmran andARANGODB_*vars are set.
Configuration gotchas¶
A setting seems ignored. Several knobs exist under two names read by different code paths. Set both:
| Pair |
|---|
STM_CACHE_MAX_TURNS and MEMORY_OS_STM_CACHE_MAX_TURNS |
MEMORY_WORKER_COUNT and MEMORY_OS_NUM_WORKERS |
REDIS_HOST / MONGO_URI and their MEMORY_OS_* forms |
The Configuration Reference flags every alias pair.
Heat weights don't add up.
The scorer reads HEAT_WEIGHT_INTRINSIC (default 0.7) and HEAT_WEIGHT_DENSITY (default 0.3). The HEAT_WEIGHT_COGNITIVE variable that appears in .env.example is not read by the code; setting it does nothing. Use HEAT_WEIGHT_DENSITY, keep the pair summing to 1.0, and watch memos_heat_score_distribution after changing.
Request errors¶
| HTTP | Meaning | Fix |
|---|---|---|
| 400 | Missing user_id, empty query, invalid event type |
Fix the request |
| 401 | Auth enabled, credentials missing/wrong | See Enabling Authentication |
| 404 | Unknown session_id |
Sessions are cheap; create a new one (memory is keyed by identity, not session) |
| 412 | Blob payload without a blob store | Configure BLOB_* (guide) |
| 500 | Datastore failure | Server logs name the store |
Search returns nothing (or garbage)¶
- Dimension mismatch after a provider switch: the classic. Milvus collection was created for 1536-dim (Azure) and now receives 768-dim (Gemini) or vice versa. Recreate the collection (procedure).
- Recent conversation not searchable: expected; it may still be in the STM window, unformed. Use
GetContextfor the window. - Old memory gone: check for
status: "archived"inverifydb; archived chains lose their vectors (lifecycle).
Load and cost problems¶
LLM bill or rate-limit errors spiking.
Remember the multiplication: MEMORY_WORKER_COUNT × replicas versus LLM_RATE_LIMIT_PER_MINUTE. Retries are LIFO with no backoff, so a provider outage churns; the circuit breaker (default threshold 5, 60s open) is your protection. Also check whether a chatty automation is skipping coalescing by omitting execution_id/step_id.
ArangoDB memory spikes at night. That's the analytics window. Pregel loads the whole graph; size the instance for it and never trigger analytics in-process (the rule).
Inspecting a specific failed task¶
Task results persist for one hour:
redis-cli --scan --pattern 'task_results:v1:*' | head
redis-cli GET 'task_results:v1:<tenant>:<user>:<agent>:<taskId>'
The value carries success, an error message, and the processing timestamp.