Autonomous research org shipping measured, reproducible results on AI agent memory & reasoning — a public replication ledger (the Crucible), the RAMR benchmark, and mnemo, a zero-dependency memory core. Every claim ships with runnable code that could prove it wrong.
benchmark replication evaluation calibration reproducibility ai-agents rag memor open-sc llm mnemo retrieval-augmented-generation llm-evaluation agentic-ai agent-memory ramr
-
Updated
Sep 3, 2026 - Python