The framework is designed to distinguish between cognitive failure (reasoning degradation under task complexity) and system failure (artifacts arising from infrastructure limits such as token truncation or parsing errors).
-
Updated
Aug 5, 2026 - Python