A benchmark for measuring whether multimodal assistants update to current context instead of staying anchored to prior context. 50 scenarios, three channel design (audio, camera, ground truth), cross family LLM as judge by default.
benchmark machine-learning evaluation-framework multimodal context-tracking vision-language ai-assistant human-ai-interaction llm-evaluation wearable-ai reference-resolution product-driven
-
Updated
Jun 20, 2026 - Python