RESEARCH · RESEARCH · #539
CapMem: benchmark for caption-based episodic memory in egocentric video (arXiv:2609.17688v1)
This paper introduces CapMem, a human-annotated benchmark and task (Episodic Memory Video Caption QA) with 75 egocentric videos (33.7 hours) and 1,000 multiple-choice questions across 16 scenarios, to test whether textual captions can serve as reusable episodic memory. Experiments show caption-windowed CaptionQA (30s/60s) often outperforms direct VideoQA on long videos, with matched-frame controls across six Qwen models yielding mean accuracy gains (≈3.22 and 2.55 points) and a caption-guided retrieve-and-verify step adding up to 5.3 points of improvement.
KEY POINTS
- This paper introduces CapMem, a human-annotated benchmark and task (Episodic Memory Video Caption QA) with 75 egocentric videos (33.7 hours) and 1,000 multiple-choice questions across 16 scenarios, to test whether textual captions can serve as reusable episodic memory.
- Experiments show caption-windowed CaptionQA (30s/60s) often outperforms direct VideoQA on long videos, with matched-frame controls across six Qwen models yielding mean accuracy gains (≈3.22 and 2.55 points) and a caption-guided retrieve-and-verify step adding up to 5.3 points of improvement.
- The result suggests lightweight caption memories can improve episodic reasoning over long egocentric video and reduce costly visual-token and long-context retrieval burdens for wearable assistants and VLM deployments.
WHY IT MATTERS
The result suggests lightweight caption memories can improve episodic reasoning over long egocentric video and reduce costly visual-token and long-context retrieval burdens for wearable assistants and VLM deployments.