Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #539

CapMem: benchmark for caption-based episodic memory in egocentric video (arXiv:2609.17688v1)

This paper introduces CapMem, a human-annotated benchmark and task (Episodic Memory Video Caption QA) with 75 egocentric videos (33.7 hours) and 1,000 multiple-choice questions across 16 scenarios, to test whether textual captions can serve as reusable episodic memory. Experiments show caption-windowed CaptionQA (30s/60s) often outperforms direct VideoQA on long videos, with matched-frame controls across six Qwen models yielding mean accuracy gains (≈3.22 and 2.55 points) and a caption-guided retrieve-and-verify step adding up to 5.3 points of improvement.

KEY POINTS

  1. This paper introduces CapMem, a human-annotated benchmark and task (Episodic Memory Video Caption QA) with 75 egocentric videos (33.7 hours) and 1,000 multiple-choice questions across 16 scenarios, to test whether textual captions can serve as reusable episodic memory.
  2. Experiments show caption-windowed CaptionQA (30s/60s) often outperforms direct VideoQA on long videos, with matched-frame controls across six Qwen models yielding mean accuracy gains (≈3.22 and 2.55 points) and a caption-guided retrieve-and-verify step adding up to 5.3 points of improvement.
  3. The result suggests lightweight caption memories can improve episodic reasoning over long egocentric video and reduce costly visual-token and long-context retrieval burdens for wearable assistants and VLM deployments.

WHY IT MATTERS

The result suggests lightweight caption memories can improve episodic reasoning over long egocentric video and reduce costly visual-token and long-context retrieval burdens for wearable assistants and VLM deployments.

SOURCES & TIMELINE

1