Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
| Authors |
|
|---|---|
| Publication date | 2025 |
| Book title | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition : CVPR 2025 |
| Book subtitle | Nashville, Tennessee, USA, 11-15 June 2025 : proceedings |
| ISBN |
|
| ISBN (electronic) |
|
| Event | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 |
| Pages (from-to) | 3262-3271 |
| Publisher | Los Alamitos, California: IEEE Computer Society |
| Organisations |
|
| Abstract |
This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing concerns that they learn spurious correlations between videos and likely answers, reinforced by their black-box nature and remaining benchmarking biases. Our method explicitly grounds VQA tasks to video fragments in four steps: entailment tree construction, video-language entailment verification, tree reasoning, and dynamic tree expansion. A vital benefit of the method is its generalizability to current video- and image-based VLMs across reasoning types. To support fair evaluation, we devise a de-biasing procedure based on large-language models that rewrites VQA benchmark answer sets to enforce model reasoning. Systematic experiments on existing and de-biased benchmarks highlight the impact of our method components across benchmarks, VLMs, and reasoning types.
|
| Document type | Conference contribution |
| Note | With supplemental file |
| Language | English |
| Published at |
https://doi.org/10.48550/arXiv.2501.05069
(Accepted author manuscript)
https://doi.org/10.1109/CVPR52734.2025.00310
(Final published version)
|
| Published at | |
| Other links | |
| Downloads |
Liu_Commonsense_Video_Question_Answering_through_Video-Grounded_Entailment_Tree_Reasoning_CVPR_2025_paper
(Accepted author manuscript)
Commonsense_Video_Question_Answering_through_Video-Grounded_Entailment_Tree_Reasoning
(Final published version)
|
| Supplementary materials | |
| Permalink to this page | |
