TAT-DQA
Freea large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly in financial reports.
About TAT-DQA
TAT-DQA is a large-scale Document Visual Question Answering (VQA) dataset constructed by extending TAT-QA. It aims to stimulate progress in QA research over complex and realistic visually-rich documents that contain both tabular and textual content, especially those requiring numerical reasoning. The dataset is sampled from real-world high-quality financial reports, with an average of 550 words per document. It features diverse answer forms (single span, multiple spans, free-form) and various numerical reasoning capabilities (addition, subtraction, multiplication, division, counting, comparison, sorting, and their compositions). TAT-DQA contains 16,558 questions associated with 2,758 documents (3,067 document pages).
Key Features
Pros & Cons
- Large-scale annotated dataset with rich financial document content
- Covers multiple numerical reasoning operations and answer types
- Provides derivation steps for answers, enabling explainability evaluation
- Open-source and freely available for download
- Limited to the financial domain, reducing generalizability to other document types
- Download requires Google Drive, no API or direct model access
- Document language is English only (implied by examples)