ERNIE Layout
Unknown
ERNIE-Layout enhances document understanding by reorganizing tokens using layout knowledge and applying spatial-aware disentangled attention in a multi-modal transformer.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
ERNIE-Layout enhances document understanding by reorganizing tokens using layout knowledge and applying spatial-aware disentangled attention in a multi-modal transformer.
Unknown
DocFormer introduces an encoder-only transformer with a CNN backbone that fuses visual, textual, and spatial features via a novel multi-modal self-attention layer for document understanding.
Unknown
LayoutLMv2 integrates text, layout, and image in a single multi-modal Transformer pre-training framework with new cross-modal tasks.
Khanam, Zeba, Achari, Vejey Pradeep Suresh, Boukhennoufa, Issam, et al.
A distributed IoT framework for urban traffic control that uses a two-stage vehicle detector and acoustic emergency vehicle detection.