Document-level information extraction
**Document-level information extraction (doc-IE)** pulls structured facts—entities, relations, events—from **whole documents**, not only single sentences. Zheng, Wang, and Huang (FuturED 2024) survey contemporary doc-IE, analyze errors o...
What it is
Document-level information extraction (doc-IE) pulls structured facts—entities, relations, events—from whole documents, not only single sentences. Zheng, Wang, and Huang (FuturED 2024) survey contemporary doc-IE, analyze errors of strong systems, and highlight lingering blockers. SLP3’s information extraction chapter (relations, events, time) plus coreference chapters are the conceptual base.
<!-- IMAGE: multi-paragraph doc → entity graph with cross-sentence links -->
Visual Spec & Architecture Diagram
Document-level IE: multi-paragraph fake medical/business note on left; extracted entity nodes (Person, Org, Drug) and relation edges (works_at, prescribed) forming a graph on right. Cross-sentence dashed arrow labeled 'coreference'.
Why it matters
Contracts, scientific papers, incident reports, and news stories spread facts across paragraphs. Sentence-level NER/RE misses “who did what to whom” when arguments are pages apart. Chatbots that “remember the vibe” are not a substitute for structured IE with audit trails.
How it works (plain)
- Detect entities (and keep coreference clusters: “she” = “Dr. Ng”).
- Extract relations/events that may cross sentence boundaries.
- Use discourse context / document encoders / graph methods as systems evolve.
- Evaluate with precision/recall on document-annotated sets.
- Analyze errors: coreference failures and weak reasoning show up repeatedly in the 2024 survey’s findings.
Everyday example
In a 12-page vendor contract, the liability cap appears in section 9 while the party nickname is introduced in section 2. Doc-IE must link them; sentence NER alone will not.
Try it
Take a one-page news story. Write three triples (subject, relation, object) that require crossing sentence boundaries. Note which hinge on pronouns.
Myths
- ⚠️ Myth: Long-context LLMs solved doc-IE.
- ✓ Reality: They help draft candidates; surveys still find coreference and reasoning failures—and you still need schemas + eval.
- ⚠️ Myth: Sentence IE + concatenation is enough.
- ✓ Reality: Cross-sentence arguments and relation transitivity create distinct failure modes.
- ⚠️ Myth: Perfect entity tags imply perfect relations.
- ✓ Reality: Relation errors persist after entity detection.
Sources
- Doc-IE survey: https://aclanthology.org/2024.futured-1.6/ ↗
- NER survey (foundations): https://aclanthology.org/C18-1182/ ↗
- SLP3 IE / coreference chapters: https://web.stanford.edu/~jurafsky/slp3/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
