Named entity recognition
**Named entity recognition (NER)** finds and labels entities in text—people, organizations, places, dates, products, and domain types you define—so systems can structure messy language. SLP3 places sequence labeling for parts of speech a...
What it is
Named entity recognition (NER) finds and labels entities in text—people, organizations, places, dates, products, and domain types you define—so systems can structure messy language. SLP3 places sequence labeling for parts of speech and named entities in its linguistic-structure volume; Yadav & Bethard (COLING 2018) survey how deep models changed NER after decades of feature-engineered systems.
<!-- IMAGE: sentence with highlighted PERSON / ORG / LOC spans -->
Visual Spec & Architecture Diagram
Annotated sentence with colored entity spans: 'Alice Chen [PERSON] joined Acme Corp [ORG] in Paris [LOC] on March 3, 2024 [DATE]'. BIO/IOB tag row underneath: B-PER I-PER O B-ORG I-ORG ... Legend for PERSON/ORG/LOC/DATE. Fake names only.
Why it matters
Search, compliance, CRM enrichment, and analytics often need entities more than chat paragraphs. Span mistakes break pipelines even when a chatbot sounds fluent.
How it works (plain)
- Split text into tokens (or characters/subwords).
- Label each token with a scheme such as BIO (Begin/Inside/Outside).
- Decode spans from tags.
- Evaluate with span-level precision/recall/F1—not only token accuracy.
- Adapt to your domain (legal, medical, tickets) with labeled examples or careful LLM extraction + review.
Neural NER (CNNs/RNNs/Transformers + CRF layers historically) largely replaced heavy hand-crafted features—but lessons from feature-based systems still help (gazetteers, cascading rules, schema design).
Everyday example
Pulling company names out of news to build a watchlist. “Jordan” might be a person, a country, or a brand—context decides.
Try it
Highlight entities in one email you wrote. Mark ambiguous cases. Write the label scheme you’d need (types + nesting rules) before picking a model.
Myths
- ⚠️ Myth: LLMs make NER evaluation unnecessary.
- ✓ Reality: Span mistakes still break pipelines—measure them.
- ⚠️ Myth: CoNLL-style person/org/loc covers every business.
- ✓ Reality: Custom types and nested entities are common.
- ⚠️ Myth: Token accuracy equals product quality.
- ✓ Reality: One bad boundary can destroy a whole span.
Sources
- Yadav & Bethard NER survey: https://aclanthology.org/C18-1182/ ↗
- SLP3 (sequence labeling chapters): https://web.stanford.edu/~jurafsky/slp3/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
- Doc-level IE survey (next step beyond sentence NER): https://aclanthology.org/2024.futured-1.6/ ↗
