NER Datasets for Named Entity Recognition

Domain-specific named entity corpora with fine-grained entity types, nested entity support, and inter-annotator agreement verification. Built for production NLP and information extraction systems.

Abstract data visualization representing named entity recognition
20+
Years in AI training data
1,000+
Language locales
1M+
Hours of video annotated
ISO
27001 certified

Beyond General NER Benchmarks

CoNLL-2003 and OntoNotes established strong NER baselines. But production information extraction requires domain-specific entity types, specialized vocabulary, and text genres those benchmarks never covered.

Clinical NLP must recognize drug names, dosages, and medical procedures. Legal NLP must identify parties, statutes, and citations. Financial NLP must extract companies, instruments, and regulatory bodies. General NER models fail on these specialized entity types.

Off-the-shelf NER datasets suffer from entity type gaps for domain-specific classes and genre mismatch: models trained on news wire text perform poorly on clinical notes, contracts, or social media.

LXT builds custom NER datasets with entity taxonomies designed for your information extraction pipeline. We annotate your target text genres, entity types, and nesting structures, with inter-annotator agreement metrics and adjudication processes that ensure production-grade ground truth.

Limitations of Public NER Datasets

Standard benchmarks serve research well. Production deployments need more.

DatasetPrimary LimitationImpact
CoNLL-2003Only 4 entity types (PER, ORG, LOC, MISC); news wire genre only; no domain-specific entities for medical, legal, or financial applicationsType coverage
OntoNotes 5.018 entity types but broad and shallow; mixed genres with inconsistent annotation quality; limited named entity nestingShallow types
ACE 20057 entity types across news and broadcast; complex annotation schema creates inter-annotator disagreement; aging corpusAging data
WNUT-17Social media text only; noisy and informal register; emerging entity focus limits applicability to formal domain NLPGenre narrow
TACREDRelation-focused rather than entity-focused; limited entity type coverage; news domain onlyRelation-only

Not sure which specs you need?

Our data specialists help you scope the right dataset for your model architecture.

Talk to a Specialist

Specs Built Around Your Model

Public datasets come fixed. Yours is configured for your architecture, environment, and use case.

Entity Taxonomy

Type Coverage

  • Core Types: Custom entity types designed for your information extraction use case
  • Nested Entities: Overlapping and nested entity spans for complex extraction tasks
  • Normalization: Entity linking to ontologies (UMLS, Wikidata, custom KB) where required

Text Genre

Corpus Composition

  • Source Types: Clinical notes, contracts, financial filings, social media, news, or mixed
  • Volume: 5K to 500K+ annotated tokens matched to your pipeline requirements
  • Language: Monolingual and cross-lingual corpora across 50+ languages

Annotation Quality

Ground Truth Standards

  • IAA Protocol: Inter-annotator agreement (Cohen's Kappa, F1) measured at each stage
  • Adjudication: Disagreements resolved by senior annotator review
  • Guidelines: Detailed annotation guidelines document edge cases and boundary decisions

Need a custom configuration?

We've built datasets across dozens of domains and use cases. Let's scope yours.

Get a Custom Quote

Edge Cases in NER Annotation

High-accuracy models handle rare attributes that public datasets miss.

Nested and Overlapping Entities

Clinical text contains entities like medication names nested within treatment descriptions. Nested annotation schemas capture these hierarchical relationships for complex extraction models.

Ambiguous Entity Boundaries

Compound proper nouns and multi-word entities have ambiguous spans. Consistent boundary guidelines and annotator calibration prevent systematic boundary errors in training data.

Domain-Specific Abbreviations

Medical abbreviations, legal shorthand, and financial codes require specialized annotator knowledge. Domain expert annotators prevent systematic mislabeling of abbreviations.

Code-Switched and Multilingual Text

Customer-facing text mixes languages within sentences. Cross-lingual annotation protocols and bilingual annotators handle code-switching in social media and support ticket corpora.

Human-in-the-Loop Annotation

Precise annotation bridges raw data and learnable signal. Expert annotators deliver precision automated tools can't match.

🏷️

Span Annotation

Expert annotators mark entity spans with precise character-level boundaries. BIO/BIOES tagging schemes output in CoNLL, spaCy, and Hugging Face formats.

🔗

Entity Linking and Normalization

Entities linked to target knowledge bases (UMLS, Wikidata, custom ontologies) for downstream disambiguation and knowledge graph population tasks.

📋

Relation Annotation

Co-occurrence and semantic relationship annotation between entity pairs for joint NER and relation extraction model training.

NER Datasets for Your Domain

Custom taxonomies and collection protocols for specific deployment contexts.

🏥

Clinical NLP

Drug extraction, diagnosis coding, clinical trial matching

⚖️

Legal NLP

Party identification, statute citation, contract analysis

💰

Financial NLP

Company mention, instrument extraction, regulatory filing analysis

📰

Media Monitoring

Brand mention, event detection, person tracking

🤖

Chatbots and Dialogue

Entity extraction for intent and slot filling

🔍

Search and Retrieval

Named entity indexing for enterprise search

📊

Knowledge Graph

Entity extraction for KG construction and population

🌍

Multilingual IE

Cross-lingual entity extraction across global markets

Secure and Ethical Data Collection

Data collection involving people and sensitive content requires robust security, compliance, and ethical protocols at every stage.

🌐

Global Demographic Reach

Collection across 1,000+ locales and diverse demographics to prevent algorithmic bias in your deployed models.

🔒

ISO 27001 Certified

Sensitive projects processed in certified secure facilities meeting the highest information security standards.

✅

GDPR & Privacy Compliance

All collection and annotation protocols vetted for consent and privacy. Legally robust for global deployment.

NER Dataset FAQs

Can you build a custom entity taxonomy for our domain?+
Yes. We work with your NLP team to define the entity type hierarchy, boundary rules, and normalization targets. We produce detailed annotation guidelines before collection begins.
What inter-annotator agreement do you target?+
We target Cohen's Kappa above 0.80 for standard entity types and above 0.70 for ambiguous boundary cases. IAA reports are included in delivery with per-type breakdowns.
Do you support nested entity annotation?+
Yes. We support nested and overlapping entity spans using standoff annotation formats. This is essential for clinical and legal text where entities frequently contain other entities.
What output formats do you support?+
CoNLL-2003 IOB2, BIOES, spaCy JSONL, Hugging Face datasets format, BRAT standoff, and custom formats. Entity linking outputs in JSON with knowledge base IDs.
Can you annotate multilingual or code-switched text?+
Yes. We have annotators for 50+ languages and bilingual annotators for code-switched text. Cross-lingual annotation uses shared guidelines with language-specific edge case documentation.
What does a custom NER dataset cost?+
Projects range from $8K for focused single-domain corpora (10K-50K tokens) to $80K+ for large multi-domain, multi-language datasets exceeding 500K annotated tokens.
How do you handle sensitive entities like personal names and medical identifiers?+
Sensitive entities are annotated but de-identified or pseudonymized in delivery unless your use case requires real entity values with appropriate data agreements.

Scope Your Custom NER Dataset

Share your target entity types, text genres, and volume requirements. An NLP data specialist will provide a detailed annotation plan within 48 hours.

Contact us.

Please provide us with the details of your inquiry and one of our team members will be in touch.

Join our global team of contributors today

Apply here to be considered for future projects including data collection, annotation and transcription
Start application
(opens in a new tab)