RESEARCH RESOURCE

A Country-Aware Atlas of NLP Datasets

AtlasNLP tracks who creates NLP datasets and whose language they cover — mapping 13,000+ dataset records across countries to reveal the geographic gaps that language labels alone cannot show.

13,462 Core dataset records
158 Countries explicitly represented
30 Core task categories
1,241 Audited language labels
Global coverage

Dataset coverage by country

Coverage follows a pronounced long-tail: the United States leads with 631 explicitly-attributed dataset records, while 39 countries have none and 121 of 197 have ten or fewer.

0 1–5 6–25 26–100 101–500 500+
Loading map…

Hover a country to see its dataset count · View full interactive map →

Key findings

What AtlasNLP reveals

79.2%
of country-task pairs have no explicit country-attributed datasets

Coverage is highly uneven: 39 countries have no explicitly attributed dataset records at all, and 121 of 197 countries have ten or fewer. Even under the broader explicit+inferred layer, 75.0% of country-task pairs remain empty.

36.4%
of country associations come from just 5 countries

The United States, China, India, the United Kingdom, and Germany account for over a third of country-record associations. Production is asymmetric too: among well-represented countries, 62.9% have most of their datasets produced by outside institutions.

74.0%
of Core records have no recoverable represented country

Even combining explicit and inferred evidence, 9,956 of 13,462 Core dataset records could not be attributed to a country — not because they lack geographic context, but because it isn't documented in a recoverable form.

Read the full analysis
Datasets

Explore AtlasNLP data

View full dataset explorer
AtlasNLP Core

Automated large-scale collection

13,462 primary dataset-contribution records extracted from ACL Anthology papers (1952–2025) and vetted through a post-extraction audit, with task, language, content-country, and producer-country metadata.

13,462 entries Automated + audited ACL-derived
Browse
AtlasNLP Gold

Human-validated reference set

1,480 entries (989 normalized dataset-name groups) curated and cross-validated by contributors from diverse geographic backgrounds, prioritising underrepresented regions.

1,480 entries (989 groups) Human-curated Cross-validated
Browse