AtlasNLP tracks who creates NLP datasets and whose language they cover — mapping 13,000+ dataset records across countries to reveal the geographic gaps that language labels alone cannot show.
Dataset coverage by country
Coverage follows a pronounced long-tail: the United States leads with 631 explicitly-attributed dataset records, while 39 countries have none and 121 of 197 have ten or fewer.
Hover a country to see its dataset count · View full interactive map →
What AtlasNLP reveals
Coverage is highly uneven: 39 countries have no explicitly attributed dataset records at all, and 121 of 197 countries have ten or fewer. Even under the broader explicit+inferred layer, 75.0% of country-task pairs remain empty.
The United States, China, India, the United Kingdom, and Germany account for over a third of country-record associations. Production is asymmetric too: among well-represented countries, 62.9% have most of their datasets produced by outside institutions.
Even combining explicit and inferred evidence, 9,956 of 13,462 Core dataset records could not be attributed to a country — not because they lack geographic context, but because it isn't documented in a recoverable form.
Explore AtlasNLP data
Automated large-scale collection
13,462 primary dataset-contribution records extracted from ACL Anthology papers (1952–2025) and vetted through a post-extraction audit, with task, language, content-country, and producer-country metadata.
Human-validated reference set
1,480 entries (989 normalized dataset-name groups) curated and cross-validated by contributors from diverse geographic backgrounds, prioritising underrepresented regions.