Everyone who reads research has the same impression: the large language model papers arrived all at once and never stopped. We wanted the number. How many papers mention LLMs now, how fast did that happen, which fields and countries are writing them, and how does the wave compare with the last one? OpenAlex, the open catalog of scholarly works run by the nonprofit OurResearch, can answer all of that without a key, and its API is in the index as openalex. The monthly chart alone costs about half a day of its free budget; the whole post took two.
Every number below comes from a single pattern: ask /works for a filter, read meta.count, and never page through results. The definition of an LLM paper is a work whose title or abstract mentions “large language model” or “LLM”; the search is stemmed, so the plurals match too.
One request per month
GET https://api.openalex.org/works?per_page=1
&filter=from_publication_date:2026-06-01,to_publication_date:2026-06-30,
title_and_abstract.search:"large language model" OR LLM
Fifty-one of those, one per month from July 2022, give the curve.
In November 2022, the month ChatGPT was released, OpenAlex has 111 works mentioning large language models. In June 2026 it has 18,768, 169 times as many. The climb is steady rather than a single spike: June 2023 was 857, June 2024 was 3,952, June 2025 was 7,318, and the series roughly doubled again in the following year.
Two things in the chart are artifacts of the data, not the literature. The January spikes exist because a work with only a year on it is dated January 1, so every January holds its own month plus the year's undated works (about 17,000 of January 2025's 21,977). And the last three months are low because indexing lags publication; July to September 2026 average 14% below the preceding quarter today, and will fill in.
One in a hundred papers
| Year | Works mentioning LLMs | Share of papers |
|---|---|---|
| 2022 | 1,107 | |
| 2023 | 14,093 | 0.18%, one in 557 |
| 2024 | 49,348 | |
| 2025 | 100,613 | 0.97%, one in 103 |
| 2026, January to September | 160,749 |
The share counts journal articles, conference papers, preprints and reviews, which is the denominator that behaves (more on that below). One paper in a hundred published in 2025 mentions large language models. The first nine months of 2026 already hold 1.6 times all of 2025, and month by month the share has kept climbing.
In the month ChatGPT shipped, one paper in about 4,900 mentioned a large language model. In the second quarter of 2026 it was 1.63%, about one in 60. The vocabulary is moving on as well. Works mentioning “AI agents” first passed 100 a month in May 2024 and 1,000 a month in March 2026. The word “agentic” cannot be counted with the normal search, which stems it to “agent” and returns every multi-agent and reducing-agent paper ever written, but an exact-match search puts it at 788 works in 2022, 6,630 in 2025 and 26,086 in the first nine months of 2026.
For scale, the previous wave in the same field was deep learning. Line the two up from the first year each term passed 1,000 works a year and the difference is speed.
Deep learning crossed 1,000 works in 2014 and took until 2023 to pass 100,000. Large language models crossed 1,000 in 2022 and passed 100,000 in 2025, in three years instead of nine. Deep learning is still the bigger literature in absolute terms, at 181,354 works in 2025, but on the current trajectory that lasts about one more year.
Where the papers are
Add group_by=primary_topic.field.id to the same filter and the API returns the count per field in one response. Divide by the same request without the search term and you get each field's share.
Two thirds of the LLM literature is filed under computer science, where one paper in 14 mentioned the technology in 2025, up from one in 68 two years earlier. Nowhere else is it above one in 68. Medicine is at one in 260, but medicine is so large that this still means 7,311 papers, more than any field except computer science and the social sciences.
The type breakdown says something about how this literature is being published. Preprints were 54% of LLM works in 2023 and 31% in 2025, against 4.5% of all works. The 2025 split is almost even: 31,431 preprints, 31,170 conference papers, 25,294 journal articles. One source in particular carries it: arXiv holds 23,603 of the 2025 works, nearly a quarter, followed at a distance by Zenodo, SSRN, and the proceedings of NeurIPS, EMNLP and ACL.
Finally, group_by=authorships.institutions.country_code. Here we counted only journal articles and conference papers, because preprints often arrive without parsed affiliations, and a paper with authors in two countries counts in both.
In 2023 the United States had an author on 35% of published LLM papers and China on 14%. In 2025 the order is reversed: China 27%, the United States 22%. The United Kingdom and Germany slipped a point or two; Hong Kong rose. The denominator grew tenfold in between, from 5,542 published papers to 56,464, so the American total still more than quadrupled. It just grew more slowly than everyone else's.
Notes for anyone building on it
- No key and no signup. There is a budget, though: the response headers report
X-RateLimit-Limit: 1000credits per day, reset at midnight UTC, and each response states its cost inmeta.cost_usd. A plain filter orgroup_byrequest costs one credit; any request with a.searchfilter costs ten. The 267 requests behind this post cost 1,921 credits, two days' worth; the documentation says a free API key raises the budget tenfold. ReadX-RateLimit-Remainingand stop before it hits zero rather than retrying into a 429. group_by=publication_dateis rejected, which is why monthly resolution means one request per month. Agroup_byreturns at most 200 groups and hides the unknown bucket unless you append:include_unknown.- Only 64% of 2025 works carry an abstract, so a title-and-abstract search undercounts whatever is said only in the body. Treat the counts as floors.
meta.countwithper_page=1is the cheapest way to count, andmeta.x_queryshows how the server parsed your filter in plain English, which is how the stemming of “agentic” came to light.- Pick your denominator carefully. “All works” includes datasets, and bulk ingests make it jump: March 2026 alone holds 9.3 million works, 5.7 million of them datasets from two repositories. Restrict to
type:article|conference-paper|preprint|reviewfor anything that looks like a share. - Works dated only by year sit on January 1. January 1, 2025 carries 2.6 million works, 82% of the month. Treat January totals as year-plus-month.
- A field
group_byomits works with no topic assigned (about 11% of 2025), and country groups overlap because multi-country works count in each; their sum can exceed the total. title_and_abstract.searchis narrower thandefault.search, which also matches full text where OpenAlex has it. Phrase searches are stemmed.- Responses took 0.4 seconds at the median and 0.8 at the 90th percentile; the slowest was 14 seconds.
meta.db_response_time_mstells you how much of that was the query. The documentation now lives at help.openalex.org; the old docs.openalex.org links redirect. - Indexing lags publication by a few months. Stop a monthly series one quarter before today, or draw the tail differently.
The listing at /api/openalex has the base URL and the docs link. It is unclaimed. If you work at OurResearch, it's yours.
Source: OpenAlex API, /works with title_and_abstract.search:"large language model" OR LLM, counts from meta.count and group_by, pulled October 8, 2026. Shares use works of type article, conference paper, preprint or review. Country shares use articles and conference papers only.