INDEX / BLOG / ONE PAPER IN A HUNDRED NOW MENTIONS...
2026-10-09
FROM THE REGISTRY

One paper in a hundred now mentions LLMs, by the OpenAlex API's count

BY DYLAN ROY · OCTOBER 9, 2026

Everyone who reads research has the same impression: the large language model papers arrived all at once and never stopped. We wanted the number. How many papers mention LLMs now, how fast did that happen, which fields and countries are writing them, and how does the wave compare with the last one? OpenAlex, the open catalog of scholarly works run by the nonprofit OurResearch, can answer all of that without a key, and its API is in the index as openalex. The monthly chart alone costs about half a day of its free budget; the whole post took two.

Every number below comes from a single pattern: ask /works for a filter, read meta.count, and never page through results. The definition of an LLM paper is a work whose title or abstract mentions “large language model” or “LLM”; the search is stemmed, so the plurals match too.

One request per month

GET https://api.openalex.org/works?per_page=1
    &filter=from_publication_date:2026-06-01,to_publication_date:2026-06-30,
            title_and_abstract.search:"large language model" OR LLM

Fifty-one of those, one per month from July 2022, give the curve.

Bar chart of works mentioning large language models per publication month from July 2022 to September 2026. About 100 a month before ChatGPT, rising to 18,768 in June 2026. January bars are taller because works dated only by year land on January 1, and the last three months are drawn hollow because indexing is still catching up.
Works whose title or abstract mentions large language models, by publication month. January bars also hold works dated only by year; the last three months are still being indexed.

In November 2022, the month ChatGPT was released, OpenAlex has 111 works mentioning large language models. In June 2026 it has 18,768, 169 times as many. The climb is steady rather than a single spike: June 2023 was 857, June 2024 was 3,952, June 2025 was 7,318, and the series roughly doubled again in the following year.

Two things in the chart are artifacts of the data, not the literature. The January spikes exist because a work with only a year on it is dated January 1, so every January holds its own month plus the year's undated works (about 17,000 of January 2025's 21,977). And the last three months are low because indexing lags publication; July to September 2026 average 14% below the preceding quarter today, and will fill in.

One in a hundred papers

YearWorks mentioning LLMsShare of papers
20221,107
202314,0930.18%, one in 557
202449,348
2025100,6130.97%, one in 103
2026, January to September160,749

The share counts journal articles, conference papers, preprints and reviews, which is the denominator that behaves (more on that below). One paper in a hundred published in 2025 mentions large language models. The first nine months of 2026 already hold 1.6 times all of 2025, and month by month the share has kept climbing.

Line chart of works mentioning large language models as a share of all journal articles, conference papers, preprints and reviews published each month, July 2022 to September 2026. From 0.02 percent in November 2022 to 1.63 percent in the second quarter of 2026, about one paper in 61.
LLM papers as a share of the month's journal articles, conference papers, preprints and reviews. Two requests per month, one with the search filter and one without. Januaries hold works dated only by year; the last three months are still being indexed.

In the month ChatGPT shipped, one paper in about 4,900 mentioned a large language model. In the second quarter of 2026 it was 1.63%, about one in 60. The vocabulary is moving on as well. Works mentioning “AI agents” first passed 100 a month in May 2024 and 1,000 a month in March 2026. The word “agentic” cannot be counted with the normal search, which stems it to “agent” and returns every multi-agent and reducing-agent paper ever written, but an exact-match search puts it at 788 works in 2022, 6,630 in 2025 and 26,086 in the first nine months of 2026.

Bar chart of works mentioning AI agent or AI agents per publication month from January 2024 to September 2026, rising from under 100 a month to about 1,500 a month, with the first month over 100 in May 2024 and the first over 1,000 in March 2026.
Works whose title or abstract mentions “AI agent” or “AI agents”, by publication month. The January bars also hold works dated only by year.

For scale, the previous wave in the same field was deep learning. Line the two up from the first year each term passed 1,000 works a year and the difference is speed.

Line chart comparing works per year mentioning deep learning, starting from 2014, and large language models, starting from 2022, aligned on the first year each passed 1,000 works. The LLM line reaches 100,613 in its third year; the deep learning line took until its ninth year to pass that figure.
Works per year mentioning each term, aligned on the first year the term passed 1,000 works. One group_by=publication_year request per term.

Deep learning crossed 1,000 works in 2014 and took until 2023 to pass 100,000. Large language models crossed 1,000 in 2022 and passed 100,000 in 2025, in three years instead of nine. Deep learning is still the bigger literature in absolute terms, at 181,354 works in 2025, but on the current trajectory that lasts about one more year.

Where the papers are

Add group_by=primary_topic.field.id to the same filter and the API returns the count per field in one response. Divide by the same request without the search term and you get each field's share.

Dot plot of the share of each field's works mentioning large language models in 2023 and 2025. Computer science leads at 7.17 percent in 2025, then decision sciences at 1.47 percent, psychology at 0.64, social sciences 0.43, neuroscience 0.39, medicine 0.38, engineering 0.37, and the rest below 0.3 percent.
Share of each field's works mentioning large language models, 2023 and 2025. Field is OpenAlex's primary topic; the 12 fields with the most LLM papers in 2025.

Two thirds of the LLM literature is filed under computer science, where one paper in 14 mentioned the technology in 2025, up from one in 68 two years earlier. Nowhere else is it above one in 68. Medicine is at one in 260, but medicine is so large that this still means 7,311 papers, more than any field except computer science and the social sciences.

The type breakdown says something about how this literature is being published. Preprints were 54% of LLM works in 2023 and 31% in 2025, against 4.5% of all works. The 2025 split is almost even: 31,431 preprints, 31,170 conference papers, 25,294 journal articles. One source in particular carries it: arXiv holds 23,603 of the 2025 works, nearly a quarter, followed at a distance by Zenodo, SSRN, and the proceedings of NeurIPS, EMNLP and ACL.

Finally, group_by=authorships.institutions.country_code. Here we counted only journal articles and conference papers, because preprints often arrive without parsed affiliations, and a paper with authors in two countries counts in both.

Dot plot of the share of published LLM papers with at least one author in each of the top ten countries, 2023 and 2025. China rose from 14 to 27 percent and the United States fell from 35 to 22 percent; the United Kingdom, Germany and India fell slightly; Hong Kong rose.
Share of journal articles and conference papers mentioning large language models with at least one author affiliated in each country, 2023 and 2025. Multi-country papers count in each country.

In 2023 the United States had an author on 35% of published LLM papers and China on 14%. In 2025 the order is reversed: China 27%, the United States 22%. The United Kingdom and Germany slipped a point or two; Hong Kong rose. The denominator grew tenfold in between, from 5,542 published papers to 56,464, so the American total still more than quadrupled. It just grew more slowly than everyone else's.

Notes for anyone building on it

The listing at /api/openalex has the base URL and the docs link. It is unclaimed. If you work at OurResearch, it's yours.

Source: OpenAlex API, /works with title_and_abstract.search:"large language model" OR LLM, counts from meta.count and group_by, pulled October 8, 2026. Shares use works of type article, conference paper, preprint or review. Country shares use articles and conference papers only.