The arXiv Two-a-Month Cap Turns Paper Flood Control Into a Bet on the Heavy Tail
arXiv now caps each submitter at two papers a month, and a harvest of September’s 38,908 announced papers shows a limit that reaches roughly 2 to 5 percent of them, leaving most of the volume with...
On October 1, arXiv began limiting every submitter to two submissions per calendar month, with no more than three active at any time. The preprint server says it needs the limit because September brought 40,363 submissions, up from 20,569 in September 2024 and 9,869 in September 2016, and because the month generated almost 9,000 support tickets for its staff and moderators. arXiv’s stated cause is AI tools, which it says have put the old practical limit on one person’s output up for debate.
Table Of Content
- What arXiv changed on October 1
- The cap is the third gate
- The volume curve behind it
- What the 40,363 counts
- The cs.AI claim
- How many papers sit above the line?
- Step 1: harvest one month
- Step 2: count who filed what
- Reading the output
- Where the cap bites
- Reading the cap as a rate limiter
- The key is the submitter
- The window is a calendar month
- The unit is a submission, and replacements are unanswered
- What this analysis cannot show
- What to do with it
- If you file papers
- If you design a quota
- What to watch next
- Related reading on this site
- Method notes
A per-account quota in front of a human review queue is a familiar design, so the interesting question is what the quota can reach. arXiv publishes its monthly counts and lets anyone harvest its metadata, which means the question can be tested. I harvested every announced paper whose first version is dated September 2026 and counted who filed them. The short version:
- arXiv’s headline figure equals the highest identifier sequence number issued for the month (2609.40363), which suggests the series counts papers that reached announcement, so rejected submissions are probably not in it.
- The announcement’s “over 6X” rise for cs.AI holds for papers whose primary category is cs.AI (392 to 2,361) and does not hold for papers merely listed there (2,650 to 5,923, a 2.2-fold rise).
- At most 2,075 of September’s 38,908 announced papers (5.3 percent) were a submitter name’s third or later paper of the month, and as few as 822 (2.1 percent) if a shared name is split by primary category.
- 96.1 percent of the 32,194 submitter names filed one or two papers, so the limit sits above what almost everyone did.
- The moderation load that arXiv says sits with a small share of submitters includes rejected and held submissions, which public data does not show. The cap rests on a bet that cannot be checked from outside.
What arXiv changed on October 1
The rule fits in one sentence. In arXiv’s words, it “now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.” The announcement calls this “a stopgap while we determine what may be the new best practice for authors employing more and more advanced AI tools.” It also says why a number replaced judgment: before AI tools, there was “an easily discernible, practical limit” to how fast independent submissions could be produced, and volunteer moderators would limit submitters who were well above it. That practical limit is now, in arXiv’s phrase, “up for debate.”
The announcement’s FAQ settles most of the edge cases:
| Question | What arXiv says |
|---|---|
| Does a rejected paper count? | Yes. The limit is for submissions, not announced papers, because submissions are what consume moderator time. |
| Does it matter which category? | No. The limit applies across all categories. |
| What about papers still on hold next month? | They do not count against the new month’s limit, but they do count toward the three active submissions. |
| Do co-authors share the count? | No. The policies apply only to submitters, so a paper counts toward the submitter’s rate and not the co-authors’. arXiv adds that it is “important for co-authors to coordinate their submissions accordingly.” |
| What if I delete a paper before it is announced? | It counts toward neither limit. |
| What about a backlog or a conference deadline? | The three-active limit has been in place since 2024 and stays. Authors are asked to plan. |
The Register adds one definition from an arXiv spokesperson: an active paper is one that has been submitted and is under consideration and review but is not yet published on the site.
The justification comes from Thomas Dietterich, whom the announcement identifies as Distinguished Professor Emeritus at Oregon State University and Chair of the arXiv Editorial Advisory Council: “a relatively small proportion of authors are submitting a large number of low-quality papers and consuming a disproportionate fraction of the moderators’ time.” The cap is built on that sentence. It is a heavy-tail claim: a small share of submitters is said to account for a large share of the load. The rest of this piece tests how far a per-submitter limit can follow it.
The cap is the third gate
The limit arrives after three earlier moves aimed at the same flood. On October 31, 2025, arXiv’s computer science category began requiring review articles and position papers to be accepted at a journal or conference after peer review, citing an “unmanageable influx” of both. On January 21, 2026, arXiv stopped accepting an institutional email address “as the sole qualifier of endorsement for new authors.” New submitters now need either an institutional email plus a previously accepted paper in the same endorsement domain, or a personal endorsement from an established author. In May, Nature reported that arXiv would ban researchers for one year when a submission contains hallucinated references or other incontrovertible signs of unchecked generative AI output.
In infrastructure terms those are three different controls: an identity cost (who may submit), validation rules (what may be submitted), and a quota (how much one identity may submit). The quota is the cheapest of the three to enforce, because it needs no judgment about any paper, and the least informed, because it cannot see quality at all.
The volume curve behind it
arXiv’s monthly submissions page publishes the series as a CSV. Here are the Septembers:
| September | Count | Change from the prior September |
|---|---|---|
| 2016 | 9,869 | n/a |
| 2022 | 15,640 | +2.1 percent |
| 2023 | 17,453 | +11.6 percent |
| 2024 | 20,569 | +17.9 percent |
| 2025 | 26,646 | +29.5 percent |
| 2026 | 40,363 | +51.5 percent |
The yearly increase has grown every year since 2022. arXiv’s arithmetic holds: 40,363 divided by 9,869 is 4.09, the 309 percent rise The Register reports, and 40,363 divided by 20,569 is 1.96, which rounds to the “doubled” in the announcement. September 2026 also beat the previous record month, June 2026 at 32,040, by 26.0 percent. Some of that is seasonal. September was 1.295 times August in 2026, against 1.221 in 2025, 1.179 in 2024, and between 1.03 and 1.17 from 2016 to 2023. The seasonal bump has been growing along with the baseline.
What the 40,363 counts
The statistics page says the chart shows “the number of new submissions received during each month.” I checked the figure against arXiv’s identifiers. An arXiv identifier has the form YYMM.NNNNN, “with 5-digits for the sequence number within the month,” according to arXiv’s identifier documentation. In my harvest the highest sequence number among identifiers beginning 2609 is 40,363, and a wider harvest that also covers August shows 31,173 for 2608. Those are exactly the official September and August counts.
The identifier month follows announcement, not submission. In a wider harvest that also kept August first versions, 1,625 papers first submitted in August carry a 2609 identifier, and 819 papers first submitted in September carry a 2610 identifier. That wider harvest found 39,714 announced papers with a 2609 identifier, 98.4 percent of the 40,363 sequence numbers. The other 649 belong to papers outside the harvest (first version before August) or to papers removed after announcement, and public data cannot say which.
Our reading: the monthly series counts identifiers issued, which tracks papers that reached announcement. Submissions that moderators rejected before announcement do not appear to be in it, so the real intake is larger than 40,363 by an amount nobody outside arXiv can see. arXiv’s word “received” is broader than what the number seems to measure. I did not find the counting rule documented, and I did not ask arXiv.
The cs.AI claim
The announcement says: “Increases over the past two years in cs.AI (over 6X increase) illustrate the rate at which researchers are submitting to arXiv. Other arXiv categories have seen submissions double over the same period of time.” Whether that is true depends on what “in cs.AI” means, because arXiv lists a paper under every category it carries, not only its primary one.
| Basis | Sept 2024 | Sept 2025 | Sept 2026 | 2024 to 2026 |
|---|---|---|---|---|
| cs.AI as primary category | 392 | 896 | 2,361 | 6.0 times |
| cs.AI listed (primary or cross-list) | 2,650 | 4,262 | 5,923 | 2.2 times |
| cs.LG as primary category | 1,382 | 2,333 | 3,233 | 2.3 times |
| cs.CL as primary category | 1,059 | 1,587 | 1,864 | 1.8 times |
The “over 6X” figure is real on a primary-category basis and absent on a listing basis. I read the 2024 and 2025 counts from the public API and the 2026 counts from the harvest below. The API’s own count for cs.AI as primary in September 2026 is also 2,361, so the two routes agree. cs.AI as a primary category went from 1.9 percent of announced papers in September 2024 to 6.1 percent in September 2026. Its 1,969 added papers are 10.7 percent of the 18,460 papers added over the same two years (from 20,448 to 38,908 announced papers).
The second half of the sentence, that other categories doubled, is a fair average with a wide spread. On a listing basis the API counts below roughly doubled across all categories (1.90 times) and grew faster than that in mathematics, while astronomy, quantitative biology and electrical engineering grew far less. Listing counts overlap, because a cross-listed paper counts in every category it carries, so the rows do not sum.
| Listed in | Sept 2024 | Sept 2026 | Change |
|---|---|---|---|
| All categories | 20,448 | 38,908 | 1.90 times |
| cs.* | 9,825 | 19,399 | 1.97 times |
| math.* | 3,833 | 9,112 | 2.38 times |
| cs.SE | 293 | 673 | 2.30 times |
| cs.CR | 574 | 1,218 | 2.12 times |
| quant-ph | 1,118 | 2,284 | 2.04 times |
| astro-ph.* | 1,696 | 2,357 | 1.39 times |
| eess.* | 1,953 | 2,104 | 1.08 times |
| q-bio.* | 401 | 435 | 1.08 times |
The surge is not a one-category story. Primary cs.AI is the steepest single line, but it supplies about a tenth of the added papers.
How many papers sit above the line?
The public arXiv API and the OAI-PMH interface both serve metadata. The OAI interface is the one that carries the submitter’s name, in its arXivRaw format, which the documentation describes as very close to the internal format arXiv stores, including the history of previous versions. Two properties of that interface shape the method. The documentation says it “does not support selective harvesting based on submission date,” and that every record’s datestamp is its last modification time. A record cannot be modified before its first version exists, so a harvest that starts on the first of the month contains every paper first submitted that month. The script keeps only records whose first version is dated in the month.
Step 1: harvest one month
arXiv’s terms of use ask for “no more than one request every three seconds” on one connection at a time, so the script waits three seconds between pages. It keeps three fields per paper: the identifier, the lower-cased submitter name, and the primary category (the first one listed).
# harvest_month.py
# First versions of arXiv papers submitted in one month, with the submitter field.
# usage: python harvest_month.py 2026-09 2026-10-03 sept.jsonl
import json
import sys
import time
import urllib.error
import urllib.request
import xml.etree.ElementTree as ET
from email.utils import parsedate_to_datetime
BASE = "https://oaipmh.arxiv.org/oai"
O = "{http://www.openarchives.org/OAI/2.0/}"
R = "{http://arxiv.org/OAI/arXivRaw/}"
month, until, outfile = sys.argv[1:4]
url = f"{BASE}?verb=ListRecords&metadataPrefix=arXivRaw&from={month}-01&until={until}"
kept = 0
with open(outfile, "w", encoding="utf-8") as out:
while url:
req = urllib.request.Request(url, headers={"User-Agent": "arxiv-rate-limit-demo/1.0"})
try:
xml = urllib.request.urlopen(req, timeout=180).read()
except urllib.error.HTTPError as err:
if err.code not in (429, 503):
raise
time.sleep(int(err.headers.get("Retry-After", 10)) + 1)
continue
root = ET.fromstring(xml)
for rec in root.iter(O + "record"):
raw = rec.find(f"{O}metadata/{R}arXivRaw")
if raw is None:
continue
v1 = parsedate_to_datetime(raw.find(f"{R}version/{R}date").text)
if v1.strftime("%Y-%m") != month:
continue
submitter = " ".join((raw.findtext(R + "submitter") or "").casefold().split())
category = raw.findtext(R + "categories").split()[0]
out.write(json.dumps({"id": raw.findtext(R + "id"), "sub": submitter, "cat": category}) + "\n")
kept += 1
token = root.find(f"{O}ListRecords/{O}resumptionToken")
url = f"{BASE}?verb=ListRecords&resumptionToken={token.text}" if token is not None and token.text else None
time.sleep(3) # arXiv asks for at most one request every three seconds, on one connection
print(f"{kept:,} first versions dated {month}")
python harvest_month.py 2026-09 2026-10-03 sept.jsonl
38,908 first versions dated 2026-09
The API, which dates a paper by its first submission, agrees. A query with submittedDate:[202609010000 TO 202609302359] returns 38,908 results. These are announced papers: 1,455 fewer than the 40,363 in the statistics table, which counts by identifier month.
Step 2: count who filed what
The counting script groups papers by submitter name and then counts the papers beyond each name’s second. It repeats the count with one change, treating the same name in two different primary categories as two people, which gives a stricter identity. It also prints the two checks used above: the highest identifier sequence number and the primary cs.AI count.
# submitters.py
# How many papers come from submitters who filed more than two in the month.
# usage: python submitters.py sept.jsonl 2609.
import json
import sys
from collections import Counter
path, id_prefix = sys.argv[1:3]
rows = [json.loads(line) for line in open(path, encoding="utf-8")]
per = Counter(r["sub"] for r in rows)
print(f"{len(rows):,} papers from {len(per):,} submitter names")
for label, lo, hi in [("1", 1, 1), ("2", 2, 2), ("3-5", 3, 5), ("6-10", 6, 10), ("11-20", 11, 20), ("21+", 21, 10**9)]:
counts = [c for c in per.values() if lo <= c <= hi]
print(f"{label:>6} papers: {len(counts):6,} names, {sum(counts):6,} papers ({sum(counts) / len(rows):5.1%})")
above = sum(max(0, c - 2) for c in per.values())
print(f"papers above the two-a-month line: {above:,} ({above / len(rows):.1%})")
# stricter identity: one name counts as one person only within one primary category
per_cat = Counter((r["sub"], r["cat"]) for r in rows)
above_cat = sum(max(0, c - 2) for c in per_cat.values())
print(f"same count with names split by primary category: {above_cat:,} ({above_cat / len(rows):.1%})")
# the official monthly figure is the highest identifier sequence number issued for that month
seq = max(int(r["id"].split(".")[1]) for r in rows if r["id"].startswith(id_prefix))
print(f"highest {id_prefix}NNNNN sequence number in this harvest: {seq:,}")
print(f"papers whose primary category is cs.AI: {sum(1 for r in rows if r['cat'] == 'cs.AI'):,}")
python submitters.py sept.jsonl 2609.
38,908 papers from 32,194 submitter names
1 papers: 27,555 names, 27,555 papers (70.8%)
2 papers: 3,397 names, 6,794 papers (17.5%)
3-5 papers: 1,156 names, 3,920 papers (10.1%)
6-10 papers: 82 names, 578 papers ( 1.5%)
11-20 papers: 3 names, 40 papers ( 0.1%)
21+ papers: 1 names, 21 papers ( 0.1%)
papers above the two-a-month line: 2,075 (5.3%)
same count with names split by primary category: 822 (2.1%)
highest 2609.NNNNN sequence number in this harvest: 40,363
papers whose primary category is cs.AI: 2,361
Reading the output
Of 32,194 submitter names, 27,555 filed one paper (85.6 percent of names, 70.8 percent of papers) and 3,397 filed two. Together that is 96.1 percent of names inside the limit. The other 1,242 names, 3.9 percent, filed three or more and account for 4,559 papers, 11.7 percent of the month. The papers above the line number 2,075 (5.3 percent), or 822 (2.1 percent) when a name is split by primary category.
Three cautions apply, and the first two pull in opposite directions.
- Names are not people. A submitter name shared by several researchers counts as one heavy submitter, which inflates the tail. In a longer harvest that kept author lists, I checked the 9 names with at least 10 papers in September. Their papers spread over a median of 5 primary categories, only one of the nine had a co-author on at least half its papers, and only two had 80 percent or more of their papers in one category. The top of the name-based tail looks mostly like different people sharing a name, so the maximum of 21 papers should not be read as one prolific researcher.
- Rejections are invisible. Rejected submissions, and ones still on hold, count toward the cap and are absent from the harvest, which pushes the true number up.
- The tail is not a stable roster at this threshold. Of the 1,242 names above two in September, 121 (9.7 percent) were also above two in August, and 513 (41.3 percent) filed anything in August.
So the honest range for the share of September’s announced papers that the cap would have blocked is 2.1 to 5.3 percent. If every one of them had never been filed, September would have had 36,833 announced papers on the loose count or 38,086 on the strict one. On the same announced-paper basis, September 2025 had 26,158 and September 2024 had 20,448. The remaining volume would still be 41 to 46 percent above last year and 1.8 to 1.9 times the level of two years ago.
Our reading: arXiv’s bet is that the harm is concentrated where the quota bites. The public record supports a modest version of that. Names filing three or more papers are 3.9 percent of names and 11.7 percent of papers, a skew of about three to one. It does not support a large version among announced papers, where the top 1 percent of names (321) filed 4.3 percent of papers. Whether the skew is steeper among rejected submissions is something arXiv can see and we cannot.
Where the cap bites
With names split by primary category, 822 papers sit above the line. Computer science supplies 45.3 percent of September’s announced papers and 44.9 percent of the papers above the line (369 of 17,610 cs.* papers, 2.1 percent), so the cap is not an AI-category tax. The three big machine learning categories together have 141 papers above the line out of 7,458 (1.9 percent). The largest counts are in cs.LG (70), cs.CV (52), cs.RO (47) and cs.AI (43). Among categories with at least 300 papers the highest rates are in mathematics and signal processing: math.OC and math.FA at 4.9 percent (31 of 629 and 17 of 350), eess.SP at 4.0 percent (19 of 471) and math.NT at 4.0 percent (25 of 632).
Reading the cap as a rate limiter
The key is the submitter
arXiv is precise about the unit: “All arXiv rate limit policies apply only to submitters.” A quota on an account is only as strong as the cost of another account, and the astronomy blog In the Dark predicted the workaround on day two: “large collaborations will find a way around it by having different co-authors do the submissions.” The data say the room exists. About 82.9 percent of September’s papers list two or more authors (parsed from the author string, so approximate), and so do 79.0 percent of the papers from names above the line. Whether groups will rotate the filing is behaviour, but arXiv’s FAQ already asks co-authors to coordinate.
Our reading: the quota raises the value of an endorsed account, which makes the January endorsement change the control that matters most. A cap with cheap identities buys a second account. A cap behind an identity cost buys a conversation about who files.
The window is a calendar month
The counter resets on the first of the month, which is a fixed window rather than a rolling one. In principle a submitter could file two papers on the last day of a month and two more on the first, and the three-active limit would be the only brake across the boundary, for as long as the first two stay unannounced. That is our reasoning from the published rules, not an arXiv statement. The three-active limit is a concurrency cap, the same idea as the bulkhead in our tutorial on isolating a slow dependency, and it is the quieter of the two limits: arXiv has had it since 2024, and held papers can keep it full.
The accounting has two design details worth copying. The cost is charged at submission, because submission is what consumes moderator time, and it is refunded when a paper is deleted before announcement, which gives authors a way to fix a mistake without losing a slot. Papers held over from a previous month use concurrency but not the new month’s quota, so a backlog in arXiv’s own queue can keep a submitter at the concurrency limit even when they have filed nothing new that month.
The unit is a submission, and replacements are unanswered
In the Dark also raised a question the FAQ does not answer: “whether the limit applies to replacements as well as new submissions.” I could not find an answer in the announcement either. The scale matters. Of September’s 38,908 new papers, 3,694 (9.5 percent) already had a later version by October 3, 4,100 later versions in all (10.5 per 100 papers). For August’s 31,552 papers, 4,844 (15.4 percent) had a later version, 5,720 in all (18.1 per 100), with a month longer to accumulate them. If replacements counted, revision would compete with new work for the same two slots. If they do not, the cap leaves the revision channel open. Either way, researchers should ask before assuming.
What this analysis cannot show
- It cannot see rejected, held or withdrawn submissions. arXiv counts them toward the cap and says submissions, not announced papers, are what consume moderator time, so part of the load it describes sits outside this data.
- It cannot tell people from names. The strict variant (name plus primary category) and the loose one (name only) bracket the answer, and neither is exact.
- It says nothing about quality or about how much of any paper is AI-written. arXiv attributes the growth to AI tools and describes an increase in “dense, AI-written papers” and in “salami” papers, where one piece of work is split into several. Those are moderators’ observations, and I did not test them.
- It covers the last two months. Older months cannot be rebuilt this way, because the OAI datestamp moves forward whenever a record is revised, so a past month’s harvest would miss papers revised later.
- The author counts are parsed from free text, and the OAI documentation warns that datestamps can also reflect administrative and bibliographic updates. The harvest still matched the API’s count exactly.
What to do with it
If you file papers
- Count submissions, not announcements. A rejected paper uses a slot, and a paper on hold counts toward the three active ones until it clears.
- Decide the slots before a deadline. Papers held from last month do not use the new month’s two, but they do use the three.
- If you spot a mistake after submitting, delete the submission before it is announced. arXiv says that frees the slot.
- In a group, agree who files what. The limit follows the filing account, not the author list.
- Ask arXiv whether replacements count.
If you design a quota
- Measure the distribution first. arXiv’s limit of two sits above what 96.1 percent of September’s names did, so it is a tail limit by construction.
- Count what costs you, and refund what never reached the queue. arXiv counts every submission because even rejected ones consume moderator time, and it does not count one deleted before announcement.
- Pair a rate with a concurrency cap, and say how each one treats items that are stuck.
- Put an identity cost behind the quota, or the quota becomes a reason to open another account.
- Publish the distribution after launch so outsiders can check whether the bet paid off.
What to watch next
arXiv says it will keep monitoring the effect on authors, readers and the corpus “and modify as needed.” The statistics page shows each month as it fills, so October and November will show whether the curve bends. The signals that matter more are not on that page: how many submitters hit the limit, how many submissions are rejected, whether replacements count, and whether anyone at arXiv publishes the distribution that would let the tail claim be checked. Terence Tao’s one-sentence post of October 1 said the change “restricts submissions to at most two a month for each submitter.” The measurement that would settle the question is arXiv’s to publish.
Related reading on this site
- Red Hat’s Faster-Coding Warning Turns AI Productivity Into a Measurement Problem covers the same asymmetry in software delivery: producing artifacts gets faster while review capacity stays fixed.
- Stanford’s Paper2Agent Turns Published Papers Into an Attribution Problem looks at what happens to a paper after it is published.
- OpenAI’s Research Acceleration Report Turns Recursive Self-Improvement Into a Self-Graded Metric is a lab’s own account of how AI speeds up its research.
- How to Build a Bulkhead in Python to Stop a Slow Dependency From Starving the Rest has the concurrency-cap code behind the three-active idea.
Method notes
The figures above come from arXiv’s statistics CSV, the public API and the OAI-PMH interface, read on October 3, 2026. The API counts use submittedDate ranges, and the category counts use cat: queries, which match primary and cross-listed papers. For the primary-category counts I paged through the API results and counted each entry’s primary category with the script below.
# api_primary.py
# Primary-category counts for big CS categories, September of 2024, 2025, 2026, by paging the public
# arXiv API (cat: matches primary and cross-listed papers; each entry carries its primary category).
import json
import re
import time
import urllib.parse
import urllib.request
UA = "sxz.io-research-script/1.0 (+https://sxz.io)"
MONTHS = [("2024-09", "202409010000", "202409302359"),
("2025-09", "202509010000", "202509302359"),
("2026-09", "202609010000", "202609302359")]
CATS = ["cs.AI", "cs.LG", "cs.CL"]
PAGE = 2000
def fetch(cat, a, b, start):
q = "cat:%s AND submittedDate:[%s TO %s]" % (cat, a, b)
url = ("https://export.arxiv.org/api/query?search_query=" + urllib.parse.quote(q) +
"&sortBy=submittedDate&sortOrder=ascending&start=%d&max_results=%d" % (start, PAGE))
for _ in range(6):
try:
req = urllib.request.Request(url, headers={"User-Agent": UA})
with urllib.request.urlopen(req, timeout=180) as resp:
return resp.read().decode("utf-8", "replace")
except Exception as e:
print("retry", cat, a, start, repr(e), flush=True)
time.sleep(10)
raise RuntimeError("failed")
res = {}
for label, a, b in MONTHS:
res[label] = {}
for cat in CATS:
total = None
primaries = []
start = 0
while total is None or start < total:
text = fetch(cat, a, b, start)
if total is None:
total = int(re.search(r"<opensearch:totalResults[^>]*>(\d+)<", text).group(1))
primaries += re.findall(r'<arxiv:primary_category[^>]*term="([^"]+)"', text)
start += PAGE
time.sleep(3.2)
res[label][cat] = {"any_listing": total, "entries_read": len(primaries), "primary_is_cat": sum(1 for p in primaries if p == cat)}
print(label, cat, res[label][cat], flush=True)
with open("out/api_primary.json", "w", encoding="utf-8", newline="\n") as f:
json.dump(res, f, indent=1)
print("DONE", flush=True)
python api_primary.py
2024-09 cs.AI {'any_listing': 2650, 'entries_read': 2650, 'primary_is_cat': 392}
2024-09 cs.LG {'any_listing': 2966, 'entries_read': 2966, 'primary_is_cat': 1382}
2024-09 cs.CL {'any_listing': 1502, 'entries_read': 1502, 'primary_is_cat': 1059}
2025-09 cs.AI {'any_listing': 4262, 'entries_read': 4262, 'primary_is_cat': 896}
2025-09 cs.LG {'any_listing': 4203, 'entries_read': 4203, 'primary_is_cat': 2333}
2025-09 cs.CL {'any_listing': 2234, 'entries_read': 2234, 'primary_is_cat': 1587}
2026-09 cs.AI {'any_listing': 5923, 'entries_read': 5923, 'primary_is_cat': 2361}
The script stopped after the line for cs.AI in 2026. The API had started answering 429 to the next request, the September 2026 cs.LG count, and the retries ran out. I took the September 2026 cs.LG and cs.CL figures from the harvest instead, where the primary category is a field of every record (3,233 and 1,864). The API and the harvest agree on the one 2026 count that both produced: 2,361 for cs.AI.
arXiv’s terms ask for at most one request every three seconds on one connection. The shown scripts pace themselves that way, but in places I ran an API script and the harvest side by side, which went past that guidance, and the 429 and 500 errors followed. I ran the final harvest alone, and its output matched my first run exactly. The checks that need more than the shown fields (the August identifier check, the persistence count, the name-collision check, the category breakdown, the revision counts and the author counts) used a longer harvest from August 1 that kept author lists and version counts. Everything in this piece is aggregate: no submitter name appears in it, and the harvest files were deleted after counting.








No Comment! Be the first one.