GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
GitHub says its Spokes design makes read scale tax every write. The replacement splits durable storage from caching workers and keeps only the ref update under agreement, and a local Git experiment...
GitHub says the system that stores Git repositories has a built-in ceiling for the agent era, and that it is rebuilding that system while the service keeps running. In a post published on October 6 and updated on October 7, principal software engineer Brian Celenza writes that “the mechanism we use for durability is the same one we use for scale.” Each repository sits on several fileservers, so every extra read replica is also an extra copy that every push has to reach. The replacement keeps the authoritative copy in Azure Blob Storage, serves reads from lightweight caching workers, and limits agreement to the one step of a push that truly needs it: the reference update.
Table Of Content
- The workload GitHub is building for
- Why the current design hits a ceiling
- What Spokes does today
- How GitHub described the same design in 2016 and 2017
- What the replacement changes
- Why the reference update is the step that needs agreement
- What a push sends
- What happens when two agents race
- A local experiment: what a hot ref costs
- Two bets for agent-era Git
- What the post leaves open
- What to do with this today
GitHub reports “up to 35 times higher write throughput” in internal benchmarks and promises a follow-up post on the architecture. This piece reads the announcement against GitHub’s own 2016 and 2017 write-ups of the current design, Git’s protocol documentation and a local experiment of our own. The experiment runs plain Git 2.55.0 on a Windows 11 machine, so it says nothing about GitHub’s implementation. What it shows is why the step GitHub wants to keep coordinating is the one where concurrent agents collide.
The workload GitHub is building for
The post opens with volume. Between September 2025 and August 2026, GitHub says, Git activity “increased to more than 2x its previous level, from 218.2 billion events per month to 473.3 billion.” The table lists the figures the post gives, with an average per second that we computed ourselves: the monthly figure divided by the seconds in the stated month, or in a 30-day month where the post names none.
| Figure in the post | Value | Average per second (our arithmetic) |
|---|---|---|
| Git events per month | 218.2 billion to 473.3 billion between September 2025 and August 2026 | about 177,000 (assuming the latest figure is August, 31 days) |
| Commits in September 2026 | 7.38 billion, “more than five times as many as a year earlier” | about 2,850 |
| Pushes per month | 0.69 billion to 3.35 billion, “4.9x year over year” | about 1,290 (30-day month) |
| GitHub Actions runs in September | 3.26 billion, “more than 4x as many as a year ago” | about 1,260 |
| Requests to the busiest repository in August | “roughly a billion” | about 370 |
| Pull request merges | “nearly 4x their volume a year ago” | no absolute figure given |
Two cautions. Averages hide peaks, and the post gives no peak rates. And if the push and commit figures describe the same month, they imply about 2.2 commits per push, which fits the post’s picture of agents that commit “after nearly every action” and are limited by “how fast a single push completes.” That ratio is our inference; the post does not pair the two figures.
The post names five pressure points: commit turnaround per agent, write throughput, merges that “contend on one reference,” the fan-out of reads from CI and code scanning, and the compaction and cleanup that every new write adds to. Reads, it says, are “relatively easy to scale.” Writes are “way harder”: “Every push has to be stored durably and made visible consistently before the next agent or CI job can build on it.”
Why the current design hits a ceiling
What Spokes does today
According to the post, every repository is stored by Spokes, “which keeps a full copy on the local disks of several fileservers, five by default.” When a push updates a reference, “a three-phase commit protocol uses a quorum” so that CI, the web UI and API clients see a consistent state, and that pairing “serves a billion repositories today.” (Wikipedia says three-phase commit “improves upon the two-phase commit protocol” by removing one way it can block indefinitely; our two-phase commit tutorial builds the original from scratch.)
The trouble is the coupling. In the post’s words: “The copies on disk are the source of truth, so adding read capacity means adding another durable replica. Every replica participates in every write, so a push is only as fast as the slowest replica in its set.” The net effect: “adding replicas to absorb read load makes writes slower.” At the highest activity levels, “losing a replica reduces read capacity, and losing quorum stops writes entirely.”
A rough sense of the write amplification: at the default of five copies, 3.35 billion pushes a month would mean 16.75 billion writes to individual replicas a month. That is our arithmetic for the default; the post does not say how many copies each repository has.
How GitHub described the same design in 2016 and 2017
| Date | GitHub post | What it says |
|---|---|---|
| April 5, 2016 | Introducing DGit (Patrick Reynolds) | The design uses Git’s distributed nature “to keep three copies of every repository, on three different servers.” DGit was later renamed Spokes. |
| September 7, 2016 | Building resilience in Spokes (Patrick Reynolds) | At least three copies of every repository. Spokes serializes writes by ensuring that “every write acquires an exclusive lock on a majority of replicas,” and so “eliminates conflicts by eliminating concurrent writes entirely.” |
| October 13, 2017 | Stretching Spokes (Michael Haggerty) | Replication across distant datacenters uses “the three-phase commit protocol to update the replicas,” which “costs four round-trips to the distant replicas.” |
| October 6, 2026 | Building Git infrastructure for agent-scale development (Brian Celenza) | Five copies by default, a three-phase commit with a quorum, and the ceiling described above. |
Read together, the posts show the pattern the 2026 post wants to leave behind. The 2017 post puts a network round trip across the continental US at “something like 60-80 milliseconds,” so four of them cost 240 to 320 milliseconds before any other work. That is 2017 arithmetic, not a measurement of today’s system, and the post also says GitHub planned “to reduce the number of round trips through the use of a more advanced consensus algorithm.” The 2026 post frames the replacement the other way around: coordinating only the reference update “shrinks the critical path of a push to the small step that needs coordination.”
What the replacement changes
The post names two tenets, minimize coordination and decouple storage from compute. Set against the current design as the post describes it:
| Question | Spokes today | The replacement |
|---|---|---|
| Where is the authoritative copy? | “a full copy on the local disks of several fileservers, five by default” | “Authoritative repository data lives in Azure Blob Storage” |
| How does read capacity grow? | “adding read capacity means adding another durable replica” | “read capacity comes from lightweight workers that cache data to serve requests” |
| What does a push wait for? | A three-phase commit with a quorum; “a push is only as fast as the slowest replica in its set” | Agreement on the reference update; storing objects, validating connectivity and secret scanning, “most of it can happen in parallel to other writes” |
| Where does maintenance run? | Compaction and garbage collection “run on the same hosts that answer live Git requests” | “separate workers handle maintenance directly against durable storage” |
| What does losing a host cost? | Read capacity, and “recovery means rebuilding a full repository copy” | “closer to a cache miss” |
GitHub adds that there is no maintenance window in the plan and that it will not ask people to change how they build software while the work proceeds. The one number it offers for the result is the 35-times write-throughput figure, which we come back to below.
Why the reference update is the step that needs agreement
Git’s own design explains why GitHub singles out the reference update. Pro Git opens its section on Git objects with “Git is a content-addressable filesystem,” a store where each object is filed under a key derived from its content. Writing the same object twice is idempotent, and different objects have different names. A branch name is the opposite: one mutable pointer that every writer wants to move. (Counters can be merged without any agreement, as our CRDT counter tutorial shows. A branch pointer cannot, because two commits cannot both be the tip.)
What a push sends
The wire protocol shows the split. After the server advertises its current references, the client sends update requests of the form old-id SP new-id SP name, the grammar in Git’s gitprotocol-pack documentation, followed by the pack of objects. The server answers with an unpack status and a result per reference: “ok [refname]” or “ng [refname] [error].” The git update-ref documentation describes the same compare-and-swap in plumbing form: given an old id, it moves the branch to the new id “only if its current value is” the old one. We captured one real push with GIT_TRACE_PACKET=1 (capability lists shortened):
exit code: 0
push< 42d8732a7f39bb76fc3868f5c868d171d788c990 refs/heads/main\0 [capabilities shortened]
push< 0000
push> 42d8732a7f39bb76fc3868f5c868d171d788c990 667a9bda155552f8b273bc791d117c7348fcbc5c refs/heads/main\0 [capabilities shortened]
push> 0000
push< unpack ok
push< ok refs/heads/main
push< 0000
The client names the old value it expects, sends the objects, and the server reports unpack ok before it reports ok refs/heads/main. The objects are stored first and the reference moves last.
What happens when two agents race
To see the compare half fail, we gave a bare repository a pre-receive hook that only sleeps for a second, so that two pushes overlap. Agent A and agent B each made one commit from the same starting tip, and B started 150 milliseconds after A. Both had already read the same branch tip when the hook slept.
branch tip before the race: fef4967c09
--- agent A: exit 0, 2.49 s
To C:\labtmp\race\origin.git
fef4967..a76f853 HEAD -> main
--- agent B: exit 1, 2.54 s
remote: error: cannot lock ref 'refs/heads/main': is at a76f8531e46042733e1f026a08e5281d9f9608dc but expected fef4967c090b1a881f70527a8673ba1556972531
To C:\labtmp\race\origin.git
! [remote rejected] HEAD -> main (incorrect old value provided)
error: failed to push some refs to 'C:\labtmp\race\origin.git'
branch tip after the race: a76f8531e4
loose objects in the server repo: count: 9
B's commit object stored on the server: commit
B's commit reachable from main: False
A won. B’s push was refused by the server’s check that the branch is still where B expected it to be, and Git reported incorrect old value provided. Yet B’s commit object is in the server repository: it was unpacked and stored before the reference update failed, and it is simply unreachable from main. That is the property GitHub’s first tenet leans on. The expensive part of a push can overlap with other pushes, and only the final compare-and-swap has to be serialized. It also means every lost race leaves unreachable objects behind, which is the kind of cleanup GitHub wants off the serving path.
A local experiment: what a hot ref costs
The post says merges “contend on one reference.” We measured what that contention costs in plain Git. Each run creates a bare repository, then one clone per agent. Every agent commits one new file and pushes at the same instant; a barrier releases all the threads together. A rejected agent fetches, rebases onto the new tip and pushes again at once, with no backoff. Three modes: every agent pushing to its own branch; every agent pushing to main and retrying; and every agent pushing to main one at a time behind a lock, the same ordering a merge queue imposes. Each cell below is the median of three runs on one Windows 11 machine with Git 2.55.0.windows.3 and a repository on local disk. This is the core of the script (the rm helper, the thread runner and the reporting are left out):
# hotref.py (core of the lab script; the rm helper, thread runner and reporting are left out)
def git(cwd, *args):
p = subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True)
return p.returncode, p.stdout + p.stderr
def setup(n):
rm(LAB)
LAB.mkdir(parents=True)
origin = LAB / "origin.git"
git(LAB, "init", "--bare", "-q", "-b", "main", str(origin))
seed = LAB / "seed"
git(LAB, "clone", "-q", str(origin), str(seed))
git(seed, "config", "user.name", "seed")
git(seed, "config", "user.email", "[email protected]")
(seed / "README.md").write_text("shared repo\n")
git(seed, "add", "-A")
git(seed, "commit", "-q", "-m", "seed")
git(seed, "push", "-q", "origin", "HEAD:main")
agents = []
for i in range(n):
d = LAB / f"agent-{i:02d}"
git(LAB, "clone", "-q", str(origin), str(d))
git(d, "config", "user.name", f"agent-{i:02d}")
git(d, "config", "user.email", f"a{i}@example.com")
agents.append(d)
return origin, agents
def land(i, d, mode, barrier, lock, out):
(d / f"task-{i:02d}.txt").write_text(f"agent {i} output\n")
git(d, "add", "-A")
git(d, "commit", "-q", "-m", f"agent {i}")
dest = f"agent-{i:02d}" if mode == "own-branch" else "main"
barrier.wait()
t0 = time.perf_counter()
attempts, msgs = 0, []
while True:
if mode == "serialized":
lock.acquire()
try:
if mode == "serialized":
git(d, "pull", "--rebase", "-q", "origin", "main")
attempts += 1
rc, text = git(d, "push", "-q", "origin", f"HEAD:refs/heads/{dest}")
finally:
if mode == "serialized":
lock.release()
if rc == 0:
break
msgs.append(text.strip())
git(d, "pull", "--rebase", "-q", "origin", "main")
out[i] = (t0, time.perf_counter(), attempts, msgs)
The full script starts one thread per agent behind a barrier, checks that the origin ends with one more commit than agents (or one more branch), and prints the table. Run as python hotref.py 4,8,16,32 3, it covers four team sizes with three runs each.
| Agents | How they land one commit each | Median time until all have landed | Pushes sent | Pushes rejected | Most tries by one agent |
|---|---|---|---|---|---|
| 4 | Each agent on its own branch | 0.4 s | 4 | 0 | 1 |
| 4 | All on main, retry at once | 2.6 s | 10 | 6 | 4 |
| 4 | All on main, one at a time | 2.8 s | 4 | 0 | 1 |
| 8 | Each agent on its own branch | 0.4 s | 8 | 0 | 1 |
| 8 | All on main, retry at once | 6.5 s | 36 | 28 | 8 |
| 8 | All on main, one at a time | 6.0 s | 8 | 0 | 1 |
| 16 | Each agent on its own branch | 0.6 s | 16 | 0 | 1 |
| 16 | All on main, retry at once | 15.4 s | 136 | 120 | 16 |
| 16 | All on main, one at a time | 13.3 s | 16 | 0 | 1 |
| 32 | Each agent on its own branch | 1.2 s | 32 | 0 | 1 |
| 32 | All on main, retry at once | 48.1 s | 527 | 495 | 32 |
| 32 | All on main, one at a time | 30.3 s | 32 | 0 | 1 |
With no shared reference, every push succeeded on the first try: 4, 8, 16 and 32 pushes for 4, 8, 16 and 32 agents, none rejected, and the whole group landed in 0.4 to 1.2 seconds. With one shared branch and immediate retries, the number of pushes was exactly n(n+1)/2 for n agents (10, 36 and 136) at 4, 8 and 16 agents in every run, and 527 in every run at 32 agents, one short of the 528 the formula gives. Each round lets exactly one push through. Every loser fetches, rebases and tries again, so the last agent to land has pushed n times, and 495 of the 527 pushes at 32 agents were rejected. Landing the same 32 commits took about 48 seconds, roughly 40 times as long as on separate branches.
Serializing the agents behind a lock removed the waste: one push per agent and none rejected. At 4 agents the queue was slightly slower than the retry storm (2.8 seconds against 2.6), because every landing inside the lock includes a fetch, a rebase and a push. From 8 agents on it was faster: 6.0 seconds against 6.5, 13.3 against 15.4, and 30.3 against 48.1 at 32 agents, with 32 pushes instead of 527. The queue saves wasted work first and time second, and its time still grows a little faster than the number of commits (0.7 seconds per commit at 4 agents, 0.95 at 32) because everything inside the lock is serial. In our lab the lock covers a fetch, a rebase and a push. GitHub’s first tenet is to shrink what sits on the serialized path to the reference update alone.
The rejections came in two forms. Some pushes were refused by the client itself, after it read the server’s advertisement and found a branch tip it had not fetched, which Git reports as fetch first. The rest passed that check and lost at the server, as in the race demo above. Across the twelve shared-branch runs, 1,543 of the 1,947 rejections (79 percent) were lost at the server and 404 (21 percent) were refused by the client. A push refused by the server had already sent its objects.
The limits of the experiment matter. A local transport has no network latency, while on a real network every rejected push also pays a round trip. There are no server-side checks such as secret scanning. The agents retried at once and in lock step, which is the worst case; a jittered backoff would break the lock step, and we did not test one. The shape of the push counts, n for the queue against n(n+1)/2 for the retry storm, follows from how a compare-and-swap on one reference behaves when every loser retries at once. The seconds are a property of this machine.
GitHub’s existing answer for a busy trunk is the merge queue, which its documentation says helps “by automating pull request merges into a busy branch and ensuring the branch is never broken by incompatible changes.” It builds each queued pull request on “temporary branches that are created on your behalf by a merge queue” (the names start with gh-readonly-queue/), and a Build concurrency setting “between 1 and 100” throttles concurrent CI builds, which “affects the velocity of merges that a merge queue can complete.” The 2026 post does not say how the new architecture changes that queue.
Two bets for agent-era Git
GitHub says it is building for workloads that include “an organization running thousands of agents against a single codebase.” Cloudflare’s Artifacts documentation takes the opposite starting point. Its best-practices page says “Create one repo for each unit of autonomous work. If you have 10,000 agents, create 10,000 repos,” and “Do not use one shared repo as a queue for many autonomous agents.” The stated reason is that separate repositories avoid “turning one shared repo into a hot spot for conflicts, large diffs, and accidental overwrites.”
Our experiment is the plain-Git version of that advice. When agents do not share a reference, nothing is rejected. When they do, the number of pushes rises with the square of the team. But work that is split across repositories still has to be merged somewhere, and, as our earlier look at Artifacts noted, Cloudflare leaves the question of how those repositories get reviewed and merged back to others, and paired its launch with a contest to build the next Git platform on top of the service. The size limits point the same way: Artifacts documents a 1 GB maximum per repository, while GitHub’s repository limits page recommends staying within 10 GB on disk, and the two pages do not say they measure the same thing. Cloudflare partitions the work so that there is no hot reference. GitHub keeps one codebase and aims to make the shared reference cheap. Most teams will need both: partition where work is independent, and make the shared reference cheap where it is not.
What the post leaves open
- The 35 times figure. The phrase is “up to,” from “internal benchmarks,” and the post gives no workload, repository size, replica comparison, hardware or latency percentile. It does not say whether the comparison is one busy repository or the whole fleet.
- When. GitHub says it is “already putting that foundation in place” and that the next post will “dive deeper into our future architecture,” but it names no date for any repository to move.
- How agreement on the reference update is reached. The old design used a quorum; the post does not say what replaces it. Nor does it say how read workers that “cache data” learn about a new tip before the next CI job reads, although the post itself sets the requirement that a push be “made visible consistently before the next agent or CI job can build on it.”
- Which Azure redundancy tier holds the copy. Microsoft’s storage redundancy documentation gives “at least 99.999999999% (11 9s)” durability of objects over a given year for locally redundant storage, “at least 99.9999999999% (12 9s)” for zone-redundant storage and “at least 99.99999999999999% (16 9s)” for geo-redundant storage. The post says only “durability and replication at Azure scale.” Between the lowest and highest of those tiers the annual loss probability differs by five orders of magnitude, and durability is a different property from the write latency GitHub is trying to cut.
- Who gets it. The post describes GitHub’s platform and says nothing about GitHub Enterprise Server or other deployment types.
What to do with this today
- Keep agent work off the shared reference. One branch or repository per agent, with a queue in front of trunk. In our runs, the mode with no shared reference never had a push rejected.
- Do not let a fleet retry against one reference in lock step. Our runs spent a quadratic number of pushes on it. Backoff with jitter is standard practice, but we did not test it, so measure it on your own setup.
- Size the merge queue to your CI budget. The Build concurrency setting is the documented throttle, and the documentation says it also affects how fast merges can complete.
- Keep repositories lean. GitHub’s limits page recommends 10 GB on disk and warns that “Large repositories can slow down fetch operations and increase clone times for developers and CI.” Fan-out reads are the other half of the agent workload.
- Watch for the follow-up post. The observable test of the 35 times claim is write latency and rejection rates for a busy repository under concurrent agents, before and after a move.
The announcement is a design statement, not a rollout. Its most useful claim is also its smallest: only the reference update needs agreement. Git’s push protocol already separates the two halves, and a few dozen lines of Python show what happens when a fleet of agents shares one branch.








No Comment! Be the first one.