TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 8, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 8, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
08 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Articles/GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Articles

GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems

GitHub says its Spokes design makes read scale tax every write. The replacement splits durable storage from caching workers and keeps only the ref update under agreement, and a local Git experiment...

October 8, 2026 14 Min Read
7

GitHub says the system that stores Git repositories has a built-in ceiling for the agent era, and that it is rebuilding that system while the service keeps running. In a post published on October 6 and updated on October 7, principal software engineer Brian Celenza writes that “the mechanism we use for durability is the same one we use for scale.” Each repository sits on several fileservers, so every extra read replica is also an extra copy that every push has to reach. The replacement keeps the authoritative copy in Azure Blob Storage, serves reads from lightweight caching workers, and limits agreement to the one step of a push that truly needs it: the reference update.

Table Of Content

  • The workload GitHub is building for
  • Why the current design hits a ceiling
  • What Spokes does today
  • How GitHub described the same design in 2016 and 2017
  • What the replacement changes
  • Why the reference update is the step that needs agreement
  • What a push sends
  • What happens when two agents race
  • A local experiment: what a hot ref costs
  • Two bets for agent-era Git
  • What the post leaves open
  • What to do with this today

GitHub reports “up to 35 times higher write throughput” in internal benchmarks and promises a follow-up post on the architecture. This piece reads the announcement against GitHub’s own 2016 and 2017 write-ups of the current design, Git’s protocol documentation and a local experiment of our own. The experiment runs plain Git 2.55.0 on a Windows 11 machine, so it says nothing about GitHub’s implementation. What it shows is why the step GitHub wants to keep coordinating is the one where concurrent agents collide.

The workload GitHub is building for

The post opens with volume. Between September 2025 and August 2026, GitHub says, Git activity “increased to more than 2x its previous level, from 218.2 billion events per month to 473.3 billion.” The table lists the figures the post gives, with an average per second that we computed ourselves: the monthly figure divided by the seconds in the stated month, or in a 30-day month where the post names none.

Figure in the post Value Average per second (our arithmetic)
Git events per month 218.2 billion to 473.3 billion between September 2025 and August 2026 about 177,000 (assuming the latest figure is August, 31 days)
Commits in September 2026 7.38 billion, “more than five times as many as a year earlier” about 2,850
Pushes per month 0.69 billion to 3.35 billion, “4.9x year over year” about 1,290 (30-day month)
GitHub Actions runs in September 3.26 billion, “more than 4x as many as a year ago” about 1,260
Requests to the busiest repository in August “roughly a billion” about 370
Pull request merges “nearly 4x their volume a year ago” no absolute figure given

Two cautions. Averages hide peaks, and the post gives no peak rates. And if the push and commit figures describe the same month, they imply about 2.2 commits per push, which fits the post’s picture of agents that commit “after nearly every action” and are limited by “how fast a single push completes.” That ratio is our inference; the post does not pair the two figures.

The post names five pressure points: commit turnaround per agent, write throughput, merges that “contend on one reference,” the fan-out of reads from CI and code scanning, and the compaction and cleanup that every new write adds to. Reads, it says, are “relatively easy to scale.” Writes are “way harder”: “Every push has to be stored durably and made visible consistently before the next agent or CI job can build on it.”

Why the current design hits a ceiling

What Spokes does today

According to the post, every repository is stored by Spokes, “which keeps a full copy on the local disks of several fileservers, five by default.” When a push updates a reference, “a three-phase commit protocol uses a quorum” so that CI, the web UI and API clients see a consistent state, and that pairing “serves a billion repositories today.” (Wikipedia says three-phase commit “improves upon the two-phase commit protocol” by removing one way it can block indefinitely; our two-phase commit tutorial builds the original from scratch.)

The trouble is the coupling. In the post’s words: “The copies on disk are the source of truth, so adding read capacity means adding another durable replica. Every replica participates in every write, so a push is only as fast as the slowest replica in its set.” The net effect: “adding replicas to absorb read load makes writes slower.” At the highest activity levels, “losing a replica reduces read capacity, and losing quorum stops writes entirely.”

A rough sense of the write amplification: at the default of five copies, 3.35 billion pushes a month would mean 16.75 billion writes to individual replicas a month. That is our arithmetic for the default; the post does not say how many copies each repository has.

How GitHub described the same design in 2016 and 2017

Date GitHub post What it says
April 5, 2016 Introducing DGit (Patrick Reynolds) The design uses Git’s distributed nature “to keep three copies of every repository, on three different servers.” DGit was later renamed Spokes.
September 7, 2016 Building resilience in Spokes (Patrick Reynolds) At least three copies of every repository. Spokes serializes writes by ensuring that “every write acquires an exclusive lock on a majority of replicas,” and so “eliminates conflicts by eliminating concurrent writes entirely.”
October 13, 2017 Stretching Spokes (Michael Haggerty) Replication across distant datacenters uses “the three-phase commit protocol to update the replicas,” which “costs four round-trips to the distant replicas.”
October 6, 2026 Building Git infrastructure for agent-scale development (Brian Celenza) Five copies by default, a three-phase commit with a quorum, and the ceiling described above.

Read together, the posts show the pattern the 2026 post wants to leave behind. The 2017 post puts a network round trip across the continental US at “something like 60-80 milliseconds,” so four of them cost 240 to 320 milliseconds before any other work. That is 2017 arithmetic, not a measurement of today’s system, and the post also says GitHub planned “to reduce the number of round trips through the use of a more advanced consensus algorithm.” The 2026 post frames the replacement the other way around: coordinating only the reference update “shrinks the critical path of a push to the small step that needs coordination.”

What the replacement changes

The post names two tenets, minimize coordination and decouple storage from compute. Set against the current design as the post describes it:

Question Spokes today The replacement
Where is the authoritative copy? “a full copy on the local disks of several fileservers, five by default” “Authoritative repository data lives in Azure Blob Storage”
How does read capacity grow? “adding read capacity means adding another durable replica” “read capacity comes from lightweight workers that cache data to serve requests”
What does a push wait for? A three-phase commit with a quorum; “a push is only as fast as the slowest replica in its set” Agreement on the reference update; storing objects, validating connectivity and secret scanning, “most of it can happen in parallel to other writes”
Where does maintenance run? Compaction and garbage collection “run on the same hosts that answer live Git requests” “separate workers handle maintenance directly against durable storage”
What does losing a host cost? Read capacity, and “recovery means rebuilding a full repository copy” “closer to a cache miss”

GitHub adds that there is no maintenance window in the plan and that it will not ask people to change how they build software while the work proceeds. The one number it offers for the result is the 35-times write-throughput figure, which we come back to below.

Why the reference update is the step that needs agreement

Git’s own design explains why GitHub singles out the reference update. Pro Git opens its section on Git objects with “Git is a content-addressable filesystem,” a store where each object is filed under a key derived from its content. Writing the same object twice is idempotent, and different objects have different names. A branch name is the opposite: one mutable pointer that every writer wants to move. (Counters can be merged without any agreement, as our CRDT counter tutorial shows. A branch pointer cannot, because two commits cannot both be the tip.)

What a push sends

The wire protocol shows the split. After the server advertises its current references, the client sends update requests of the form old-id SP new-id SP name, the grammar in Git’s gitprotocol-pack documentation, followed by the pack of objects. The server answers with an unpack status and a result per reference: “ok [refname]” or “ng [refname] [error].” The git update-ref documentation describes the same compare-and-swap in plumbing form: given an old id, it moves the branch to the new id “only if its current value is” the old one. We captured one real push with GIT_TRACE_PACKET=1 (capability lists shortened):

exit code: 0
   push< 42d8732a7f39bb76fc3868f5c868d171d788c990 refs/heads/main\0 [capabilities shortened]
   push< 0000
   push> 42d8732a7f39bb76fc3868f5c868d171d788c990 667a9bda155552f8b273bc791d117c7348fcbc5c refs/heads/main\0 [capabilities shortened]
   push> 0000
   push< unpack ok
   push< ok refs/heads/main
   push< 0000

The client names the old value it expects, sends the objects, and the server reports unpack ok before it reports ok refs/heads/main. The objects are stored first and the reference moves last.

What happens when two agents race

To see the compare half fail, we gave a bare repository a pre-receive hook that only sleeps for a second, so that two pushes overlap. Agent A and agent B each made one commit from the same starting tip, and B started 150 milliseconds after A. Both had already read the same branch tip when the hook slept.

branch tip before the race: fef4967c09
--- agent A: exit 0, 2.49 s
    To C:\labtmp\race\origin.git
       fef4967..a76f853  HEAD -> main
--- agent B: exit 1, 2.54 s
    remote: error: cannot lock ref 'refs/heads/main': is at a76f8531e46042733e1f026a08e5281d9f9608dc but expected fef4967c090b1a881f70527a8673ba1556972531
    To C:\labtmp\race\origin.git
     ! [remote rejected] HEAD -> main (incorrect old value provided)
    error: failed to push some refs to 'C:\labtmp\race\origin.git'
branch tip after the race: a76f8531e4
loose objects in the server repo: count: 9
B's commit object stored on the server: commit
B's commit reachable from main: False

A won. B’s push was refused by the server’s check that the branch is still where B expected it to be, and Git reported incorrect old value provided. Yet B’s commit object is in the server repository: it was unpacked and stored before the reference update failed, and it is simply unreachable from main. That is the property GitHub’s first tenet leans on. The expensive part of a push can overlap with other pushes, and only the final compare-and-swap has to be serialized. It also means every lost race leaves unreachable objects behind, which is the kind of cleanup GitHub wants off the serving path.

A local experiment: what a hot ref costs

The post says merges “contend on one reference.” We measured what that contention costs in plain Git. Each run creates a bare repository, then one clone per agent. Every agent commits one new file and pushes at the same instant; a barrier releases all the threads together. A rejected agent fetches, rebases onto the new tip and pushes again at once, with no backoff. Three modes: every agent pushing to its own branch; every agent pushing to main and retrying; and every agent pushing to main one at a time behind a lock, the same ordering a merge queue imposes. Each cell below is the median of three runs on one Windows 11 machine with Git 2.55.0.windows.3 and a repository on local disk. This is the core of the script (the rm helper, the thread runner and the reporting are left out):

# hotref.py (core of the lab script; the rm helper, thread runner and reporting are left out)
def git(cwd, *args):
    p = subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True)
    return p.returncode, p.stdout + p.stderr


def setup(n):
    rm(LAB)
    LAB.mkdir(parents=True)
    origin = LAB / "origin.git"
    git(LAB, "init", "--bare", "-q", "-b", "main", str(origin))
    seed = LAB / "seed"
    git(LAB, "clone", "-q", str(origin), str(seed))
    git(seed, "config", "user.name", "seed")
    git(seed, "config", "user.email", "[email protected]")
    (seed / "README.md").write_text("shared repo\n")
    git(seed, "add", "-A")
    git(seed, "commit", "-q", "-m", "seed")
    git(seed, "push", "-q", "origin", "HEAD:main")
    agents = []
    for i in range(n):
        d = LAB / f"agent-{i:02d}"
        git(LAB, "clone", "-q", str(origin), str(d))
        git(d, "config", "user.name", f"agent-{i:02d}")
        git(d, "config", "user.email", f"a{i}@example.com")
        agents.append(d)
    return origin, agents


def land(i, d, mode, barrier, lock, out):
    (d / f"task-{i:02d}.txt").write_text(f"agent {i} output\n")
    git(d, "add", "-A")
    git(d, "commit", "-q", "-m", f"agent {i}")
    dest = f"agent-{i:02d}" if mode == "own-branch" else "main"
    barrier.wait()
    t0 = time.perf_counter()
    attempts, msgs = 0, []
    while True:
        if mode == "serialized":
            lock.acquire()
        try:
            if mode == "serialized":
                git(d, "pull", "--rebase", "-q", "origin", "main")
            attempts += 1
            rc, text = git(d, "push", "-q", "origin", f"HEAD:refs/heads/{dest}")
        finally:
            if mode == "serialized":
                lock.release()
        if rc == 0:
            break
        msgs.append(text.strip())
        git(d, "pull", "--rebase", "-q", "origin", "main")
    out[i] = (t0, time.perf_counter(), attempts, msgs)

The full script starts one thread per agent behind a barrier, checks that the origin ends with one more commit than agents (or one more branch), and prints the table. Run as python hotref.py 4,8,16,32 3, it covers four team sizes with three runs each.

Agents How they land one commit each Median time until all have landed Pushes sent Pushes rejected Most tries by one agent
4 Each agent on its own branch 0.4 s 4 0 1
4 All on main, retry at once 2.6 s 10 6 4
4 All on main, one at a time 2.8 s 4 0 1
8 Each agent on its own branch 0.4 s 8 0 1
8 All on main, retry at once 6.5 s 36 28 8
8 All on main, one at a time 6.0 s 8 0 1
16 Each agent on its own branch 0.6 s 16 0 1
16 All on main, retry at once 15.4 s 136 120 16
16 All on main, one at a time 13.3 s 16 0 1
32 Each agent on its own branch 1.2 s 32 0 1
32 All on main, retry at once 48.1 s 527 495 32
32 All on main, one at a time 30.3 s 32 0 1

With no shared reference, every push succeeded on the first try: 4, 8, 16 and 32 pushes for 4, 8, 16 and 32 agents, none rejected, and the whole group landed in 0.4 to 1.2 seconds. With one shared branch and immediate retries, the number of pushes was exactly n(n+1)/2 for n agents (10, 36 and 136) at 4, 8 and 16 agents in every run, and 527 in every run at 32 agents, one short of the 528 the formula gives. Each round lets exactly one push through. Every loser fetches, rebases and tries again, so the last agent to land has pushed n times, and 495 of the 527 pushes at 32 agents were rejected. Landing the same 32 commits took about 48 seconds, roughly 40 times as long as on separate branches.

Serializing the agents behind a lock removed the waste: one push per agent and none rejected. At 4 agents the queue was slightly slower than the retry storm (2.8 seconds against 2.6), because every landing inside the lock includes a fetch, a rebase and a push. From 8 agents on it was faster: 6.0 seconds against 6.5, 13.3 against 15.4, and 30.3 against 48.1 at 32 agents, with 32 pushes instead of 527. The queue saves wasted work first and time second, and its time still grows a little faster than the number of commits (0.7 seconds per commit at 4 agents, 0.95 at 32) because everything inside the lock is serial. In our lab the lock covers a fetch, a rebase and a push. GitHub’s first tenet is to shrink what sits on the serialized path to the reference update alone.

The rejections came in two forms. Some pushes were refused by the client itself, after it read the server’s advertisement and found a branch tip it had not fetched, which Git reports as fetch first. The rest passed that check and lost at the server, as in the race demo above. Across the twelve shared-branch runs, 1,543 of the 1,947 rejections (79 percent) were lost at the server and 404 (21 percent) were refused by the client. A push refused by the server had already sent its objects.

The limits of the experiment matter. A local transport has no network latency, while on a real network every rejected push also pays a round trip. There are no server-side checks such as secret scanning. The agents retried at once and in lock step, which is the worst case; a jittered backoff would break the lock step, and we did not test one. The shape of the push counts, n for the queue against n(n+1)/2 for the retry storm, follows from how a compare-and-swap on one reference behaves when every loser retries at once. The seconds are a property of this machine.

GitHub’s existing answer for a busy trunk is the merge queue, which its documentation says helps “by automating pull request merges into a busy branch and ensuring the branch is never broken by incompatible changes.” It builds each queued pull request on “temporary branches that are created on your behalf by a merge queue” (the names start with gh-readonly-queue/), and a Build concurrency setting “between 1 and 100” throttles concurrent CI builds, which “affects the velocity of merges that a merge queue can complete.” The 2026 post does not say how the new architecture changes that queue.

Two bets for agent-era Git

GitHub says it is building for workloads that include “an organization running thousands of agents against a single codebase.” Cloudflare’s Artifacts documentation takes the opposite starting point. Its best-practices page says “Create one repo for each unit of autonomous work. If you have 10,000 agents, create 10,000 repos,” and “Do not use one shared repo as a queue for many autonomous agents.” The stated reason is that separate repositories avoid “turning one shared repo into a hot spot for conflicts, large diffs, and accidental overwrites.”

Our experiment is the plain-Git version of that advice. When agents do not share a reference, nothing is rejected. When they do, the number of pushes rises with the square of the team. But work that is split across repositories still has to be merged somewhere, and, as our earlier look at Artifacts noted, Cloudflare leaves the question of how those repositories get reviewed and merged back to others, and paired its launch with a contest to build the next Git platform on top of the service. The size limits point the same way: Artifacts documents a 1 GB maximum per repository, while GitHub’s repository limits page recommends staying within 10 GB on disk, and the two pages do not say they measure the same thing. Cloudflare partitions the work so that there is no hot reference. GitHub keeps one codebase and aims to make the shared reference cheap. Most teams will need both: partition where work is independent, and make the shared reference cheap where it is not.

What the post leaves open

  • The 35 times figure. The phrase is “up to,” from “internal benchmarks,” and the post gives no workload, repository size, replica comparison, hardware or latency percentile. It does not say whether the comparison is one busy repository or the whole fleet.
  • When. GitHub says it is “already putting that foundation in place” and that the next post will “dive deeper into our future architecture,” but it names no date for any repository to move.
  • How agreement on the reference update is reached. The old design used a quorum; the post does not say what replaces it. Nor does it say how read workers that “cache data” learn about a new tip before the next CI job reads, although the post itself sets the requirement that a push be “made visible consistently before the next agent or CI job can build on it.”
  • Which Azure redundancy tier holds the copy. Microsoft’s storage redundancy documentation gives “at least 99.999999999% (11 9s)” durability of objects over a given year for locally redundant storage, “at least 99.9999999999% (12 9s)” for zone-redundant storage and “at least 99.99999999999999% (16 9s)” for geo-redundant storage. The post says only “durability and replication at Azure scale.” Between the lowest and highest of those tiers the annual loss probability differs by five orders of magnitude, and durability is a different property from the write latency GitHub is trying to cut.
  • Who gets it. The post describes GitHub’s platform and says nothing about GitHub Enterprise Server or other deployment types.

What to do with this today

  • Keep agent work off the shared reference. One branch or repository per agent, with a queue in front of trunk. In our runs, the mode with no shared reference never had a push rejected.
  • Do not let a fleet retry against one reference in lock step. Our runs spent a quadratic number of pushes on it. Backoff with jitter is standard practice, but we did not test it, so measure it on your own setup.
  • Size the merge queue to your CI budget. The Build concurrency setting is the documented throttle, and the documentation says it also affects how fast merges can complete.
  • Keep repositories lean. GitHub’s limits page recommends 10 GB on disk and warns that “Large repositories can slow down fetch operations and increase clone times for developers and CI.” Fan-out reads are the other half of the agent workload.
  • Watch for the follow-up post. The observable test of the 35 times claim is write latency and rejection rates for a busy repository under concurrent agents, before and after a move.

The announcement is a design statement, not a rollout. Its most useful claim is also its smallest: only the reference update needs agreement. Git’s push protocol already separates the two halves, and a few dozen lines of Python show what happens when a fleet of agents shares one branch.

Tags:

AI Coding AgentsDeveloper InfrastructureDistributed SystemsGitGitHub

Share

A lugworm lying on wet sand and mud at low tide
Previous Post

A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK

Two orange safety relief valves on grey pressure vessels in an industrial plant
Next Post

How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Blue-lit server racks in a modern data center, illustrating the compute infrastructure behind the AI boom.
Articles

The AI Boom Is Spending Real Money Before Proving Real Returns

June 7, 2026
Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A person with a laptop and smartphone, representing digital attention and AI-assisted work
Articles

AI Chatbots Are Making Attention a Design Problem

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026