Why would a public chain stop? From Cosmos' 25-hour halt, truly understanding blockchain consensus
- Core viewpoint: Cosmos Hub, due to the Neutron governance attack incident, had more than one-third of validators voluntarily stop producing blocks to prevent the flow of hacker funds, and recovered about a day later through a coordinated validator upgrade. The incident reveals that decentralization of a public chain does not equal never going down, and the consensus mechanism faces a trade-off between Safety and Liveness.
- Key elements:
- The attacker exploited Neutron's chain-level governance permissions to tamper with the admin of application contracts such as Astroport through the wasmd privileged instruction, and about 1.7 million ATOM were bridged into Cosmos Hub.
- Validators with more than one-third of voting power stopped running, Cosmos Hub lost Liveness, halted at block height 33,086,740, and users' ATOM transfers could not be confirmed.
- Recovery plan: Gaia v28.3.0 executed a one-time state modification at the recovery height, transferring 1,227,121 ATOM from the attacker's address to a 4-of-6 multisig address; after more than 67% of voting power installed the patch, block production resumed in about 6 minutes.
- Under BFT consensus, block commitment requires more than two-thirds of voting power. When this is insufficient, the network would rather pause to preserve Safety than sacrifice consistency and continue producing blocks.
- Historical comparisons: Bitcoin's client rule disagreement in 2013, Ethereum's The DAO hard fork in 2016, and Solana's out-of-memory halt in 2021 all reflect different failure paths involving consensus disagreement or loss of Liveness.
- Private key autonomy safeguards asset control, but the underlying consensus power determines whether the chain can process transactions and whether state can be collectively modified. The two are not the same thing.
On the evening of September 22, a user initiated an ATOM transfer.
Overnight, the transaction remained stuck in "pending confirmation."
The private key was not lost, and the wallet showed no signing anomalies. Upon checking again the next day, multiple public RPCs all showed that Cosmos Hub had stalled at block height 33,086,740.
With no new blocks being produced, there was naturally nowhere that could package this transaction.
Only about a day later, after Cosmos Hub resumed block production, did this ATOM transfer—which had been stuck in a pending state all along—finally succeed.

For ordinary users, this may be the most intuitive lesson in understanding blockchain consensus.
We are used to saying "no central authority can shut down a public chain," but reality is clearly far more complex. A sufficiently decentralized blockchain indeed typically has no "shutdown button" in a server room, but it can still stop.
This pause of Cosmos Hub happened to fully expose this mechanism—usually hidden beneath the surface—to ordinary users.
1. Why Did Cosmos Suddenly "Stop Producing Blocks"?
First, we need to clarify a potentially confusing point: what was directly attacked in this incident was not Cosmos Hub.
The event first occurred on Neutron.
On September 22, a Neutron governance proposal called "AIATO: AI Agent Takeover" was passed. The attacker exploited a loophole in chain-level governance authority, using privileged instructions natively provided by the wasmd framework to change the contract admin of applications such as Astroport and Drop to addresses controlled by the attacker.
This is not what we typically understand as a "code vulnerability" or "protocol flaw."
To put it simply, the applications themselves have their own "door locks," but Neutron's chain-level governance holds a higher-privilege "master key." When the attacker controlled the governance outcome, they effectively obtained this key, allowing them to reassign admins, migrate contracts, and further transfer the assets within.
What truly dragged Cosmos Hub into this was the cross-chain fund movement that followed.
According to Cosmos Labs' post-mortem, before Neutron halted operations, the attacker had already moved some assets to multiple networks, with approximately 1.7 million ATOM transferred into Cosmos Hub and beginning to be swapped through cross-chain liquidity.
In other words, Cosmos Hub itself was not directly attacked, and ordinary Hub users' funds were not directly stolen due to the Neutron vulnerability.
However, the ATOM obtained from the attack had already entered the Hub. To prevent the remaining ATOM from continuing to flow out, some Cosmos Hub validators began to halt their node operations.
By around 19:18 on September 22 (SGT), the validators that had stopped operating represented more than one-third of total voting power. As a result, Cosmos Hub could no longer form new blocks and ultimately stalled at 33,086,740.

This step is crucial.
It means that Cosmos Hub does not have a "Pause" button that some company can simply click, nor did it first go through an on-chain governance vote. What actually stopped the network was that enough validators stopped participating in consensus formation.
But what is more noteworthy is actually the recovery process that followed.
About 4 hours after the chain halt, validators received a complete recovery plan: execute a one-time state modification at the halt height, transferring the remaining ATOM in the attacker's address to a multisig address jointly managed by community validators.
Cosmos Labs then produced the Gaia v28.3.0 patch based on the plan validators had already agreed upon, tested it, and distributed it to validators.
This version of Gaia would execute a one-time state change at the designated recovery height, transferring 1,227,121 ATOM from the attacker's address to a 4-of-6 multisig address composed of Nansen, Keplr, Enigma, Silknodes, Kiln, and Polkachu.
By the early hours of September 23, validators confirmed to have installed v28.3.0 exceeded 67% of total voting power. Therefore, at 12:00 UTC that day, Cosmos Hub coordinated a restart. About 6 minutes later, this one-time state modification was executed at block height 33,086,741, and the network resumed normal block production.
Ultimately, throughout the entire process from Cosmos halting block production to resuming operations, it was validators who first caused the network to lose Liveness, and then more than two-thirds of voting power accepted a new set of state transition rules, ultimately making this rule set the canonical state after recovery.
At this point, a seemingly simple question emerges: Since it is a decentralized public chain, why can more than one-third of validator power stop it, while restoring the network requires enough validators to jointly accept and run the same software?
The answer is actually hidden in the word "consensus."
2. So-Called Consensus Was Never About "Never Stopping"
One of the most easily misunderstood things about blockchain is equating "decentralization" with "never going down."
In reality, what consensus mechanisms truly solve is how many nodes can agree on transaction ordering and ledger state without a central bookkeeper.
However, different public chains implement this in different ways.
Bitcoin's most classic approach is PoW, or proof of work—miners compete with computing power to produce blocks. When two valid branches briefly appear on the network, nodes choose one to continue building on based on cumulative work.
So Bitcoin does not have a clear moment where "after 67% voting, this block is forever finalized." It is closer to probabilistic finality: the more subsequent blocks there are, the more computing power cost is required to reorganize previous transactions.
This is also why people have long said that a Bitcoin transaction is best waited on for 6 block confirmations. After all, even with high computing power, one cannot simply bypass the consensus rules that nodes are executing.
Of course, this does not mean Bitcoin's state is "absolutely unmodifiable" under all circumstances. Theoretically, if the entire ecosystem accepts a new client and new consensus rules, through a Hard Fork, state changes that were invalid under the old rules could also be made valid.
But here lies the problem: who has the ability to get enough miners, Full Nodes, exchanges, wallets, and users to accept such a new rule set together?
Almost no one.
The development team cannot decide consensus rules on behalf of the entire Bitcoin network, and miners and exchanges find it difficult as well, because the consensus threshold that must be crossed is extremely high. When Binance was hacked for 7,000 BTC, some suggested that CZ contact major miners to intervene, but it ultimately came to nothing.
Ethereum provides another classic example.
After transitioning to PoS, Ethereum now uses the Gasper consensus, composed of Casper FFG and LMD-GHOST. Simply put, one part of the mechanism determines "which chain to follow now," while the other part is responsible for giving blocks true Finality.
When validators representing at least two-thirds of staked ETH agree on the relevant checkpoint, a block can move further toward finalization. Conversely, if more than one-third of the stake does not participate in correct voting for a long time, the network may temporarily fail to form Finality. However, Ethereum also designed an inactivity leak, which gradually reduces the effective weight of offline validators when finalization cannot occur for a long time, giving the network a chance to eventually restore Finality.
To truly change such an outcome, one would similarly need to change protocol rules and clients.
As in the 2016 The DAO incident, the Ethereum community ultimately carried out a Hard Fork, executing at block 1,920,000 a special state modification that the Ethereum Foundation at the time directly called an irregular state change, transferring the relevant ETH into a recovery contract.
However, some miners and community members who refused to upgrade and continued maintaining the original state eventually formed Ethereum Classic (ETC), leading to the well-known ETH and ETC split, showing that not everyone accepted this rule set.

Cosmos Hub is different again. It uses CometBFT, closer to a typical BFT consensus.
It can be understood as a more typical BFT consensus, meaning that for a block to truly commit, it needs a Commit from more than two-thirds of voting power.
Its advantage is that Finality is very clear. Once a block is committed after sufficient validator voting, there is no need to keep waiting for more and more blocks as in PoW, trading probability for security.
But the other side is also very direct: if one-third or more of voting power no longer provides the votes needed to form a Commit, the remaining validators cannot muster more than two-thirds no matter how hard they try.
At that point, the safest choice for the network is precisely what happened here: "pausing block production." So from the perspective of distributed systems, this brief halt of Cosmos Hub is actually not mysterious at all.
In a nutshell, after a group of validators holding sufficient voting power stop participating, the consensus protocol, according to its own rules, would rather lose availability than continue confirming new blocks without sufficient consensus.
Behind this are actually two concepts in distributed systems that ordinary users often conflate:
- Safety: Different nodes must not simultaneously confirm two conflicting final states;
- Liveness: Whether the network can still continue running forward and processing new transactions;
For BFT systems, when there are insufficient nodes participating in consensus, pausing is sometimes precisely the cost paid to maintain Safety. To put it bluntly, this decentralized ledger would rather stop there first than let the remaining participants each keep their own record.
Looking back from this angle, one finds that many seemingly completely different incidents in public chain history actually revolve around the same thing:
When distributed nodes can no longer form a consistent opinion on the "correct state," what should the network do?
3. From Bitcoin to Solana, Where Is the Real Risk Boundary of Public Chains?
This is not the first time Cosmos has brought this question to the table.
As early as 2013, Bitcoin experienced a very classic chain fork incident.
At that time, Bitcoin 0.8 switched its underlying database from Berkeley DB to LevelDB. Subsequently, a block containing a large number of transaction inputs appeared. New-version nodes could process it normally, but some old-version nodes judged this block invalid due to Berkeley DB lock count limitations.
Thus a very awkward situation emerged: everyone was running Bitcoin, but old and new clients began to produce different answers to the question of "whether this block is valid at all."
The network therefore split into two chains, and the new 0.8 side once had about 60% of hashrate, unable to rely on normal hashrate competition to quickly converge on its own.
In the end, large mining pools coordinated to switch back to the old version, regained more hashrate on the old-rule side, and only then did the network reconverge. Bitcoin later specifically reviewed this incident through BIP 50.
By 2016, Ethereum's The DAO incident pushed the problem one step further.
As mentioned above, in The DAO incident, the Ethereum community ultimately carried out a Hard Fork, executing at block 1,920,000 a special state modification explicitly called an irregular state change by the Ethereum Foundation, transferring the relevant ETH into a recovery contract.
But not everyone agreed with this handling. Some miners and community members who refused to accept the state modification continued to maintain the original rules, which led to the long-standing Ethereum Classic (ETC).
This DAO Fork was also a classic event, effectively telling everyone that when extreme events occur, there is social consensus beyond code consensus. If sufficiently consistent opinion cannot be formed, one chain really can split into two.
Solana in 2021 demonstrated another completely different failure path.
In September of that year, a large number of bot transactions flooded the network, causing validator node memory exhaustion and the crash of many nodes. Ultimately, the entire network could not form a consistent opinion on the current state and stopped confirming new blocks for about 17 hours. Validators then coordinated together to restore the network.

Putting these incidents together, one finds that they are not the same thing:
- The 2013 Bitcoin issue was that different clients began executing different validity rules;
- The 2021 Solana issue was that a large number of validator nodes could no longer normally participate in consensus, and the network lost Liveness;
- What Ethereum DAO faced was closer to the question of whether a community should actively modify state through new protocol rules;
- And this Cosmos Hub incident carries another layer of特殊性: the network first actively lost Liveness through validator coordination to stop the attacked assets from continuing to move; afterward, a sufficiently high proportion of validator power jointly accepted new software and the recovery state, allowing the network to form consensus again;
So, rather than simply reducing these events to "so blockchains can shut down after all" or "decentralization is all fake," it is better to acknowledge a more truthful fact:
A consensus mechanism is never a machine that cannot break. What it truly provides is actually a set of decentralized rules, such as who decides the correct chain when disagreement arises; how many participants are needed for a state to gain finality; whether the network chooses to keep running or stop when failure occurs; and under extreme circumstances, what kind of collective action can change the rules of operation going forward.
This also leaves this Cosmos incident with a question more worth ordinary users' consideration than "whether the chain should have been halted."
Final Thoughts
We often say, not your keys, not your coins.
This statement certainly still holds, but it emphasizes asset control—as long as the private key is in your own hands, wallets, exchanges, or other third parties cannot sign a transfer on your behalf.
The premise is that the blockchain you are on must be capable of processing that signature at any time.
On the day Cosmos Hub stopped producing blocks, users still held their own private keys, and assets did not disappear into thin air because of it. It was just that even if you correctly signed a transaction, there was no new block available to accept it.
The recovery process further showed that if enough consensus participants accept a new set of state rules, the on-chain state of specific accounts can also change without a signature from the original address's private key.
This does not invalidate "Not your keys, not your coins," but it reminds us that private key sovereignty and underlying consensus power are never the same thing.

The same is true for wallets.
Wallets can ensure that private keys and signing rights remain in users' own hands. They can identify chain-level anomalies as quickly as possible, accurately display transaction status, establish RPC and node redundancy, and reconfirm the final transaction result after the network recovers.
But a wallet cannot restore consensus on behalf of a public chain, nor can it guarantee that the underlying network will never be interrupted, let alone guarantee that on-chain rules and states will never change at the consensus level.
Therefore, what a mature decentralized system truly needs to pursue may never have been "nothing can ever be changed under any circumstances." On the contrary, it should clarify these imperfect boundaries as much as possible: Who can pause consensus? How much weight is needed? Under what circumstances is emergency intervention allowed?
Because true decentralization cannot make a system never encounter incidents. The key is that even if an incident does occur, we can still know who, according to what rules, and with how much consensus, decided how this ledger should be recorded next.


