复杂的代价:为什么系统不是越大越稳
From the Bronze Age collapse to modern infrastructure — when the returns on growth turn into fragility
A Concrete Puzzle: Why Solutions Grow Into Problems
Every modern institution was built to solve a specific problem. Regulation exists to prevent the last crisis from repeating. Approval workflows exist to prevent accidents. Shared internal platforms exist to stop every team from rebuilding the same thing. Multi-tier supply chains exist to cut cost. Decades later, the same institutions are the things being complained about: regulatory arbitrage, approval gridlock, platforms and business units dragging each other down, a supply chain where one break stops everything.
This is not random decay. It is a structure worth treating as an intellectual problem. It asks: why does the effort to solve problems eventually grow into part of the problem?
The cheapest explanations are "bureaucracies expand on their own" or "entrenched interests block reform." Sometimes both are true, but neither explains why the same curve appears where there is no bureaucracy and no entrenched interest — open-source projects, scientific communities, and fast-growing startups all travel the same road from nimble to ponderous.
The claim this essay defends is this: complexity is not a symptom of failure but a by-product of success. Any system with spare capacity will almost always choose to add another layer in order to solve the problem in front of it, and every layer borrows maintenance capacity from the future. The real risk is not that a system becomes complex. It is that complexity grows faster than the capacity to maintain it, year after year.
Complexity Is an Investment, and Its Returns Decline
To see this, one implicit assumption has to go: that complexity is free. It is not. Complexity requires a continuous supply of energy to sustain — taxation, management hours, training, coordination, audit, and a large volume of explanation cost that nobody can quite itemise.
The anthropologist Joseph Tainter, in The Collapse of Complex Societies, advanced an argument still cited today: when societies face problems, their dominant response is to increase sociopolitical complexity — more levels, finer divisions of labour, more specialised institutions, more elaborate resource allocation. This strategy works for a long time. The catch is that it obeys diminishing returns: the problem-solving capacity bought by each additional layer falls as layers accumulate.
Once the cost of further complexity exceeds what it returns, the system no longer has slack to absorb the next shock. Tainter's conclusion is counterintuitive: collapse is not a cataclysm but a rapid simplification — the system falls back to a cheaper level of complexity. It looks like catastrophe because the people inside it bear the entire cost, while the benefit — relief from the cost burden — is diffuse and delayed.
The least welcome corollary is this: for an over-complex structure, simplification may be rational at the level of the system. Maintaining the structure is itself consuming the thing the structure was built to protect.
Three Mechanisms That Make It Brittle
Complexity alone does not kill systems. Three specific structures do.
1. Tight Coupling: Local Failure Acquires a Global Path
Eric Cline's 1177 B.C. deals with the collapse of the Late Bronze Age. The point of his argument is not to identify a single culprit but to show that the palace economies had become densely interdependent through gift exchange, dynastic marriage, and long-distance trade in tin and copper. Drought, earthquakes, migration pressure, local war — the system had absorbed each of these individually. When they arrived together, interconnection itself fused them into a single collapse.
The mechanism deserves to be remembered on its own: interconnection raises efficiency and simultaneously merges previously independent failure modes into one shared failure mode. Financial systems, power grids, global supply chains, and a single company's technology stack all obey the same logic.
2. Hidden Dependency: Legibility Is Bought With Vision
Donella Meadows, in Thinking in Systems, stresses that a system's behaviour comes mainly from its structure, and that structure includes stocks, flows, feedback loops and delays. Delay is the most neglected of these: it separates cause from effect in time, so feedback fails and misjudgement accumulates.
Part of what complexity does is hide dependencies. An abstraction layer makes the local view clean, swappable, manageable — at the price of pushing real dependencies out of sight. When something breaks, the causal chain is no longer inside anyone's field of view. This is not negligence; it is what the structure produces. The more elegantly a system encapsulates its complexity, the less diagnosable it becomes when it fails.
3. The Maintenance Deficit: Benefits Arrive Now, Costs Arrive Later
Paul Kennedy's The Rise and Fall of the Great Powers argues that great powers drift into a long-run imbalance between their relative economic base and their strategic commitments: commitments are rigid, long-term and public, while the economic base that funds them shifts. The same structure applies to any system: the benefits of complexity are immediate, visible and demonstrable; maintenance costs are deferred, dispersed and unclaimed.
So every decision tilts toward adding another layer and cutting maintenance. This is not short-sightedness; it is what the incentives produce. By the time the maintenance deficit becomes visible, it is usually no longer a technical problem. It is a political one.
Complexity is a loan secured against future maintenance capacity, and the bill does not come due within the same term of office, the same budget, or the same person's career.
The Case Against: Scale Is Not the Disease; Missing Redundancy Is
The argument so far slides easily into a wrong position — anti-complexity, anti-scale, nostalgic for small communities. That position does not hold, for at least four reasons.
- Large systems can be extremely robust. The internet, a properly zoned power grid, and the biological immune system are all enormous and highly resilient. What they share is modularity, redundancy, diversity and loose coupling. Fragility comes from structure, not from size. Conflating the two mistakes the symptom for the disease.
- Some systems gain from disorder. Taleb, in Antifragile, separates three states: fragile, robust, and better off for volatility. For the third, redundancy is not waste but optionality, and small errors are not accidents but information. For such systems, reducing complexity means reducing the capacity to learn — the strongest counterexample to simplification as a policy.
- Tainter's model explains backwards and suffers survivorship bias. It is good at saying why a complex society that already collapsed did so; it struggles to give usable early warning in advance. And societies whose complexity rose without collapsing have no clear place in it. A theory with strong descriptive power and weak predictive power should not be used as a prescription.
- Simplification kills. Complexity in public health, food safety, aviation management and financial supervision corresponds to real deaths and real losses. The "efficiency gains" from dismantling it are often settled years later, in another currency. Anti-complexity romanticism is usually a luxury that only affluent societies can afford.
All four hold, so the thesis must be narrowed. The defensible proposition is not that complexity is harmful, but that fragility comes not from the absolute level of complexity but from the gap between complexity and the capacity to maintain it.
Limits and Tests: Three Kinds of Complexity
Once narrowed, the question stops being "more or less complexity" and becomes an operational one: which kind of layer is this?
- Load-bearing complexity: each layer removes a class of failure. Redundancy, isolation, validation, rate limiting, fallback paths. It raises cost and raises resilience at the same time.
- Ornamental complexity: it adds coordination, explanation and compliance cost without adding capability. Layered approvals, stacked metrics, alignment meetings that exist in order to align, abstractions added so that the abstraction can be understood.
- Disguised complexity: it looks load-bearing and is ornamental. It pushes cost downstream or into the future under the names of "risk control," "standardisation," or "governance."
Three tests separate them:
- The removal test: take this layer away — does the failure mode disappear, or does it merely relocate out of view?
- The modularity test: can any component be replaced without stopping the whole? If not, the modularity is nominal.
- The maintenance-ownership test: who carries this layer's maintenance cost, and do they hold the matching resources and authority? If the bearer has no voice, the layer will eventually be rediscovered as an outage.
Pass all three and the layer is worth adding. Fail any one and it is probably an invoice handed to the future.
The Judgement
Three levels follow.
For organisations: maintenance has to be a budget line, not the first thing cut when money is tight. A more honest metric than "how much technical debt do we carry" is this: if all new work stopped today, how long could the existing system keep running? The speed at which that number falls says more about an organisation's true position than any architecture diagram.
For individuals: complexity amplifies capability and dependency in the same measure. The criterion for choosing a system is not how advanced it looks but whether you understand how it fails. Not understanding the failure mode means handing control to someone else — and when it breaks, you cannot tell whether to fix it, route around it, or replace it.
For the age of AI: this is urgent. Models have driven the marginal cost of "add another layer of automation" sharply down, and that is precisely the danger. When an organisation can add another abstraction, another agent, another automated pipeline at almost no cost, it will. The crucial point: AI lowers the cost of building, not the cost of maintaining. If maintenance still has to be carried by people, then cheaper building only pushes complexity up to a level that maintenance capacity can never match. What appears is not a leap in efficiency but a population of systems that nobody fully understands and nobody is able to switch off.
The price of complexity is not complexity itself. It is that complexity is always secured against future maintenance capacity. So the question that matters was never "should we be more complex." It is: is maintenance capacity growing in step with complexity? As long as the answer is no, the system is already borrowing — however well it happens to be running today.