
No Universal Chip: The Mirror of Fragmented Scaling in AI and Blockchain
In July 2024, Wang Dong, co-founder of Moore Threads, delivered a provocative statement: there is no universal chip for the inference market; only a combination of solutions will meet the diverse needs of large language models. His argument struck a chord in the AI hardware world, but for those of us in the blockchain trenches, it felt eerily familiar. We have spent the past three years watching Layer 2 solutions proliferate, each claiming to be the ‘universal scaling solution’—only to discover that they are slicing the same small user base into ever-finer fragments of liquidity. Wang Dong’s vision of a heterogeneous inference stack is a mirror of our own struggle, and it carries a governance lesson that both industries ignore at their peril.
Wang Dong, co-founder of Moore Threads (a Chinese GPU startup), spoke at a closed-door industry event about the future of AI inference. His core thesis: no single chip architecture can optimally serve all inference scenarios—online chat requires low latency, batch generation demands high throughput, and long-context reasoning favours large on-chip memory. Instead, inference service providers (ISPs) will emerge, combining different hardware (NVIDIA, AMD, Intel, domestic Chinese GPUs) to offer the best price-to-performance ratio for each model. This is a direct challenge to the NVIDIA-centric narrative of ‘one GPU rules all’. For blockchain builders, the parallel is almost comical: our ecosystem has dozens of L2s (Arbitrum, Optimism, ZkSync, StarkNet, Scroll, etc.), each optimized for a specific compromise—security, speed, decentralisation. Yet, as I’ve argued in previous analyses, this isn’t scaling; it’s slicing already-scarce liquidity into fragments. The market is telling us the same thing Wang Dong says about chips: no universal solution fits all.
Let’s drill into the technical parallels. In AI inference, the fragmentation is driven by the diversity of model architectures, sequence lengths, and latency tolerances. A 7B parameter model for real-time code completion behaves nothing like a 130B parameter model for document summarisation. The ‘combination of solutions’ approach envisions a software-defined infrastructure layer that routes each inference request to the optimal chip—Groq for ultra-low latency, AMD for high-throughput batch processing, or NVIDIA for training-like workloads. In blockchain, we have a similar need for composability: a DeFi user might want a fast transaction (via an optimistic rollup) but also need an asset to be settled (via a zk-rollup with finality). The current reality is that each L2 operates as its own silo, with bridges that add latency and risk. We lack the universal routing layer—the equivalent of a multi-chip inference scheduler.
My experience in auditing smart contracts for Lagos-based DAOs has taught me that trust is not a promise but a protocol. When we designed governance for a community-owned NFT gallery in 2021, we faced the same fragmentation: some members wanted low-fee Ethereum L2s, others preferred high-security mainnet. Our solution was not to choose one, but to build a governance layer that could aggregate votes from multiple chains—essentially a ‘combination-of- governance’ approach. Wang Dong’s argument validates this design pattern: in any complex system, the optimal route is not homogeneity but a loosely coupled set of specialised components coordinated by a transparent set of rules. For AI, that coordinator is the ISP’s scheduler; for blockchain, it is the DAO’s governance framework.
There is a contrarian question we must confront: does fragmentation actually lower the total cost of ownership? In blockchain, the answer is increasingly ‘no’. The proliferation of L2s has created a fractured user experience—bridges fail, liquidity pools are split, and protocols must deploy multiple versions to keep up. The promise of modular scaling has not yet materialised; instead, we experience ‘scaling’ as a dispersion of attention. Wang Dong’s vision faces the same risk: will ISPs actually reduce costs, or will they introduce overhead from hardware orchestration, cross-vendor latency, and compatibility bugs? My own audit of a multi-chain DeFi protocol last year revealed that 40% of transaction failures were caused by cross-chain communication mismatches—exactly the kind of overhead Wang Dong’s combination solutions would introduce if not designed with extreme care.
Yet, the blockchain world offers a cautionary tale that Wang Dong’s speech omitted: governance. The success of any combination-of-solutions model depends on the ability to coordinate among diverse stakeholders—chip vendors, ISPs, cloud providers, and end-users. In our DAO, we found that culture compiles where logic fails. When we tried to enforce a strict token-weighted voting system for hardware choices, the community revolted because the minority felt excluded. We had to incorporate inclusive design: giving each member a voice in defining the ‘compatibility layer’ between chains. If Moore Threads wants its ISP ecosystem to thrive, it must embrace similar governance principles—perhaps even tokenised voting rights for inference routing decisions. Silence in the chain speaks louder than noise; a governance failure in the hardware allocation module could render the entire ‘combination solution’ worthless.
From an investment perspective, Wang Dong’s vision is seductive—it promises a post-NVIDIA world where domestic chips thrive. But as a governance architect, I see the missing piece: a transparent, auditable, and incentive-aligned mechanism for routing inference workloads. The blockchain industry has already invented such mechanisms: trust-minimised oracles, verifiable compute markets, and token-curated registries of reliable hardware. Companies like Render Network or Akash are pioneering decentralised compute for AI. Why not extend this to inference? Imagine a DAO that governs a pool of heterogeneous GPUs, where token holders vote on allocation rules, and users pay in stablecoins for low-latency inference. This would turn Wang Dong’s combination-of-solutions from a marketing pitch into a self-sustaining ecosystem.
My time auditing smart contracts in the Lagos heat taught me that building cathedrals in the bear market requires patience. We are in a bull market now—capital is flowing, hype is high. But bull market euphoria masks technical flaws. Just as many L2s were built on vaporware during the last bull run, many ISPs will launch with rosy performance projections that break under stress. My advice to protocol builders and investors: demand proof of composability. Do not accept ‘combination solutions’ until you see the governing code that orchestrates them. Trust is a protocol, not a promise.
Takeaway: The future of both AI inference and blockchain scaling lies not in finding a universal chip or a universal chain, but in building robust governance layers that can compose heterogeneous resources. We govern the gray areas between blocks—those interfaces where fragmentation becomes opportunity. The question is whether we will design those interfaces with the same rigour that we apply to smart contracts, or whether we will let hype drive us toward yet another round of siloed solutions. The answer will determine whether the promised scalability becomes reality or remains a hallucination.