Shepherd Newsletter #7: The String Quartet Problem
Welcome to this month's edition of what I've been thinking about in the technology space. This issue is about an old rule of computing, Amdahl's Law: why AI escaped it, and what it says about the parts of the world that are still slow. Shepherd runs a long-only global technology strategy, and we build conviction by working through problems like these ourselves.
The End of Amdahl's Tax
Most of the software that runs a bank or an airline is decades old, and nobody dares to touch it. Yet AI gets measurably better every few months. The gap puzzled me for a while, and I think the answer is a rule from 1967. Amdahl's Law says you can parallelize the friendly parts of a program all you want, but the serial parts set the ceiling. With that much legacy underneath, hardware and software could only evolve separately, and slowly.
Transformers broke the pattern almost by accident. The core of a large model is remarkably little code, mostly matrix multiplication and attention, with no legacy behind it. When the hot loop is that small, every layer of the stack can be redesigned around it at once: the algorithm, the compiler, the chip, the network.
And the gains stack. A better attention algorithm, a compiler that squeezes harder, lower precision, a chip laid out for exactly this computation: each multiplies the others instead of waiting on the slowest part. For the first time in decades, software and hardware can be designed together.
The loop is even starting to close on itself. Google used AI to design parts of its new TPU. OpenAI used its own models to write its chip's core kernels, beating expert-written versions by half, and went from first design to a finished blueprint in nine months.
What Remains Sequential
Amdahl's Law was never repealed. When one part of a system speeds up, the part that cannot sets the pace. So what is still sequential?
Inside the data center, it is inference itself. A model writes one token at a time, and each token has to wait for the one before it. Within each token, the data is tossed between pools of specialized experts on different chips, a design called Mixture of Experts (MoE), dozens of times. Every switch along the way is a traffic light, and the delays add up.
Outside the data center, it is people and the physical world. The economist William Baumol noticed in the 1960s that a string quartet needs the same four musicians and the same time to play a piece as it did in Beethoven's day, yet their wages rose with everyone else's. Work that cannot be sped up becomes relatively more expensive as everything around it gets cheaper.
I suspect AI replays this at a larger scale. The work a model can absorb, writing code, drawing, first-pass analysis, customer service, will see sharp deflation. What it cannot absorb will not: an approval, a regulatory review, a decision someone must be accountable for, a building or a power line that has to be built. Those steps become the most expensive part of the system, and they dictate its speed. The code for a new data center can be written in days. The permit and the grid connection still take years.
I hold this loosely, but it changes what I look for: less at how fast the parallel part is getting, more at who owns the steps that cannot be hurried.
What do you see as the most underweighted topic in tech? As always, I'd welcome your thoughts on any of this.
Until next month,
Jess Xu
