Key Takeaways
- "Small" used to measure risk, not effort. AI broke that link: builds got faster, but the number of assumptions riding on each one didn't shrink.
- Teams still give an hour-long AI build the same light validation a two-day story used to get, even though it can carry a dozen assumptions instead of one.
- Small should describe how many assumptions ride on a piece of work, rather than how long it took to build.
- A better name for it might be a single bet: one thing you're willing to be wrong about and check before you move to the next one.
I've written about end user validation, about why we need product-focused engineers, and about specs turning into a new kind of waterfall. All three pieces keep bumping into the same word, and I think the word itself is starting to lie to us. That word is small.
Small used to be a proxy for effort. A small story was one a team could build in a day or two, ship, and get in front of a user before too much time or money was on the line. If the assumption behind it was wrong, the cost of being wrong was small too, because so little had been built. Effort and risk moved together. That's why small worked as a unit of measure. It wasn't measuring the feature at all. It was measuring how much you were willing to lose if you were wrong.
AI broke that relationship. An agent can build in an hour what used to take a team two weeks, which means a team can now call something small because it only took an hour, while the actual feature it produced carries just as many decisions and assumptions as a two-week effort used to. The build got smaller. The bet did not.
That's the trap. A work item that takes an hour to build still gets treated with the same light validation a small story always got, a quick look at the diff, maybe a demo, out the door. But the thing being deployed might bundle a dozen assumptions about what the user needs instead of one. We kept the word small and the habits that came with it, while the thing the word was supposed to describe kept getting bigger.
I don't think small is the right word anymore, or at least not the right thing to measure by. If we're going to keep working in small increments, and I think we still should, small needs to describe the number of assumptions riding on a piece of work rather than how long it took to build. One assumption, tested with a real user before the next one gets built on top of it, is small in the way that actually matters. An hour of agent time that bakes in ten assumptions about what someone wants is not small, no matter how fast it shipped.
Maybe the word we need isn't small at all. Maybe it's something closer to single bet, one thing we're willing to be wrong about and check before we move to the next one. Whatever we call it, the size has to be measured in what we're guessing about the user rather than in how many minutes it took an agent to produce the code.