Shannph Wong

Shannph Wong

When execution stops being the constraint

The Judgment Reservoir

Sustaining Judgment When Execution No Longer Builds It

Execution built judgment for free. That subsidy is over.

You are preparing for your annual board meeting, and for the first time in three years, you are not sure what story to tell.

The numbers are impressive. The product teams launched three new products and cleared a feature backlog that had haunted quarterly reviews for years. Engineering shipped more in the last twelve months than in the previous two years combined. The go-to-market motion that used to take a quarter to stand up now takes five weeks. Support response times dropped by 40%. By every execution metric the board tracks, this was the best year the company has ever had.

A year ago, you told the board that the AI investments would drive a 30% gain in revenue growth. You believed it could be higher. The reason was simple: if we could ship twice as much, move twice as fast, cover twice as much market surface, the results would follow. The teams executed on that logic flawlessly.

The actual gain is 6%.

In a previous era, you would have celebrated a 6% acceleration in growth from a single operational shift. But that is not what you promised, or reorganized around, or told your investors to expect. The gap between how much was built and how little impact it produced points to something none of the numbers can explain.

What bothers you is not the number. It is what the year felt like.

The place should be buzzing. The pace of shipping alone should have created momentum. At the last all-hands, you could feel something was off with the energy in the room, and you still can’t name it. Team meetings are productive but mechanical. Roadmap reviews surface priorities with sensible rationale, and nobody in the room has conviction about any of it. You can’t remember the last time someone pounded the table fighting for an idea they believed in despite the pushback. The organization is making sensible choices. It has stopped making brave ones.

You have also lost some key people this year. Not many. The ones who mattered.

The VP of Engineering had been with the company for over a decade, through the early-stage chaos, through three hard pivots, through the scaling crises that nearly broke the organization. Every time, the company emerged stronger. She resigned in September. She was not burned out from the work. If anything, the opposite. She told you she had spent the better part of the year reviewing AI-assisted proposals, product specs, and architectural recommendations, and every one of them seemed good enough: structured, defensible, needing at most a tightened trade-off or two. Dozens landed on her calendar every week, each one adding to an accumulation of accountability over outcomes she could not really shape.

She decided to leave after a product review. The team presented three approaches to an integration architecture the enterprise strategy depended on, each developed heavily with AI assistance, each arriving with evaluation criteria articulated: integration complexity, migration effort, maintainability, security posture. They asked her which direction to take. She looked through all three and knew from experience that none of them had the extensibility to handle what enterprise customers would demand. When she started to explain, the room asked whether she had data connecting her concern to one of the criteria. She did not have a metric. She had twenty years of building systems that looked just like this and seeing where they broke. The room did not reject her judgment. They added it as a footnote, decided to use the existing criteria, and moved on without her.

She told you she had stopped feeling like a builder and started feeling like a clerk with a rubber stamp. She missed the days when the team brought her a hard problem and her job was to see something they could not. Now they brought her three options and asked her to vote.

Your best senior architect left two months later. During his exit conversation he said something you keep replaying: “I used to know why we were building what we were building. Not the business case, the real reason behind it. I don’t feel that anymore, and I don’t think anyone else does either.” At the time you dismissed it as nostalgia.

Neither of them brought you a number. The gains had numbers. Cycle time, release frequency, time to market, all of it in the deck you are staring at now. They told you that people were accepting answers too quickly. That your most experienced people had become reviewers of reasoning rather than authors of it. That engineers could explain what the system was supposed to do but could not explain why it behaved that way. None of that has a line in the operating review. You had instrumentation for one side of the ledger and anecdotes for the other, and you weighted them accordingly.

Their departures left holes that are much larger than their role descriptions. When your VP of Engineering was in the room, she could look at a product approach and tell you whether it would hold together at scale, because she had built and rebuilt enough systems to feel structural weakness before she could articulate it. That instinct is not in any documentation. It walked out the door with her. When the junior engineers had their code reviewed by the architect, they learned what mattered, what was worth fighting over and what was noise. Two of those engineers have told you, separately, that they are shipping faster than ever and learning less than they have in years.

The conclusion you came to that stopped you from sending the board deck is that another year of the same approach will not compound beyond the 6%. You can feel the plateau coming. The first six months of AI adoption produced real acceleration. The next six produced refinements. The gains are flattening, and the operational advantage that felt decisive in January is now table stakes. You have started wondering whether every competitor is arriving at the same conclusion, and which answer would be worse.

You did everything right. You moved early. You built real capability, not a headline. And here you are with a board deck draft, aware that the results do not match the effort, that the trajectory is not compounding, and that you cannot explain the gap by pointing to any single decision that was wrong. The people who might have told you why, the ones whose judgment the system had no way to use, are the ones who are no longer in the building.


This pattern is not specific to any one company. It is forming across organizations at all different stages of AI adoption, and the ones that have not felt it yet are on the same trajectory. The question is not whether it arrives, but when.

When output metrics show real gains and business impact does not match, the reflex diagnosis lands on execution: the tools need tuning, the workflows need refinement, the teams need better prompts. That diagnosis is comforting because the remedy is more of the same, done better. It is also insufficient.

The deeper problem is not AI adoption. It is that AI displaces the critical work nobody accounted for without replacing what that work produced.

Under acceleration, execution scales faster than coherence, and coherence, unlike alignment, does not hold on its own. Teams perfectly aligned to their goals can still produce a fragmented result. Maintaining coherence requires deliberately designed decision infrastructure1: clear decision ownership, feedback loops that detect drift, incentives that reward coherence rather than raw output. Even the best decision infrastructure is only as good as the judgment it runs on.

That infrastructure depends on people who can tell the difference between drift to be corrected and noise to be ignored, who know when a correction is urgent and when it is merely reactive, who can feel that an integration architecture will not hold under enterprise complexity before the data confirms it.

Demand for that judgment is higher than it has ever been. Every additional release, every concurrent initiative, every expanded surface requires someone to evaluate whether what the organization is collectively building still holds together, and AI has multiplied all three. Supply is draining. People leave. People burn out. People disengage when the system stops demanding what they are best at. Each departure takes with it judgment that no document or decision framework replaces. New people arrive, but organizational judgment cannot be hired in finished form. It is built through years of operating inside the system.

Judgment behaves like a reservoir, and not every reservoir needs to stay full. The instinct for hand-tuning a query plan or knowing which server to restart at 2am was once worth having but is no longer worth maintaining. Other judgment should migrate to where it is needed rather than persist, moving from running the process to designing the system that runs it. Depletion is not the failure. Unmanaged depletion is: draining capacity before knowing which parts of the system still depend on it.

Who in the organization is still able to tell which reservoirs of judgment the system depends on? Making that assessment requires people who have carried that load, and they are the same people the depletion is losing. An organization that has already drained that capacity will run the audit and conclude that everything it sees is everything it needs.

The New Evaluation Bottleneck

Nearly every AI investment builds on the same assumption: more information leads to better decisions. That is only true up to a point. After that, the bottleneck moves, and most organizations do not notice because the new bottleneck feels nothing like the old one.

When information was scarce, the bottleneck was producing good options. Now that AI generates plausible options in seconds, the bottleneck moves to evaluating them. Anyone who has stood in a paint store staring at a hundred shades of white understands the feeling. The differences are real and narrow, and you can either become an expert in undertones or ask the clerk which one is most popular. The interior designer picks the right one in thirty seconds, because she knows your south-facing windows will wash out the cool tones, your maple floors pull warm, and two children under three means the finish has to survive grape juice. Her judgment is not about the swatches. It is about everything the swatches do not contain.

Now ask that designer to make that call for twelve houses a day. By the eighth she has stopped asking about the muddy sheepdog who greets everyone at the front door, and by the twelfth she is picking the one that looks close enough. The judgment is intact. The bandwidth to apply it is gone.

AI-generated options appear comprehensive because they arrive pre-evaluated against criteria. The integration architecture comes with a technical trade-off matrix. The product strategy comes with projected customer impact. The hiring plan comes with modeled scenarios. Every option is framed so it can be compared: cost, timeline, projected return. The decision process gravitates towards those generated criteria simply because they are there, and measurability stands in for comprehensiveness. The system never asks what good judgment would have asked: are these even the right criteria?

What gets lost is not the analysis. It is the capacity to question what the analysis operates inside. A scoring rubric is limited to the dimensions someone thought to include, and when the rubric is AI-generated, not even that. The product instinct that says the impact projections are modeled on the wrong customer segment is not a dimension the rubric includes, and you can only choose from what is presented. Organizations have collapsed defining the rubric with evaluating against it into a single step, and that collapse is frictionless. Nobody has to choose metrics over experience. The numbers simply arrive first.

Over time, the role of the experienced person narrows from shaping decisions to approving them. That shift feels like efficiency. The first to feel the loss are the people who carry the deepest understanding.

Faster Iteration Without Direction

A few years ago I ordered paella at a Spanish restaurant in Shanghai. When it arrived, I was impressed. Everything looked right: the rice, the seafood, the saffron color, the carbon steel pan. Every visual cue said paella. Then I tasted it. It tasted like fried rice. The chef had seen what paella was supposed to look like and had replicated the appearance with real skill. The chef had done everything short of tasting an authentic paella.

The gap between what looks like paella and the real thing is the gap between replication and judgment. The chef recreated the dish from its appearance. With enough reference images and enough technique, the replica looked convincing. The taste could not be reverse-engineered, because it was never in the reference. It lives in the experience of eating and preparing the real thing.

Organizations have always relied on the coherence of an argument as shorthand for the quality of the thinking behind it. If the reasoning is well-structured and internally consistent, the instinct is to assume someone did the hard work. Now AI produces output that passes that filter without the thinking behind it, and the gap is invisible for the same reason the paella looked authentic: the reference never contained what was missing.

In practice, AI gives the right answer across most of the routine decisions an organization makes, and that is ironically what makes the problem worse. It is natural to stop double-checking because usually there are no apparent consequences. Review cycles shorten, pushback softens, and the habit of scrutiny atrophies. When a problem does arrive where the old patterns do not apply and the ambiguity is genuine, the atrophied muscles cannot engage.

Imagine that Spanish restaurant now gets a rice cooker that makes the paella fast and cheap to produce. The restaurant can now run variants at will: seafood, mixed, spicy. It is tempting to believe that with enough iterations one of them converges on the real thing. It will not. The rice cooker made production faster, and the chef still does not know what paella should taste like.

A version of this belief operates at the organizational level: because AI made experimentation cheap, organizations can iterate their way to the right answer. When the direction is clear, cheap experimentation genuinely accelerates. When it is not, it only generates more options that were never worth pursuing. The speed is real. The convergence is an illusion.

Judgment is the capacity to look at a set of options that all sound right and see the one that truly is, or to recognize that none of them are, and to have the conviction to act on it before having the data to support it. That capacity cannot be bought, installed, or prompted into existence. It has to be developed. Which raises the harder question: where does it come from?

Where Judgment Comes From

We tend to romanticize judgment, describing it as a trait some people are born with. That framing is convenient because it turns judgment into a hiring problem, a problem organizations already know how to solve. But when a chess grandmaster glances at a board mid-game and sees the winning line, what looks like instinct is the residue of thousands of games of deliberate analysis, repeated under enough pressure to become automatic recognition. It was built, not born. The grandmaster is not doing a different kind of thinking. It is thinking that has been done so many times it stopped requiring effort.

Human cognition operates on two tracks2. The first is fast and automatic, always running, the one that lets you drive home while listening to a podcast without thinking about where to turn. The second is slow, deliberate, and taxing, the one that engages when you see the flashing lights of a firetruck ahead and decide to take an early exit. The boundary between them moves with training. Judgment is what happens when the deliberate work, done enough times under real conditions, compresses into automatic recognition. Losing people who have judgment is a talent problem. Losing the ability to develop it is a structural one.

Consider what judgment in action looks like. A CPO who has led product expansions into adjacent markets builds a mental model from the ones that worked and the ones that stalled. When the team presents a plan to move upmarket into enterprise, the CPO can feel that the plan underestimates the complexity, not because they analyzed this specific market but because they lived through six similar expansions that floundered on the same invisible dependencies: longer sales cycles that starve the core business of attention, support demands that do not scale linearly, a buyer whose procurement process reshapes the roadmap in ways the plan does not account for. It arrives now as pattern recognition, a conviction they would struggle to justify on a spreadsheet and yet would stake their credibility on.

What separates that CPO from someone with the same title and similar tenure is what was done with the experience rather than the experience itself. Twenty years of doing the same thing is not twenty years of judgment development. It is one year repeated twenty times. Judgment forms only when the cycle closes: the outcome is examined, the world model is updated, and the updated model shapes the next decision. The rate of judgment development is governed by how many cycles close and how much was riding on each one. What makes the cycle work is not the size of the damage. It is that the person committed to a reading of the situation before they knew whether it was right, and then found out.

That cycle does not close in the model. It closes in whoever will answer for the next decision. It can generate options, evaluate them against known criteria, and present them with a coherent rationale. The weight of an outcome has to land on whoever is accountable for the next decision, not the model. Even a model that learned perfectly from every outcome would be improving only itself, not the people the organization depends on when facing a situation the model has never seen. And even with rapid model improvement, the trajectory does not reach it, because the reasoning was never in the source material. The document only presents the outcome, not the work that produced it: the three weeks of debate, the hallway conversation that resolved the disagreement, the feel for what was signal and what was noise. The same document could have been produced by entirely different reasoning under different conditions, which is why a human trying to reverse-engineer the judgment behind it would hit the same wall.

The most significant impact of this kind of AI use is on the environment where experienced people apply their judgment and less experienced people build theirs. The proposal review that used to take two hours now takes thirty minutes, because the proposals arrive with the standard concerns pre-empted. The architecture review becomes a checklist discussed async. Quarterly planning runs smoothly because every initiative comes with defensible rationale. Each of these is useful. But when the output consistently looks competent, the reviews get shorter, the pushback gets softer, and eventually the meetings themselves start to feel unnecessary, because if the feedback is never substantive, why convene to deliver it? Deliberate thinking requires a trigger to engage, and the trigger has been filtered out before anyone sees it. The process that would have developed and maintained judgment is hollowed out from the inside, and it feels like efficiency the entire time.

An individual can decide to slow down and think more deliberately. But an organization that has restructured its processes around speed cannot make that choice the same way. The blind spots are not in the people. They are built into the processes, the cadences, the incentives. The proposal that arrived half-formed and forced someone to ask why, the review that ran long because nobody could agree, the objection that could not be answered from the deck: those were the signals that deeper thinking was needed, and they are the ones being removed. And once the process has been reduced to approving what comes in the inbox, it is not only the current generation’s judgment that goes unexercised. It is the next generation’s judgment that never forms. The reviews where people still building their judgment learned by studying how experienced people think are the same reviews being shortened, moved async, or eliminated. The organization has not only stopped using its judgment. It has dismantled the conditions that produce it.

Who Tells You the Bet is Wrong?

The loss of judgment capacity does not only make individual decisions worse. It removes something more fundamental: the ability to discover when the organization’s own direction is wrong.

Companies that recognize the need for organizational coherence under acceleration are building decision infrastructure to maintain it. That infrastructure is essential, and it solves a specific problem: detecting when decisions made inside each team’s scope stop pointing in the same direction as the company’s strategic bet, and correcting them when they do. It works as long as the bet is right, and no bet stays right, because the conditions it was placed against are always shifting. Nothing in that infrastructure detects when the bet itself has drifted from reality. The better it functions, the more precisely the organization executes a direction nobody is checking. The capacity to challenge the bet has to carry real authority, because one step in the wrong direction requires two to correct, steps that could have been spent on the right bet.

Earlier in my career I joined a GPS navigation company as a senior architect. I was employee fifty, and we grew past a thousand people in under a decade. My credibility came from scaling the architecture that got us there. Under the pressure of user growth doubling every six months and the demands of carriers who represented 80% of revenue, that architecture had become a monolith, forked repeatedly to serve each mobile carrier’s requirements. When Google announced free turn-by-turn navigation on Android, they gave away the thing carriers were paying us millions a year for. The core business had a clock on it. We needed to diversify into adjacent markets, and the architecture I had spent years building made that nearly impossible.

I proposed rebuilding the entire product on a new API platform designed to serve every product line from a single foundation. The bet was enormous. If it failed, it would take the core business down with it. I asked for seven hand-picked engineers and three months to build a proof of concept, and I led the team personally. We hit every milestone. Then the skeptics shifted their objection: a proof of concept was not a production system. I asked for three more months. We shipped the full product with all existing features and the next version’s roadmap items, cutting app load times in half and improving key response times by thirty percent.

The replatform came with a cost I did not anticipate. The new architecture eliminated the need for parallel teams maintaining forked codebases, and team leads saw their scope shrink and their advancement path close. About a dozen engineers left before I had the chance to talk with them about the opportunities this opened up. Years later that same platform enabled the two largest deals in the company’s history: embedded connected navigation across GM’s and Toyota’s entire automotive fleets. The bet was validated only in retrospect, and only because the organization absorbed the cost of making it.

Consequence develops judgment, including the consequences nobody anticipated. I had built the configuration layer that made the forks possible, and I had spent years cherry-picking bug fixes across them. I knew what each new fork cost before anyone asked and what a clean abstraction would buy. The product lines had to stop being a collection of apps and become surfaces over a common platform. I had modeled the architecture and the market. I had not modeled what eliminating the forked codebases would do to the people whose scope depended on them.

Judgment enables conviction. Once that mental model is built through consequential experience, a person can trust it even when it conflicts with the available data, and even when there is no metric to defend it in the room. The skeptics could price precisely what a failed replatform would cost the business that paid for everything. Nobody could price the deals that would not happen three years out. My conviction was not confidence that I was right. It was that I was the one person who knew what continuing down the same path would cost, because I had been paying it. My credibility was the architecture I was proposing to tear down.

The improbable choice is where new understanding comes from. The probable option, by definition, reflects the current consensus about how the world works. In this case the consensus had a team behind it: the automotive product was already being built as another fork, by engineers replicating the pattern the last decade taught them. That was the safe path, and it would have shipped. It would also have produced yet another codebase and no platform. The deals that came three years later that depended on mobile and the car running on the same foundation would not have been possible. Consensus is built on assumptions that were true when it formed, and nothing inside it registers when they stop being true. The only way to find out is to choose what it would not. Without that capacity, the organization optimizes within its existing framing rather than discovering when the framing needs to change. And when every organization draws from models trained on largely the same data, the framing is not even its own.

This is what is at stake when judgment erodes. The organization will remain coherent. It will remain efficient. And if the bet is wrong, it will converge on that wrong answer faster than any competitor, with every correction mechanism working exactly as designed to keep it on that course. The person who can tell you the bet is wrong is someone who has paid for a decision and lived long enough with the result to know what the next one will cost. That is what the organization is no longer producing.

Watching is Not Doing

When I was a senior architect, I gave an onboarding task to a recent grad engineer: build the tracer bullet, a monitoring call that would run continuously against our most critical APIs, so we would know if they were unhealthy before a customer did. Most organizations would not hand API design to a junior. That work tends to be reserved for senior engineers. I have never operated that way. Ownership should run end to end at every level, and what changes with seniority is how much ambiguity you are handed, not how much of the problem you own.

He spent a day working on it and walked me through his design. He had understood the hard part: the tracer bullet had to be identifiable as a test so it would not pollute the logs or the metrics. His solution was to add a boolean parameter on the API signature to mark the call as a tracer.

It was a reasonable answer that would work, and I spent the next hour with him on it anyway. I asked what other ways that information could be sent. We ended up moving it into the caller context instead, out of the API entirely.

He accepted the change, and then I saw the look on his face that said “was a boolean really worth an hour of discussion?” So we spent another hour answering that question. What it costs to change an API after a dozen teams have integrated against it. Why a parameter that carries context rather than intent is the wrong kind of parameter. That changing an API is always expensive, no matter how small the change, because of how it compounds.

Two things happened in that room and only one of them looked like teaching. The consequence was his. He made a real design decision on infrastructure other teams would depend on, and he had to defend it. The hour that followed is what an observer would have seen and it only worked because of what preceded it. He was not watching someone else reason about API durability. He was hearing why his own choice was the wrong one, which is why it stayed with him.

Take away the consequence and the same two hours produce something thinner. He sits in a review, hears the identical reasoning about somebody else’s flag, and leaves knowing that APIs are expensive to change. What he does not leave with is the feel for which small decisions are the durable ones, because he never made one and never had to answer for it. He can recognize the shape of good judgment without having developed the taste for it. The paella that looks right.

Take away the observation instead and he ships the boolean. Eighteen months later four teams change their integrations, and he learns at full price a lesson the organization already knew. That is judgment development, but it is the slow and expensive kind, reinventing what somebody in the building could have told him.

This is what apprenticeship was. Not the hour of explanation, which any organization can schedule, but the hour of explanation on top of a decision he had already made and could not take back. Remove either part and the other stops producing judgment.

Here is what should be uncomfortable about that scene. I had no management responsibility for that engineer and was assigned no goal that his API design served. I was the senior architect for the entire technology platform, in a year when the architecture was forking faster than we could maintain it. The two hours came out of nothing that anyone measured. If I had skipped them, no system in that company would have registered it, and he would have shipped a working tracer bullet either way.

That is what inherited judgment development looks like. It worked, and it worked exactly once. It required a specific person to be in the room, to recognize what mattered, and to spend an hour they did not have on an engineer they were not responsible for. Nothing about that arrangement can be relied on twice, let alone across a hundred engineers and a decade. An organization that develops judgment this way is not running a system. It is hoping the right people keep showing up.

Designing it means building the conditions into the work itself, so the consequence and the observation arrive together whether or not the right person is available that week.

The clearest example I have is an Amazon mechanism called the Correction of Error (COE). When something customer-impacting breaks, the accountable leader is obligated to write a COE: a structured analysis covering not just what happened and why, but how to detect the problem earlier, respond faster, resolve it with less customer impact, and prevent the entire class of error from recurring. The standard is exacting. Senior leaders read every word.

I wrote one after a failure in the Amazon catalog pipeline, the systems through which billions of daily catalog updates are applied across nearly every Amazon property: marketplaces, Prime Video, Whole Foods. The pipeline depended on AWS Kinesis, so when a major Kinesis outage hit US-EAST-1, the pipeline went down with it. All marketplace regions were blocked simultaneously, nearly 200,000 merchants could not update listings, and Whole Foods suspended orders because inventory updates could not process. Thirty-one and a half hours from impact to full recovery.

Handling the outage demanded judgment under pressure. The COE that followed is where the deeper judgment formed. The structure does not let you stop at “the dependency failed.” It demands you trace why that single dependency was a point of failure across all regions at once, why failover was not possible, why the mechanisms that existed had never been validated at that scale. Each level strips away another layer of the system’s implicit assumptions. By the fifth, you are no longer analyzing an incident. You are understanding the architecture of a system you thought you already knew. Thirty corrective actions followed, each with a named owner and a due date. You can only prescribe at that granularity because the process forced you to understand at that granularity.

Amazon invested heavily in making COEs visible in perpetuity. They are shared, searchable, referenced by senior leaders. In practice people read them retroactively, when they hit a similar problem and want to see how it was handled. That is useful. Writing one produces something else. A friend who recently joined Amazon told me they had read mine. They could describe what went wrong. Six years later, I can still feel the pattern of cascading failure that would tell me a new system carries the same structural weakness, before the metrics confirm it.

The COE is proof that judgment development can be built into the work rather than left to whoever happens to care. It develops one kind of judgment, operational, and product, strategic, and market judgment each need their own mechanism built from their own consequential experience. The COE is not what transfers. What is needed is the mechanism that forces the analysis on the person accountable for the outcome, and did it every time rather than when someone had an hour to spare.

It is also proof that the development happens through the doing, not the reading. Organizations that make judgment visible are doing necessary work. Organizations that assume visibility alone transfers judgment are confusing the output with the journey that produced it.

Delegation Without Follow-Through

Andy Grove, who ran Intel through its hardest decade, put it plainly: delegation without follow-through is abdication. The word he chose was not understanding. It was follow-through3: the monitoring the delegator retains after the work has been handed off.

That monitoring does two jobs. The obvious one is that it catches errors before they compound. The invisible job is to keep the delegator’s own model of the work current, which is how a leader stays calibrated on a domain they no longer execute in themselves. Delegate to a person and both loops run. The person doing the work learns from doing it. The person monitoring learns from watching how it goes and comparing that against what they expected.

Delegate the same work to an agent and neither loop closes. The agent has no consequence to absorb and no world model to update. And the reviewer, facing an overwhelming queue of AI-generated output that has grown faster than the hours available to them, has downgraded monitoring into approval. Approval teaches nothing. The work gets done, the output is fine, and no judgment is formed anywhere in the system.

Consider how this happens during a code review, which is one of the highest-validity feedback environments an engineering organization has.4 The cues are identifiable, the feedback arrives fast, and the judgment it builds transfers upward into architecture and design. An agent generates twelve pull requests before lunch. They are competent. The first three get read closely. By the seventh, the reviewer is checking whether the tests pass and whether anything looks obviously wrong, because twelve PRs is what the day looks like now and merge rate is what their quarterly review is measured on. The reviewer is behaving rationally. Nothing in the system is asking for anything else. Eleven of the twelve did not need close reading. The scrutiny that would have caught the twelfth has been spent, and the twelfth merges with just a glance. The returns from additional volume do not merely diminish. They reverse, because the finite resource that separates the consequential decision from the routine one has been consumed by decisions that did not require it.

What used to happen in that review does not happen anymore. The junior engineer who would have written the code and learned where the system breaks did not write it. The senior engineer who would have learned something by explaining why it breaks did not explain it. Both ends of the apprenticeship closed in the same motion, and the merge rate improved.

The pattern is not specific to engineering. A product manager who builds a working prototype in an afternoon has produced something a written spec could not, and has skipped the part where writing the spec forced them to name the edge cases the prototype does not surface. The prototype works. That is what makes the gap invisible.

Nobody in this arrangement is doing their job badly. The reviewer is clearing the queue. The product manager is moving faster than they could last year. It is the same year repeated, and neither loop closed, which is the only way judgment has ever formed. What is different now is that the failure is no longer a matter of individual reflection. The organization is handing people work that leaves nothing to reflect on, and then measuring them on how much of it they close.

The two failures are not independent. Automation reduces what enters at the bottom, because the work that formed judgment has been absorbed. Review gates increase what is drawn from the top, because every unit of generated output requires an experienced person to sign off on it. Less comes in, more goes out, and the mechanism producing the shortfall is the mechanism installed to manage it.

The Reservoir

Judgment is consumed every time a leader makes a hard call, every time a senior practitioner is spread across one more review, every time the organization expands its surface area faster than its people can keep up. When the judgment reservoir5 is drawn down without anyone noticing, the trajectory is already set. The current level tells you nothing about the direction. It can look full and already be draining from underneath.

The good news is that judgment is renewable. It is replenished through decisions people commit to and then live with the results of. For decades the balance between consumption and replenishment held without anyone designing it. Execution was slow enough that demands on senior judgment were constrained by how much the organization could ship, and juniors developed judgment through the execution work itself: writing the code, owning the customer call, sitting in the room when the trade-off was made. The judgment development pipeline ran as a byproduct of normal operations. Nobody had to think about it because nobody had to design for it.

AI is not the only pressure on that balance. Hiring freezes thinned the junior cohort, remote work removed the ambient proximity, and flatter structures widened the span between the people making decisions and the people learning from them. What separates AI from the others is that it removes the work itself.

Organizations are hiring fewer juniors because AI handles much of the output those roles used to produce. The juniors who are present are doing different work: learning to use AI effectively rather than developing deep familiarity with the systems, customers, and markets those tools operate on. When the product review surfaces a strategic tension, the junior PM who has spent a year generating specs with AI assistance has a different relationship to the customer than one who spent that year in the field learning why customers actually churn. The pipeline is not only shrinking. What flows through it has changed, and the judgment it produces is thinner.

The instinctive response when the pressure builds is to hire for it. Bring in senior people who already carry the judgment. It is the obvious move, and it pulls exactly the wrong lever. Every senior hire draws on the judgment of people who are already overextended, because onboarding at that level requires context transfer only existing seniors can provide. It does nothing for the pipeline beneath them. If the developmental mechanisms have been removed, even senior people who want to invest in the next generation have nothing to invest through. The pattern reinforces itself: depletion creates pressure, pressure triggers hiring, hiring increases the burden on people already carrying too much, and the attrition that follows deepens the deficit the hiring was meant to address. The organizations that recognize this early will design for it. The ones that do not will keep pulling the same lever until the reservoir is too depleted for any short-term fix to resolve.

Assume the hiring works anyway. The senior people land, they onboard without draining anyone, the pipeline beneath them gets rebuilt. There is still a loss that none of it addresses.

Organizations have always run on error correction across functions. Product read an opportunity one way, engineering read the feasibility another, finance read the risk a third, and the disagreements caught what no single function could see on its own. That correction worked for a structural reason: the errors were uncorrelated. Each function built its model from a different substrate, from different customers, different failure modes, different time horizons. Each was right and wrong in a different direction.

When the judgment that remains in each function is scaffolded by the same models, trained on the same corpus, surfacing the same considerations in the same order, the readings stop diverging. Product and engineering arrive at the review having consulted the same underlying distribution, and they find that they agree. The agreement looks like alignment and gets treated as confirmation.

Headcount holds. Seniority holds. The org chart still shows three functions with three perspectives. What has gone is the property that made three perspectives worth having: differentiation. That independent perspective goes before the people do.

When Everyone Reasons From the Same Place

Judgment formation is not only an organizational problem. It is a competitive one. Every major competitor in a given market now has access to the same AI capabilities. The frontier models draw on a substantial amount of overlapping material. Integration patterns are standardizing, tooling is commoditizing, and the timing advantage of early adoption is expiring.

Consider two AI-native competitors who have moved past productivity use and now rely on AI to evaluate market opportunities, frame product direction, and build the rationale for where to invest. When the model generates three strategic options and the leadership team selects the most compelling one, that choice feels like a judgment call. But when the three options were drawn from the same training data every competitor’s models draw from, and the evaluation criteria are the ones most prevalent in that corpus, the judgment has been pre-filtered to the most plausible. The compelling option is the one the model can argue for most fluently. The same filtering runs inside the organization, one decision at a time.

The antidote is supporting and amplifying the people whose judgment was built through consequence rather than through the same corpus the model draws from. People who can feel when the frame is wrong because their world model was constructed differently. The experienced leader who senses the framing is off is constructive divergence. The junior who asks why the established framework works the way it does is constructive divergence. When an organization removes the conditions under which those people develop, thrive, and stay, it removes its own capacity for the pattern-breaking that innovation requires. The improbable choice, the one the model would not have selected, is where competitive differentiation begins.

The gap that opens cannot be closed by deploying the model a competitor deployed last quarter. It is not a tooling gap. It is a judgment gap, and judgment gaps are invisible until the moment arrives that demands the answer only judgment can provide.

The Autopilot Precedent

AI is not the first technology to displace the work that built judgment without replacing it. Commercial aviation ran the same sequence across four decades, with the stakes measured in lives.

Autopilot reduced cockpit crews from five to two and changed the profession. Alongside advances in jet engines, deregulation, and hub-and-spoke network design, it enabled a massive expansion of the industry. More routes became economically viable. More aircraft flew more frequently. More pilots were needed, not fewer, because the scale of operations grew faster than the crews shrank. Air traffic control developed into a coordination layer keeping distributed, high-velocity operations pointed in the same direction. Entirely new industries emerged around the expanded system. The industry did not shrink because of automation. It expanded through it.

The transformative moment came with the Airbus A320 in 1988, the first commercial aircraft with full digital fly-by-wire. Automation handled not only cruise but the complex transitions, the approaches, the instrument procedures. The pilot’s role shifted from active operator to system monitor. Manual flying became the exception.

By the mid-1990s, the consequences were visible to anyone looking. In 1996, an FAA human factors team published a study of flight crew interfaces with highly automated cockpits. It identified the problem directly: automation dependency was creating vulnerabilities in pilots’ ability to manage flight path manually when the automation failed or behaved unexpectedly. It raised an explicit concern that investment in human expertise was being reduced under economic pressure, precisely when nearly three-quarters of incidents involved flight crew errors. The problem was documented. The industry noted it.

Thirteen years later, Air France Flight 447 crashed into the Atlantic, killing all 228 aboard. The airspeed sensors iced over at high altitude, the autopilot disconnected, and the pilots had to fly manually. They could not. The training regime had not prepared them for manual flight at altitude under those conditions, and the instincts that would have told them what was happening had atrophied through years of automation-managed flight. The aircraft was functional. The skills to fly it were not. Air France’s own internal assessments had already flagged that airmanship among some long-haul pilots was weakening, and an FAA review afterward found the same pattern industry-wide: in over 60 percent of accidents, pilots struggled to fly manually or mishandled the automated controls. The co-chair of an FAA advisory committee on pilot training summarized it in five words: “We’re forgetting how to fly.” The warning existed before the crash. The consequences arrived anyway.

The response was not to remove autopilot. Nobody argued for returning to manual operations. The efficiency gains were real and the expansion they enabled was too valuable to reverse. The response was to re-engineer training to develop and maintain the judgment that automation was degrading. Beginning formally in 2013 and updated in 2017, the FAA issued guidance addressing the core finding: continuous use of automated systems failed to reinforce the knowledge and skills required for manual flight. Pilots were required to hand-fly regularly, practicing the skills that autopilot had made unnecessary on most flights and essential on the flights that mattered most. The industry reinvested in the capabilities the automation had made seem obsolete, because the moments when those capabilities were needed were the moments that determined whether everyone on the aircraft survived. The reinvestment was not confined to the cockpit. Aviation safety engineers, simulator designers, route network analysts: none of these roles existed in their current form before autopilot and its companion technologies made the expanded system possible. The expansion and the judgment investment were the same bet.

Aviation had decades to discover the atrophy, study it, and adjust after catastrophe. Organizations deploying AI at scale do not have that timeline. The same cycle is playing out in years rather than decades, and by the time the consequences surface, the gap between what the organization needs and what it can produce is already measured in cohorts of people who were never developed. The compression cuts the other way as well. The organizations that invest in maintaining judgment capacity will not merely avoid the atrophy. They will build the capacity to operate at scales and complexities that would be impossible without both the automation and the people capable of governing it.

Aviation did not restore everything autopilot displaced. The flight engineer’s seat never came back, and nobody argued it should. Those capabilities were drained deliberately and permanently. What the industry identified, after a documented warning and 228 deaths, was the single capability the system still depended on and could not manufacture on demand: a pilot who could fly the aircraft when the automation stopped.

The diagnostic question is not how much judgment the organization has today. It is which of that judgment the organization is still structurally dependent on, and whether anyone remaining is positioned to tell the difference. Answering it requires looking at inputs rather than inventory: how many people are making decisions whose outcomes they will personally carry, how many are close enough to those decisions to learn from them, and how many of the mechanisms that used to produce both are still standing.

The VP of Engineering who left carried something the system could not replace. The architect whose exit conversation named what nobody else would say left a gap no documentation could fill. The junior engineers who described shipping faster than ever and learning less than they had in years were reporting the early symptoms of a judgment formation pipeline that had gone dry.

None of it was anyone’s fault. The mechanisms that used to produce that judgment were never designed. They were inherited. And that inheritance is gone.

Building the replacement is now a design problem, and the organizations that treat it as one compound the advantage with every person they develop. The organizations that do not will converge toward the same plausible center as every competitor.

The autopilot handled cruise. The humans handled the moments that mattered most. Aviation answered a generation ago the question every organization now faces, and the answer was not to hold on to everything. It was to establish, before the moment arrived, which capability the system could not operate without.

Those moments arrive. They always do.

Footnotes

  1. The structural case for decision infrastructure, and the three mechanisms that compose it, is developed fully in “Engineering Coherence: Building Decision Architectures for the AI Era”.

  2. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.

  3. Grove, A. S. (1995). High Output Management. Vintage Books.

  4. Klein, G. & Kahneman, D. (2009). Conditions for Intuitive Expertise: A Failure to Disagree. American Psychologist, 64(6), 515-526. Reliable intuition requires an environment with regular patterns and the opportunity to learn them through feedback that is both rapid and unambiguous. Where those conditions are absent, experience produces confidence without accuracy.

  5. The stock and flow framing follows Donella Meadows, Thinking in Systems: A Primer (2008). A stock can appear stable while its outflow exceeds its inflow, and the level gives no information about the trajectory.


This is the kind of work I help leaders through directly. If what I've described is what's happening in your organization, reach out and we can talk it through.

Work with me →