Shannph Wong

Shannph Wong

When execution stops being the constraint

The Judgment Engine

Sustaining Judgment When Execution No Longer Builds It

Execution built judgment for free. That subsidy is over.

You are preparing for your annual board meeting, and for the first time in three years, you are not sure what story to tell.

The numbers are genuinely impressive. The product teams launched three new products this year and cleared a feature backlog that had haunted quarterly reviews for years. The engineering organization shipped more in the last twelve months than in the previous two years combined. The sales team adopted AI-assisted outreach and pipeline management. The go-to-market motion that used to take a quarter to stand up now takes five weeks. Customer support response times dropped by 40%. By every execution metric the board tracks, this was the best year the company has ever had.

A year ago, you told the board that the AI investments would drive a 30% gain in revenue growth. You believed it could be even higher, even a 50% gain was within reach. The reason was simple; if we could ship twice as much, move twice as fast, cover twice as much market surface, the results would follow. The teams executed on that logic flawlessly.

The actual gain is 6%.

In a previous era, you would celebrate that. A 6% acceleration in growth due to a single operational shift would have been made into a case study. But that’s not what you promised, or staffed for, or reorganized around, or what you told your investors to expect. The gap between how much was built and how little impact it actually produced points to something deeper, something you do not yet have language for.

What bothers you is not just the number. It is what the year felt like.

The place should be buzzing with energy. The pace of shipping alone should have created momentum. People should feel like they are winning. At the last all-hands meeting, you could feel something was off with the energy in that room, in a way you can’t trace to any single thing. Team meetings are productive but feel mechanical. Roadmap reviews surface priorities with sensible rationale, but nobody in the room has conviction about any of it. Decisions get made against options that all seem reasonable enough. You just can’t remember the last time someone pounded the table fighting for an idea they believed in despite the pushback they got in the room. The organization is making sensible choices, but it has stopped making brave ones.

You have also lost some key people this year. Not many, but the ones who really mattered.

The VP of Engineering had been with the company for over a decade, through the early-stage chaos, through three hard pivots, through a painful rearchitecture and platform migration, through the scaling crises that nearly broke the organization. And each and every time, the company and the team emerged stronger than ever. She resigned in September. She was not burned out from the work. If anything, it was the opposite. She told you she had spent the better part of the past year reviewing AI-assisted proposals, product specs, architectural recommendations, and every one of them seemed good enough; structured, defensible, needing at most a tightened trade-off or flagging a scaling concern. She had dozens of these landing on her calendar every week, each one adding to a growing accumulation of accountability over outcomes she could not really shape.

She decided to leave after a product review. The team presented three approaches to an integration architecture required for the strategy to expand to enterprise customers. The proposal was developed heavily with AI assistance, and each option was presented with evaluation criteria documented: integration complexity, migration effort, development effort, maintainability, security posture. The team asked her to weigh in on which direction to take. She looked through all three and knew from her experience that none of them had the extensibility to handle the complexity that the enterprise customers would demand. When she started to explain, the room just asked her which of the evaluation criteria reflected her concern and what data she had that would justify changing the metric. Her feedback did not neatly fit into the specified criteria. She did not have a metric. She had twenty years of building systems that looked just like this and seeing where they broke. The room did not exactly reject her judgment. They added it as a footnote, but then made the decision on the existing criteria and moved on without her.

She told you she stopped feeling like a builder and more like the owner of the rubber stamp. She missed the days when the team brought her a hard problem and her job was to see something they could not see. Now they brought her three well-formed options and her job was to vote on one. She was not tired from overwork. She was drained from the slow realization that the system no longer seemed to want what she was best at.

It reminded you of something. Years ago, you had watched the same pattern play out in design. Organizations displaced design conviction with A/B testing UX changes. The system produced optimized outcomes but could not produce original ones. The designers who carried the intuition for what a product should feel like eventually left, because the system had no way of using what they knew. At the time, you noted it and moved on. It was a UX design problem. Now the same structural pattern is reproducing itself across your entire organization, and you understand it was never isolated to one department. It was an early warning you did not recognize.

Your best senior architect left two months after the VP of Engineering. During his exit conversation, he said something you keep replaying in your mind: “I used to know why we were building what we were building. Not just the business case, the reason behind it. I don’t feel that anymore, and I don’t think anyone else does either.”

At the time, you dismissed it as nostalgia. You’re not so sure now.

The holes they left behind are larger than their role descriptions. When your VP of Engineering was in the room, she could look at the product approach and tell you in ten minutes whether it would hold together at scale because she had built and rebuilt enough systems to feel structural weakness before she could articulate it. That instinct is not in any documentation she left behind. It is not in the decision frameworks or the review templates. It walked out the door with her. The architect carried something similar. His code reviews were where junior engineers learned what mattered, what was worth fighting over and what was noise. Two of those engineers have told you, separately, that they are shipping faster than ever and yet somehow learning less than they have in years.

The sobering realization that stopped you from sending the board deck is that another year of the same approach will not make these results compound beyond the 6%. You can already feel the plateau coming. The first six months of AI adoption produced tangible acceleration. The next six produced refinements on the approach. The gains are real but they are flattening out, and the operational advantage that felt decisive in January has now become table stakes. Two of your closest competitors have closed the gap. One of them, a company you were not even tracking seriously eighteen months ago, has nearly caught up to you.

You did everything right. You moved early and boldly. You built real capability, not a headline. Yet here you are, sitting with this board deck draft, aware that the results do not match the effort, that the trajectory is not compounding, and that you cannot explain the gap by pointing to any single decision that was wrong. The people who might have told you why, the ones whose judgment the system had no way to use, are the ones who are no longer in the building.


What happened here has an explanation. The organizations that have not yet felt it are on the same trajectory. The question is not if this will happen, but when.

When the output metrics show real gains but the business impact does not match, the standard solution stays at the execution level: the tools need tuning, the workflows need refinement, the teams need better prompts or more training. That diagnosis is comforting because the solution is more of the same, done better. It is also insufficient.

The deeper problem is not how an organization uses AI. It is what is happening to the people and the systems that do the critical work nobody accounted for. This is the work that AI displaces without replacing.

Under the accelerated pace of the AI era, execution scales faster than coherence. Coherence is fundamentally different from alignment. Teams perfectly aligned to their goals can still produce a fragmented result if those goals were ambiguous, or if achieving them at AI-accelerated speed produced side effects nobody designed for. Alignment is about hitting the metric. Coherence is about staying true to the intent. To maintain that, organizations need to deliberately design their decision infrastructure1: clear decision ownership, feedback loops that detect drift from company-wide objectives, and incentives that reward coherence instead of just raw output. That infrastructure is what keeps an organization from breaking apart under its own speed. But even the best decision infrastructure is only as good as the people whose judgment it runs on.

Decision infrastructure depends on the quality of people’s judgment. People who can tell the difference between drift that matters and noise that does not. People who know when a correction is urgent and when it is merely reactive. People who can feel that an integration architecture will not hold under enterprise complexity before the data confirms it, or that the adjacent market the product team is expanding into behaves fundamentally differently at scale than the models project.

Right now, the need for that kind of judgment is higher than it has ever been, and the existing supply is shrinking. People leave. People burn out. People disengage when the system no longer demands what they are best at, and reduces them to a rubber stamp. Each departure takes with it the judgment that no document or decision framework can replace. New people arrive, but organizational judgment is not something you can hire in finished form. Judgment is built through years of operating inside the system. It cannot be learned in an onboarding video.

Regardless of where an organization is in its AI adoption, one question persists: “Are we developing judgment capacity at the rate that it is being consumed?” For organizations still early in the AI journey, this answer should shape what to design for. For organizations deep into it, this answer explains what has already started to erode. If the answer is unclear, the company may not recognize that everything they have built to keep decisions coherent depends on the very people it has stopped investing in developing. The cost of discovering that too late is not a quarter of missed targets. It is years of judgment formation that cannot be closed on demand.

The answer starts with understanding judgment itself: how it forms, why it is harder to develop than most organizations realize, and why the ability to systematically cultivate it may be the most durable advantage available in a world where every competitor has access to the same AI models.

When More is Less

Nearly every AI investment an organization makes is based on the same underlying assumption: more information leads to better decisions. That is only true up to a point. After that, the bottleneck moves, and most organizations do not notice it because the new bottleneck feels nothing like the old one.

When information was scarce, the bottleneck was producing good options. When AI generates options in seconds, the bottleneck moves to evaluating them. Anyone who has stood in a paint store staring at a hundred shades of white understands what this feels like. Every paint swatch is plausible. The differences between them are real but narrow. You are left with two options: try to instantly become an expert in color theory of undertones and light refraction by watching YouTube on your phone, or ask the store clerk which one is most popular and just go with that. The interior designer walks in and picks the right one in thirty seconds, because she knows your south-facing windows will wash out the cool tones, that your maple floors pull warm, and that two children under three means the finish needs to survive crayon and grape juice, not just look right under showroom lighting. Her judgment is not about the swatches. It is about everything the swatches do not contain.

Now imagine asking that designer to make that call for twelve houses a day. The first three get her full attention. By the eighth, she has stopped asking about the muddy sheepdog who greets everyone at the front door. By the twelfth, she is picking the one that looks close enough. The judgment is still there. The bandwidth to apply that judgment has been drained by the volume of plausible options that were all asking for the same scrutiny. Organizations are now creating this problem at scale. The returns do not just diminish. They reverse, because the finite resource that distinguishes the critical decisions from routine ones has been spent on decisions that did not require it.

AI-generated options appear comprehensive because the AI pre-evaluated them against a set of criteria. The integration architecture comes with a comparison matrix. The product strategy comes with projected customer impact estimates. The hiring plan comes with pre-modeled scenarios. Every option is framed so it can be easily compared: cost, timeline, projected return. The decision process gravitates toward these criteria because they are there, and measurability stands in for comprehensiveness. The system never prompts the question that judgment would have asked: are these even the right criteria to drive the decision?

What gets lost in this substitution of measurability for comprehensiveness is not analysis. It is the capacity to question the decision framework the analysis operates inside. A scoring rubric captures the dimensions someone thought to include. It cannot capture the dimension nobody named. The product instinct that says the impact projections are modeled on the wrong customer segment operates on a dimension the rubric does not contain. The act of choosing among well-formed options feels like judgment, but you can only choose from the options presented. Judgment questions whether the right options are on the table to begin with. Organizations have collapsed the act of defining the rubric and evaluating the criteria into one step, and that collapse is frictionless. Nobody has to explicitly choose metrics over experience. The numbers simply arrive first.

Over time, the role of the experienced person narrows from shaping decisions to approving them. That shift feels like efficiency. The people who carry the deepest understanding are the first to feel the loss.

The Dangers of Plausibility

A few years ago, I was at a Spanish restaurant in Shanghai and ordered paella, Valencia’s traditional rice dish. When it arrived, I was impressed. Everything looked right: the rice, the seafood, the saffron color, the presentation in the carbon steel pan. Every visual cue said paella. But when I tasted it, it was off, by a lot. It tasted like I was eating fried rice. The chef had clearly seen what paella was supposed to look like and had replicated the appearance with real skill. What the chef had never done was taste an authentic paella.

The gap between the visual output and the actual thing is the gap between pattern replication and judgment. The chef reverse-engineered paella from its appearance. With enough reference images and enough technique, the replication was convincing. What could not be reverse-engineered was the taste, because the taste was never in the reference material. It lived in the experience of eating and preparing the real thing.

Organizations have always evaluated the quality of thinking by its coherence. If an argument is well-structured and internally consistent, the instinct is to assume they had done the work. AI produces output that passes that filter without the underlying process, and the gap is invisible for the same reason the paella looked authentic: the reference material never contained what was missing.

In practice, AI is very often giving the right answers across the routine decisions that occupy much of the organization’s attention. Review cycles shorten, pushback softens, and the muscle for questioning plausible-sounding reasoning weakens because exercising it feels unnecessary most of the time. The habit of scrutiny atrophies. When a decision arrives where historical patterns do not apply and the ambiguity is genuine, the atrophied muscle is unprepared to engage. The people who have built something like this before know from their experience what will hold and what will not, but they cannot fully articulate why. That knowledge is activated by what feels to them like instinct, not analysis.

The paella analogy has a second layer that is pertinent. Imagine that restaurant had gotten a new rice cooker that changed the economics of the dish. The labor that made paella slow and expensive to prepare, watching the rice, adjusting the heat, judging the texture by feel, is no longer needed. The restaurant can run paella variants faster and cheaper: a seafood version, a mixed version, a spicy version. The natural conclusion is that with enough iterations, one of them will converge to the real thing. But, it will not. You can run a hundred variants of a rice dish and never produce an authentic paella, because the rice cooker only accelerated the production process, not the judgment process. Each variant is faster to produce and no closer to the goal, because speed of iteration without taste just means going in every direction faster. The tool that makes experiments cheap does not give the experiments its direction. Direction requires someone who knows what real paella tastes like.

There is a persistent narrative that because AI has made experimentation cheap, organizations can iterate their way to the right answer. Run enough experiments, test enough variants, and the data will eventually tell you what works. That narrative deserves direct challenge. If you already know where you are going, cheap experimentation is a genuine accelerant. You get there faster with fewer expensive wrong turns. But if the direction itself is unclear, or if the organization has lost the people who know what the real thing tastes like, then cheap experimentation just means generating more evidence for options that were never worth pursuing. The speed is real. The convergence is an illusion.

Judgment is the capacity to look at a set of options that all sound right and see the one that truly is, or to recognize that none of them are, and to have the conviction to act on that recognition even when the data does not yet support it. That capacity is not something you can buy, install, or prompt into existence. It has to be developed. And that raises a harder question: where does it come from?

Where Judgment Comes From

We tend to romanticize judgment, talking about it like it is a trait or instinct that some people are born with. That framing is comforting because it lets organizations treat judgment as a hiring problem. Find the people who have it. Put them in the right seats. Move on. But when a chess grandmaster glances at a board mid-game and immediately sees the winning line, what looks like instinct is actually the residue of thousands of games of deliberate analysis, repeated under enough pressure to compress into automatic recognition. It was built, not born.

Human cognition operates on two tracks2, that explains almost everything happening inside organizations right now. System 1 track is fast, automatic, always running. It is what lets you drive home while listening to a podcast without thinking about where to turn next. System 2 track is slow, deliberate, and hard. It is what kicks in when you see the flashing lights of a firetruck on the road ahead and decide to take an early exit. The boundary between them is not fixed. It moves with training. Judgment is what happens when the System 2 work, done enough times under real conditions, compresses into System 1 expertise. If judgment were innate, losing people with good judgment would be a talent problem. Because judgment is developed, losing the ability to develop it is a structural problem, and one that hiring alone cannot solve.

Consider what is actually happening when judgment operates. A CPO who has led enough product expansions into adjacent markets carries a model built from the ones that worked and the ones that stalled. When the team presents a plan to move upmarket into enterprise, the CPO can feel that the plan underestimates the complexity, not because they have analyzed this specific market but because they have watched six similar expansions flounder on the same invisible dependencies: longer sales cycles that starve the core business of attention, support demands that do not scale linearly, and a buyer whose procurement process will reshape the product roadmap in ways the plan does not account for. Their System 2 did this work years ago. It arrives now as pattern recognition, as a conviction they would struggle to defend in a spreadsheet but would stake their credibility on.

What separates the CPO from someone with the same title and similar tenure is not the experience itself. It is what was done with it. Experience and judgment are related but distinctly different. You can accumulate years of experience without developing judgment, if the experience was never compared against what was expected, and never integrated into an updated insight into how the system works. Twenty years of doing the same thing is not twenty years of judgment development. It is one year repeated twenty times. Judgment forms only when the cycle closes: the outcome is examined, the world model is updated, and the updated model is used to shape the next decision.

AI cannot close that cycle. It can generate options, evaluate them against known criteria, and present them with coherent rationale. But it cannot examine an outcome against what was expected, update a world model, and carry the weight of that update into the next decision. That is the judgment formation loop, and it runs only in people.

What AI does instead is reshape the environment in which those people operate. The proposal review that used to be two hours is now thirty minutes because the AI-generated proposals arrive with the standard concerns already addressed. The architecture review is now a checklist discussed async. The quarterly planning process moves through smoothly because every initiative comes with a defensible rationale. Each of these is genuinely useful. But when the output consistently looks competent, the reviews get shorter, the pushback gets softer, and eventually the meetings themselves start to feel unnecessary, because if the feedback is never substantive, why convene to deliver it? System 2 requires a trigger, and the trigger has been filtered out before anyone sees it. That is the trap. The process that would have developed and maintained judgment is hollowed out from the inside, and it feels like efficiency the entire time.

An individual can decide to slow down and think more deliberately. An organization that has restructured its processes around speed cannot make that choice the same way. The blind spots are not in the people. They are built into the processes, the cadences, the incentives, the infrastructure the company runs on. The friction that would have signaled deeper thinking was needed is precisely the friction being removed, and its absence registers as progress. Once the process has been reduced to a rubber stamp, it is not just the current generation’s judgment that goes unexercised. It is the next generation’s judgment that never forms. The reviews where people still building their judgment learned by watching experienced people think are the same reviews being shortened, moved async, or eliminated. The organization has not just stopped using its judgment. It has shut down the factory that produces it.

The Deeper Stake

The loss of judgment capacity does not just make individual decisions worse. It removes something much more fundamental, the ability to discover when the organization’s own direction is wrong.

Organizations that recognize the need for coherence under acceleration are building the decision infrastructure to maintain it. That infrastructure is essential, but it solves a specific problem: keeping local decisions pointed in the same direction as the company’s strategic bet. It works as long as the bet is right. But bets have a shelf life, and the world does not hold still. Coherence infrastructure is designed to detect and correct when teams drift from the thesis. It has no mechanism for detecting when the thesis itself has drifted from reality. The better it functions, the more confidently the organization stays on course, even when the course is wrong. Without something that can challenge the bet itself, coherence infrastructure becomes a system for pursuing the wrong direction with increasing precision. That capacity to challenge the thesis must be standing, not periodic, because one step in the wrong direction requires two to correct.

Earlier in my career, I joined a GPS navigation company as the senior architect as employee 50 and helped it grow to over a thousand in under a decade. My credibility came from scaling the architecture that got us there. But under the pressure of user growth doubling every six months and the demands of major mobile carriers who represented 80% of our revenue, that architecture had become a monolith, forked repeatedly to support each carrier’s specific requirements. It had become a trap, with every capability tightly coupled to the logic on the mobile device. Each additional carrier represented the risk of an additional fork. By the time Google announced free turn-by-turn navigation the week before our IPO, the threat was not just to our core product. It made visible what was already true: we needed to diversify into adjacent markets, and the architecture I had spent years building made that nearly impossible. We could not decouple the capabilities from the application they were built inside.

The opportunity was there. Our nascent automotive team could offer something nobody else had: a hybrid connected and embedded navigation experience with traffic-optimized routing, enhanced search, live reviews. On paper, we could leverage everything we had already built. In practice, the existing architecture made it infeasible to port any of it. The capabilities existed. They were locked inside a codebase that could not serve anything other than what it was originally built for.

I proposed rebuilding the entire product on a generalized API platform, designed to serve every product line from a single foundation: mobile, automotive, and whatever came next. The bet was enormous. If it failed, it would destabilize the business that still paid for everything. I asked for a tiger team of seven engineers and three months to prove it. I picked the team and led it personally. We hit every checkpoint. The skeptics shifted their objection: a proof of concept is not production. I asked for three more months. We shipped the full product to the App Store with all current features and the next version’s roadmap, cutting load times in half and improving key response times by thirty percent.

But the replatform came with a cost I did not fully anticipate. The new architecture eliminated the need for parallel teams maintaining forked codebases. Team leads saw their scope shrink and their advancement path close. About a dozen engineers left before I ever got the chance to talk to them about the opportunities I believed this would open up. The improbable choice cost something even though it turned out to be right. Years later, that same platform enabled the company to secure its two largest deals in history: embedded connected navigation across GM’s and Toyota’s entire fleets. Neither would have been possible on the old architecture. The judgment call was validated, but only in retrospect, and only because the organization absorbed the cost of making it.

Consequence develops judgment. Making decisions with real stakes and living with the outcomes, including the outcomes you did not anticipate, builds the world model that allows a person to see things others cannot.

Judgment enables conviction. Once you have built that world model through consequential experience, you develop the capacity to trust it even when it conflicts with the available data. This is not recklessness. It is the earned confidence that comes from having seen a pattern before and knowing what it predicts, even when you lack the language and metrics to defend it in the room. You understand the risk of taking the path. You also understand the risk of the missed opportunity if you do not. Conviction is what lets you weigh those two risks against each other when only one of them is visible to everyone else, and then make the improbable choice: the one the room would not select, wagering your credibility on a reading of reality the data does not yet support.

The improbable choice is where new understanding comes from. The probable option, by definition, reflects the current consensus about how the world works. When the consensus is wrong, which it eventually always is, the only way to discover that is by choosing something the consensus would not choose and learning from the result.

If any link in that chain breaks, the organization defaults to selecting among options that all reflect the same underlying model of how the world works. It optimizes within a frame rather than discovering when the frame needs to change. When every organization is drawing from models trained on largely the same data, the frame is not even its own. Each company, independently and rationally, selects the options that the shared corpus rates as most probable. They are all accelerating toward the same center.

This is what is at stake when judgment erodes. The organization will remain coherent. It will remain efficient. It will execute its thesis with increasing precision. And if the thesis is wrong, it will converge on that wrong answer faster than any competitor, with every correction mechanism working exactly as designed to keep it on that course.

What Was Never Written Down

The prevailing narrative assumes that AI’s trajectory, if extended far enough, will eventually close the gap between plausible output and genuine judgment. The models are improving on every measurable dimension. Context windows are expanding. Reasoning capabilities are advancing. The trajectory is real. But the assumption that it leads to judgment has a structural problem, and it is not about the capability of the models. It is about the material they are built from.

Consider the PRFAQ, the narrative decision document that Amazon uses to drive strategic alignment. A well-written PRFAQ is a dense, rigorously structured artifact. It opens with a press release that articulates what it uniquely solves for the customer, why announcing it matters, the approach, the trade-offs, the risks, and the expected impact. This document survives for years across the teams. What does not survive is the seven weeks of debate that preceded it. The moment when someone realized that the customer need everyone was building around was just a symptom of a deeper problem. The hallway conversation where two directors resolved a disagreement that would have blocked the project for a month. None of this is in the document.

This gap is not a documentation failure. Even a human trying to reverse-engineer the judgment behind a set of decisions would face the same problem. The same document could have been produced by entirely different reasoning processes, under entirely different pressures, with entirely different things left unsaid. The number of plausible judgment calls that could have produced the same written artifact is effectively infinite. We are asking AI to solve a problem that the artifacts themselves do not contain enough information to resolve, because the signal was never in the material.

The designer Massimo Vignelli, talking about typography in the documentary Helvetica, put it precisely: “Typography is really white… It is the space between the blacks that really makes it.” The mastery lives in the invisible relationships that determine whether the whole composition holds together or falls apart. Most people see the letters. The letters are not where the craft is.

The training data captures the blacks. The reasoning, the struggle, the political context, the felt sense of what mattered and what was noise, the relationships that shaped which arguments were heard and which were dismissed: these are the whitespace. They are where judgment lives. They are rarely written down, and when they are, the written version is a flattened summary of an experience lived in three dimensions. AI can reproduce the visible artifact with increasing skill. It cannot reproduce the invisible space that gave the artifact its meaning. That space cannot be captured in a form any model can learn from.

This is why apprenticeship keeps surviving every attempt to replace it. The mechanism is stubbornly simple and it has never changed: proximity to someone making real decisions with real consequences. A medical resident does not learn judgment from textbooks. The resident learns it from standing next to an attending who is diagnosing a patient whose symptoms do not fit the textbook, watching which plausible explanations are discarded and why, and then making their own calls with real patients whose outcomes they will carry. The content of what is taught changes completely across eras of medical practice. Yet, the structural relationship between master and apprentice has survived every previous technological transformation because it is the only mechanism that transmits the whitespace.

When the work that created that proximity is automated, the mechanism does not find a new vehicle. It breaks, unless the organization deliberately builds a replacement. This is not a narrow concern about software engineering culture. It is happening simultaneously across every knowledge profession that has adopted AI at scale. The organizations that recognize it will sustain their judgment capacity. The ones that do not, will discover the cost long after the cause, and by the time they do, the judgment development pipeline they needed will have been dry for years.

The Observation/Consequence Framework

Apprenticeship works because it does two things at once, and neither one alone is enough: observation and consequence.

Observation is absorbing how experienced people think. Sitting in a design review and studying how the room navigated ambiguity, not just the final decision. Noticing which questions a principal engineer asks about an architecture proposal, and realizing that they were all about what would break under conditions nobody else had considered. Observation is how pattern recognition develops. It is how a developing practitioner builds the vocabulary of judgment before they can exercise it themselves.

Consequence is making decisions where the outcomes are real, the stakes are felt, and the blast radius of a failure is genuine, even if bounded. It’s owning the customer call where the answer is “we got this wrong,” not merely reading the summary afterward. It’s about standing in front of the room and recommending a path when choosing the wrong one costs the team three months and you’d be on the hook to explain why.

Observation without consequence produces people who can recognize the shape of good judgment but have never developed the taste for it. The paella that looks right. They can identify the right questions in retrospect, but they never had to own an ambiguous decision and be accountable for its outcome. On the other hand, consequence without observation produces people who learn exclusively from their own mistakes. It is a slow, expensive, grueling climb to reinvent the lessons the organization already learned.

The most vivid example I have of observation and consequence working together is another Amazon mechanism called the Correction of Error (COE). When something impactful breaks at Amazon, the accountable leader writes a COE: a structured, rigorous analysis that covers not just what happened and why, but how to detect the problem earlier, respond more quickly, close on the diagnosis faster, resolve the issue in ways that minimize customer impact, and prevent the entire class of error from recurring. It is not your garden-variety post-mortem. It is an operating system for turning failure into organizational capability. The standard is exacting. The senior leaders read every word.

I wrote a COE after a failure in the Amazon Catalog Pipeline, the systems through which all 5 billion daily catalog updates are applied across practically every Amazon property worldwide: marketplaces, Prime Video, Whole Foods. This pipeline depended on AWS Kinesis. When a massive Kinesis outage hit in US-EAST-1, the pipeline went down with it. The blast radius was immediate and global: all marketplace regions blocked simultaneously, nearly 200,000 merchants unable to update their listings, Whole Foods forced to suspend orders because inventory updates could not be processed. Thirty-one and a half hours from impact to complete recovery.

Handling the outage event demanded judgment under pressure: prioritization calls when every priority was legitimate, understanding impact across systems coupled in ways the documentation did not capture, making decisions without waiting for approval because the compounding impact was happening in real time. But the COE process that followed is where the deepest judgment was formed.

The COE structure forces you to trace causal chains to a depth that few post-mortems would take you. It will not let you stop at “the dependency failed.” It demands you trace why that single dependency was a point of failure across all regions simultaneously, why failover was not possible, why the recovery took nineteen and a half hours on top of a twelve-hour outage, why the mechanisms that existed had never been validated at this scale. Each level strips away another layer of the system’s implicit assumptions. By the fifth level, you are no longer analyzing an incident. You are understanding the architecture of a system you thought you already understood.

Thirty corrective actions followed. Each with a named owner, a due date, a priority. Many of them were mine. The specificity of those actions is itself the evidence that judgment has been formed. You can only prescribe at that granularity if you understand the system at that granularity. And you only understand it at that granularity because the structured COE process forced you there.

That is the COE author’s experience. The consequence half of the equation. Now consider the reader’s.

Amazon invested heavily in making COEs visible in perpetuity. The documents were shared, searchable, referenced by senior leaders. The intention was that the organization would learn the lessons from each significant failure and not repeat them. What happens in practice is much more limited. People did read COEs retroactively on occasion, but that was primarily when they had a current problem and wanted prior art to help explain it. That was useful. It was fundamentally different from the value of the judgment the author developed. The reader received valuable information, but the author developed deep judgment. A friend who recently joined Amazon told me they had read my COE. They could describe what went wrong. As the author, I can still feel, six years later, the pattern of cascading failure that would tell me a new system was vulnerable to the same structural weakness, before the metrics confirmed it.

The COE process was transformative for everyone who went through writing one. The first COE an engineer contributed to at Amazon taught them more about the operational culture than any onboarding document could. It taught the bar. The act of being held to that standard, of having your analysis challenged and sent back for revision, of producing something that senior leaders would evaluate against a bar you knew you could not shortcut, was itself the judgment development instrument. The document was not the mechanism. The doing was the mechanism.

The warning and the promise coexist. The COE is proof that judgment can be deliberately developed through structured processes that combine real consequence with enforced analytical rigor. It is also proof that the development happens through the doing, not through the reading. Organizations that make judgment visible are doing necessary work. Organizations that assume visibility alone will transfer judgment are confusing the output with the journey that produced it.

Judgment as a Renewable Resource

The Amazon Catalog pipeline’s thirty one and a half hour outage was a massive drain on judgment capacity: senior engineers and leaders spending everything they had under pressure. The COE process that followed converted that judgment expenditure into renewal, forcing deeper understanding across everyone involved. But the COE develops a specific kind of judgment: operational. Product judgment, strategic judgment, market judgment all require different mechanisms, different consequential experiences. The COE is proof that deliberate design works. It is not a template for every kind of judgment an organization needs.

During COVID, organizations discovered what happens when observation breaks. Remote work stripped away the ambient context through which people still building their judgment absorbed how experienced people thought. But the work itself was still there. They were still building, still encountering real consequences. Remote work degraded observation. AI now threatens both.

Judgment inside an organization behaves like any other reserve3. It gets consumed every time a leader makes a hard call, every time a senior practitioner is spread across one more review, every time the organization expands its surface area faster than its people can keep up. It gets replenished when people develop through consequential experience: apprenticeship, structured processes that force analytical depth, roles that carry real stakes. When the consumption exceeds the replenishment, the reserve declines. The current level tells you nothing about the trajectory. It can look full and already be in irreversible decline.

The good news is that judgment is a renewable resource. It can be renewed through the density and intensity of consequential experience: apprenticeship, structured processes that force analytic depth, roles where real decisions carry real outcomes. It is not based on time. Someone who spends three years inside an environment where decisions connect to consequential outcomes that reshape how they see the world will develop judgment faster than someone who spends a decade in an environment where decision consequences are diffused or absorbed by someone else. The renewal rate is not measured by calendar years, but by how many cycles and how much weight each one carries.

For decades, the balance between consumption and replenishment held without anyone designing it. Execution was slow enough that the demands on senior judgment were bounded by how much the organization could ship. Juniors developed judgment through the execution work itself: writing the code, owning the customer call, sitting in the room when the trade-off was made. The judgment formation pipeline ran as a byproduct of normal operations. Nobody had to think about it because nobody had to design for it.

AI broke both sides of that balance at once. The demand for judgment has exploded alongside the output. Every additional release, every concurrent initiative, every expanded surface requires someone to evaluate whether what the organization is collectively building still makes sense as a whole. The supply has not kept pace.

On the judgment renewal side, organizations are hiring fewer juniors because AI handles much of the work those roles used to do. The juniors who are present are doing different work: learning to use AI effectively rather than developing deep familiarity with the systems, customers, and markets those AI tools operate on. When the production incident happens at 3 AM, the on-call engineer who has been prompting models for a year has a different relationship to the system than one who spent that year writing the code that runs it. When the product review surfaces a strategic tension, the junior PM who has been generating specs with AI assistance has a different relationship to the customer than one who spent that year in the field learning why customers actually churn. The apprenticeship pipeline is not just shrinking. The nature of what flows through it is changing, and the judgment it produces is thinner.

The instinctive response when the pressure builds is to hire for it. Bring in senior people who already carry the judgment the organization needs. It is the obvious move, and it pulls exactly the wrong lever. Every senior hire draws on the judgment of the people who are already overextended, because onboarding at that level requires context transfer that only they can provide. And it does nothing for the pipeline beneath them. If the developmental mechanisms have been removed, even senior people who want to invest in the next generation have nothing to invest through. The pattern reinforces itself: depletion creates pressure, pressure triggers hiring, hiring increases the burden on the people already carrying too much, and the burnout and attrition that follow deepen the deficit the hiring was meant to address.

The organizations that recognize this early will design for it. The organizations that do not will keep pulling the same lever, hiring to fill a deficit that hiring deepens, until the reserve is too depleted for any short-term fix to resolve.

What to Build

This judgment formation pipeline can be rebuilt. It requires deliberate investment on four fronts: visibility infrastructure that restores the conditions for observation, consequential role design that puts people still building their judgment in contact with real stakes, decision processes that can actually receive what experienced judgment produces, and mechanisms that contest the current generation’s judgment rather than merely reproducing it.

Visibility infrastructure

People used to learn by proximity. They sat in design reviews, overheard hallway debates, watched experienced people navigate ambiguity in real time. That learning has been disrupted twice: first by remote work, then by AI-accelerated execution that compressed or eliminated the activities where it used to happen. Returning to the office has not restored what was lost. What replaces it must be designed.

Shopify’s CEO Tobi Lütke saw this problem early. As AI adoption spread across the organization, the reasoning behind decisions was becoming invisible. Individual engineers were getting faster, but the thinking that produced their work had moved into isolated conversations with AI that nobody else could learn from. Growing up in Germany, Tobi drew on a concept from his apprenticeship years: the Lehrwerkstatt, a teaching workshop where the shop floor is the classroom and apprentices learn by working alongside experienced craftsman in the open. So Shopify built their AI-equivalent of that open classroom, an AI agent called River that lives in the company’s Slack operating under a single hard constraint: no direct messages. If you try to message it privately, it declines and asks you to create a public channel. Every interaction is searchable. When senior developers resolve complex issues through River, those conversations become a living knowledge base. Someone’s hard-won fix for an obscure error message becomes the next person’s starting point.

Nearly six thousand Shopify employees now work with River in the open. The results that matter most are not the productivity metrics. They are behavioral changes: a support engineer can see a developer’s log query in another channel to debug an issue and can leverage the same technique the next day to identify a customer’s issue. New hires scroll through existing channels to see how experienced colleagues frame product questions before sending their first one. The AI did not teach them. The visibility did. The principle extends beyond AI tooling to any process that makes the act of deciding observable: open design reviews where the reasoning is the point, not the artifact. Decision artifacts that preserve not just what was decided but what was considered, traded off, and deliberately deferred.

Visibility develops pattern recognition. It is necessary but not sufficient. The COE process demonstrated exactly this gap: the readers gained information, but the writer gained judgment. Observation alone produces people who can recognize the shape of good judgment but have never developed the taste for it. For that, you need consequence.

Consequential role design

The medical residency model is the template. The scope is bounded, but with real patients, real decisions, and real outcomes. The attending physician is close enough to intervene but far enough to let the resident carry the weight of the choice. The attending’s judgment is present as a safety net, not as a substitute. That structural relationship is how judgment has always been developed: the learner makes calls that matter, receives real feedback from real outcomes, and builds their world model by accumulating the consequences they personally carried.

For decades, technology organizations ran their own version of this without designing it. A junior engineer who shipped code that caused a production outage at 3 AM felt the consequences because it paged their tech lead who was on-call. Real customer data had to be restored, and the recovery and the post-mortem were theirs to deliver. A new PM who descoped an “edge case” for an integration felt the consequences when they sat in front of a customer to apologize for breaking the financial reporting pipeline that prevented them from closing their monthly books. A junior UX designer who tweaked the onboarding flow felt the consequences when the abandonment rate after their new third screen doubled overnight and the growth team’s quarterly number went backward. The scope was bounded, but in every case the consequences were real. They learned that operational risk is felt at 3 AM, not described in a training video. They learned that a spec is not the same as a customer commitment. They learned that good intention does not automatically make a good user experience. Judgment that can be exercised in the most critical moments is formed through the accumulation of these consequences, each one closing the gap between what the practitioner expected and the weight of what actually happened.

When AI automates the execution work, it removes the vehicle through which the consequence was delivered. The code still ships, but the junior engineer did not write it in a way that forced them to understand the system’s dependencies. The analysis still gets produced, but the new PM did not build it from raw metrics in a way that forced them to confront what the data could not tell them. The onboarding flow still gets designed, but the junior designer never iterates through fifteen failed versions in a way that builds the instinct for where users lose patience. The output is there. The developmental experience is not.

Given that, the next logical question would be: “If AI can produce the output, why would you assign a junior practitioner to produce it more slowly and with more risk?” To answer that, you need to consider what really happens when a junior engages with a senior. Take the junior engineer who designs and builds a new integration capability the old way, under supervision. When it breaks in staging, he asks the senior engineer to sit down with him and walk through why. The discussion is not just about what went wrong, but why the dependency it broke on was invisible, why the system behaves differently under load than what the documentation suggests, why this particular failure pattern has been lurking around since the first version was deployed. Halfway through the explanation, the senior engineer pauses as they realize the dependency they are describing was a real constraint two years ago, but the client the platform team delivered six months ago specifically eliminates that constraint. They realize they spent the last 10 minutes talking about an anti-pattern that no longer applies. The junior’s failure surfaced something that their own instinctual reflex had stopped questioning.

That engagement produced judgment in both directions. The junior learned what the system’s documentation could not teach. The senior was forced to articulate assumptions they had long since stopped examining and discovered that one of them had gone stale. The organization that eliminates this work in the name of efficiency is not just optimizing away its renewal pipeline. It is removing one of the few mechanisms that keeps its most experienced judgment from going unexamined.

Systems that can receive judgment

There is a classic Sherlock Holmes story, “The Adventure of Silver Blaze,” where Holmes cracks the case by noticing what did not happen. The critical clue, “the curious incident of the dog in the night-time,” was that the guard dog never barked. The killer was someone the dog already knew. The absence of barking was the signal.

Experienced judgment frequently operates in exactly this register: sensing what should be present but is not. Data-driven decision processes are, by their nature, blind to absence, evaluating only that which is presented. A risk matrix captures the risks someone identified, but it cannot capture the risk nobody thought to name. A decision process that can only accept numerical inputs creates no room for judgment that operates on the absence of signal. The experienced practitioner’s concern does not lose a debate. It never enters it to begin with.

This is what happened to the VP of Engineering. Her concern was about what was absent from all three proposals, a structural property none of the criteria covered. The room had no mechanism for weighing that kind of input. The decision criteria were already defined, the scores were already populated, and her judgment had no anchor on which it could hold. She did not need to win the argument. She needed the argument to be possible. The system closed it before it began.

As a leader, I have been on the other side of that moment, the one who had to decide whether the judgment arriving at the gate was a drift to resist or a signal to let through.

At that same GPS navigation company, in my role as chief architect, consistent architectural standards were the only thing keeping the organization from tearing itself apart. Customers, products, and headcount were all doubling every year, each with their own priorities. Half the engineering team had joined in the last six months and was still struggling to understand how to apply the patterns we already had. The other half was struggling to maintain those patterns under relentless pressure from the business to deliver more and faster. Keeping everyone on consistent foundations was the only way to turn what would otherwise be a non-linear explosion of complexity into something linear enough for the teams to catch up. That meant most proposals that introduced something fundamentally different were most often not worth the complexity they would create.

Unfortunately, our search experience was genuinely terrible. If a user misspelled a business name or the database listed it differently, the results were useless. The underlying data was unreliable: convenience stores registered as restaurants, a gym reported as an auto repair shop, stale information everywhere. Everyone knew it needed to change.

The team proposed replacing the entire approach with a new search engine built on Apache Solr and Lucene, a custom point-of-interest database built by crawling and deduplicating across dozens of sources, machine learning pipelines running overnight to improve accuracy. It was a fundamentally different stack. New open-source projects, new infrastructure that was memory-intensive and expensive, new deployment patterns, and systems that nobody in the organization knew how to operate, monitor, or validate. The outputs were non-deterministic. Nobody could tell you whether a new model was genuinely better than the last one. I pushed back hard.

My pushback was legitimate. I demanded more evidence, more demos, more proof of how it would scale and how we would keep it running. But there came a turning point where I recognized that the alternative to accepting the risk was keeping the experience we already had, the one that embarrassed us every time a user searched for a McDonald’s and got a result forty miles away. Innovation always looks unreasonable against the standards of the current system. So, I stopped listing reasons it could fail and started working on how to make it succeed: real benchmarks built from real data, early experiments with elastic cloud infrastructure for the data pipelines, test suites redesigned around probabilistic accuracy rather than deterministic pass/fail. It was a world I had to embrace because the world we were in was no longer good enough.

What made that turn possible was that the system had a way to accept a judgment signal that contradicted the established direction. In this case, I stood in for that system because I held both the authority to enforce coherence and the experience to recognize when a deviation was actually a signal about where the world was heading. The lesson is not that you need the right person at the gate, but rather the gate has to be built to enable the contradictory signal to come through. For me it was a learning moment that reshaped how I thought about authority. For an organization, it has to be more durable than one person’s ability to detect. The methodology has to define it and the decision framework has to support it. But in the AI era, it has to be built in, because the velocity at which contradicting signals arrive, and the cost of suppressing them, are both higher than they have ever been.

Metrics are essential, but decision processes must also create an explicit space for experienced judgment alongside that data. The system needs to ask “what does the data say?” and “what is your experience telling you about what the data is not showing?” It needs to treat a principal engineer’s instinct that something will not hold at scale as a signal that warrants investigation, not an opinion to be footnoted.

Judgment renewal

When you are trained as a scientist or an engineer, you are hardwired to a specific instinct: solve for the variable. When there were multiple variables, find the relationship between them. Run the experiment to validate or invalidate the hypothesis. For me, it felt like the right answer was just about solving for one more variable, and the rigor of that approach got me far. I could show demonstrably why an enterprise service bus attached to SOAP services was wrong for a near-real-time service because of the latency thresholds, and equally why a messaging system built for high-frequency trading was overkill.

The moment I felt that instinct fail was during a product review for a music streaming service. The product team was presenting a proposed feature to let users follow their favorite artists and receive notifications about new releases, concert dates, and the like. As I thought through what it might look like, my instincts reached for the scalability problem: what notification spikes would we have to handle during busy summer or holiday release windows? How would the system handle the load spikes? Then I saw the expression on the product manager’s face and realized I was answering a question nobody had asked. This was a review about whether the idea had customer value worth pursuing, not a specification review for implementation. He was looking for feedback on what we were leaving on the table for customers, not gaps in his PRD. The data point I was reaching for would not have answered his question, and asking for it would have distracted him from the one that mattered.

It was a small moment, but it named something larger. The instinct that had been right for most of my career, that one more data point leads to a better decision, had become a reflex I was applying indiscriminately. The “right” solution was not the same as the verifiably “correct” one. The right solution was the timely one, the one that moved the needle and anticipated where the world was going. Every problem has an expiry date. In a review meeting, asking the team for another data point can sound rigorous, but it can also reveal that you have not done the harder math: the opportunity cost of making a customer wait for a fix to their pain point while you seek the comfort of one more variable.

Every experienced leader carries instincts like this: earned through real consequence, validated across enough decisions that they have compressed into automatic reflexes. That is what makes them valuable, and what makes them dangerous when the context shifts, because the confidence they produce is the same whether they are right or wrong. Every correct thesis eventually becomes the wrong thesis. Principles endure. Vision can be durable. But a thesis about how the world works has a shelf life, and the more consequential the decisions it informed, the harder it is to recognize when that shelf life has expired. An organization that transmits its current leaders’ judgment with high fidelity to the next generation, without contesting the assumptions underneath those instincts, will be perfectly calibrated to the world as it was, not the world as it is. That is not adaptation. It is preservation. Judgment must not just be transmitted but contested. That is what keeps it current.

This was hard before, but it is even harder in the AI era because AI defaults to confirmation, refining the thesis rather than contesting it. Using AI as an adversary requires deliberate intent: under what conditions would this strategy fail? What evidence would tell you this thesis is wrong? What are you not seeing because your pattern recognition was trained on the last decade?

The apprenticeship pipeline feeds back into renewal in an unexpected way. New people, precisely because they lack the established world model, sometimes see things experienced people have filtered out. The senior engineer whose reflex is to throw hardware at a scalability problem does not think to ask whether an AI agent could debug the contention for a fraction of the cost. The senior strategist who has already committed to market expansion does not revisit whether the existing customers’ declining renewal rates should change the thesis. The naive question turns out to matter precisely because the established framework has learned not to ask it. The organization that cuts off the junior pipeline does not just lose future judgment capacity. It loses a source of present challenge.

The default workflow in many AI-assisted organizations has become to ask AI to generate options, then select from what it produces. That workflow feels efficient. It is also a judgment atrophy exercise. The practitioner never forms their own view, never confronts the gap between what they think and what they can articulate, never experiences the productive discomfort of having a position and discovering it is wrong.

The reverse workflow is a judgment development exercise. Form your own position first. Build your own hypothesis. Then bring AI in to challenge, extend, or disconfirm what you have produced. AI dramatically amplifies a practitioner’s ability to test their own thinking, but an amplifier is only useful if there is a signal to amplify. The practitioner who begins with their own view and uses AI to stress-test it develops stronger judgment than the practitioner who begins with AI’s view and selects from its output. The first is building a world model. The second is navigating a collage of other people’s thinking dressed in a convincing narrative. Organizations that embed this discipline into how work is structured are investing in the judgment formation pipeline every time a practitioner sits down to work.

One Infrastructure

Each of these mechanisms addresses a different part of the problem, and the temptation will be to invest in one or two and defer the rest. That produces a system with a predictable imbalance. Visibility without consequential experience produces pattern recognition without conviction. Consequential experience without visibility forces every practitioner to learn from scratch. Both together still fail if the organization’s decision processes have no surface on which judgment can land. Contesting established judgment without the formation pipeline to develop new judgment tears down conviction without building anything to take its place. The four are not independent investments. They are one infrastructure.

Competitive Convergence

Judgment formation is not only an organizational problem. It is a competitive one. Every major competitor in a given market now has access to essentially the same AI capabilities. The frontier models have been trained off the same corpus of data. The integration patterns are standardizing, the tooling is commoditizing, and being an early adopter provided a timing advantage, but that advantage is expiring.

What happens when two competitors rely on those same models to shape strategic decisions? Consider two AI-native competitors who have moved beyond using AI for productivity and now rely on it to evaluate market opportunities, frame product direction, and build the rationale for where to invest. When the model generates three strategic options and the leadership team selects the most compelling one, that selection feels like making a judgment call. But when the three options were drawn from the same training data every competitor’s models draw from and the evaluation criteria are just the ones most prevalent in the corpus, the “judgment” becomes pre-filtered to the most plausible. The “compelling” option is, by construction, simply the most probable one.

One decision shaped this way is a judgment call. Hundreds of them, accumulated across a year of leadership decisions and then the thousands an organization makes beneath them, become something else entirely. The feature prioritization that felt like product instinct, the positioning that felt like market insight, the organizational design that felt like operational wisdom were each individually defensible. Each one shaped, even partially, by models drawing from the same well. The aggregation of those thousands of calls is effectively the organization’s strategy. And if every competitor’s thousands of calls are being mediated by the same models, built on the same data, optimizing for the same patterns of plausibility, those strategies converge. Not because anyone copied anyone else, but because everyone independently selected the most probable option from the same distribution. The unique value proposition is not as unique as it appears from the inside. The competitive position drifts toward the same center as every other competitor, and the convergence is invisible because each individual decision still feels like a choice.

This convergence operates at a level even deeper than organizational decision-making. Leaders making judgment calls on those options are themselves using the same LLMs to process and filter the information that shapes their worldview. The morning coffee articles they scan are summarized by AI before their first meeting. The competitive landscape gets synthesized into an AI-generated briefing. The strategic frameworks get refined through a conversation with a model that draws on the same corpus as every other leader’s model. Each interaction feels like original thinking, because the leader is genuinely engaging, genuinely questioning, genuinely synthesizing. But the raw material being synthesized has already been filtered through the same probability distribution where the rough edges are smoothed out. The result is an echo chamber with the texture of independent thought: a logical narrative voice, data that looks like rigor, a coherent worldview that feels earned. What has actually been happening is convergence toward the same center as every peer who works the same way. It is groupthink that does not feel like groupthink, because it arrives in your own voice.

The answer is not about stopping the use of AI. The productivity gains to process the glut of information are real, and no leader can review it all in depth. The antidote for this convergent thinking is supporting and amplifying the people whose judgment was built through consequence, not through the same corpus the model draws from. People who can feel when the frame is wrong because their world model was constructed in a fundamentally different way. These people create constructive divergence. The experienced leader who senses the framing is wrong is constructive divergence. The junior who asks the naive questions about why the established framework works the way it does is constructive divergence. When an organization removes the conditions under which those people can develop, thrive and be retained, it removes its own capacity for the pattern-breaking that innovation requires. The improbable choice, the one the model would not have selected, is where competitive differentiation begins.

Employee Retention

The conventional framing treats employee retention ultimately as a compensation problem. Offer more, retain more. But that framing fails to account for what the most valuable people actually want. The VP of Engineering did not leave because the compensation was wrong. She left because the system could not use her judgment. She spent a year reviewing proposals that were all competent, all defensible, all roughly interchangeable. The distance between the options was never large enough for her judgment to be the thing that decided it. She did not leave for a competitor who paid more. She left for an environment where her judgment would matter.

The experienced people who carry the deepest judgment want to exercise it. They want to develop the next generation, because mentorship is how judgment stays alive, and people who have spent decades building it care about the legacy they leave behind. The desire to do work that is real is what retains them, and it is the same thing that attracts people still building their judgment who want to learn under real conditions. The formation infrastructure and the conditions for retention are the same investment, and the pipeline feeds itself in both directions. Build it and it attracts both. Fail to build it and the experienced people leave for environments where their judgment matters, the developing people never arrive, and the gap compounds.

The competitive gap that opens is not one that can be closed by deploying the same model a competitor deployed last quarter. It is not a tooling gap. It is a judgment gap. And judgment gaps compound in the same way that judgment itself compounds: slowly, and then all at once when the moment arrives that demands what only judgment can provide.

The Autopilot Precedent

AI is not the first technology to trigger this sequence. There is a historical precedent where judgment eroded under automation, where that erosion was visible long before it became catastrophic, and where the way out was deliberate reinvestment rather than reversal. It played out in commercial aviation when autopilot became widely deployed and where the stakes were measured in lives.

Autopilot reduced cockpit crews from five to two and changed the profession fundamentally. What autopilot enabled, alongside advances in jet engine technology, deregulation, and hub-and-spoke network design, was a massive expansion of the industry. More routes became economically viable. More aircraft flew more frequently. More pilots were needed, not fewer, because the scale of operations grew faster than the crews shrank. Air traffic control developed into a coordination layer that keeps distributed, high-velocity operations pointed in the same direction. Entirely new industries emerged: ground logistics, on-board catering services, tourism at scales that would have been impossible without the automation. The industry did not shrink because of automation. It expanded through it.

The transformative moment came with the Airbus A320 in 1988, the first commercial aircraft with full digital fly-by-wire technology. A new generation of automation fundamentally changed the pilot’s role from being an active operator to more of a system monitor. The automation handled not just cruise but increasingly the complex transitions, the approaches, and the instrument procedures. Manual flying became the exception. Monitoring became the norm.

Within roughly two decades after autopilot was broadly introduced, the consequences were visible to anyone looking. In 1996, an FAA Human Factors Team published a comprehensive study of flight crew interfaces with highly automated cockpits. The report identified the problem directly: automation dependency was creating vulnerabilities in pilots’ ability to manage flight path manually when the automation failed or behaved unexpectedly. It raised an explicit concern that investments in human expertise were being reduced due to economic pressures, precisely when nearly three-quarters of all incidents involved flightcrew errors. The problem was documented. The industry noted it. The lack of urgency had catastrophic consequences.

Thirteen years later, Air France Flight 447 crashed into the Atlantic, killing all 228 people on board. The aircraft’s airspeed sensors iced over at high altitude and the autopilot disconnected, requiring the pilots to fly the aircraft manually. They could not. The training regime had not prepared them for manual flight at high altitude under those conditions. The instincts that would have told them what was happening had atrophied through years of automation-managed flight. The plane was fully functional. The skills to fly it were not. Air France’s own internal assessments had already identified that the airmanship skills of some of its long-haul pilots were weakening. The warning was visible before the crash. The consequences arrived anyway.

Four years later, Asiana Flight 214 crashed on approach to San Francisco in clear weather, killing three and injuring over 180. The airline had explicitly encouraged maximum use of automation. The pilot had never executed this type of approach without automated glideslope guidance in a real Boeing 777. The NTSB concluded the crew had over-relied on automated systems they did not fully understand. The conditions were benign. The automation dependency was not.

An FAA study spanning 46 accidents and major incidents, 734 voluntary reports, and data from over 9,000 observed flights found that in more than 60 percent of accidents, pilots had trouble manually flying the plane or made errors with automated flight controls. The co-chair of an FAA advisory committee on pilot training summarized the finding in five words: “We’re forgetting how to fly.”

The industry’s response was not to remove autopilot. Nobody argued for returning to manual operations. The efficiency gains were real and the expansion they enabled was too valuable to reverse. The response was to deliberately re-engineer training to maintain the judgment that automation was degrading. Beginning formally in 2013 and updated in 2017, the FAA issued guidance addressing the core finding: continuous use of automated systems failed to reinforce the knowledge and skills required for manual flight. Pilots were required to hand-fly regularly, to practice the skills that autopilot had made unnecessary on most flights but essential on the flights that mattered most. Simulator training became more rigorous, not less, precisely because real-world opportunities to exercise manual judgment had decreased. The industry was forced to reinvest in maintaining the very capabilities that the automation had made seem obsolete, because it understood, after the evidence was written in wreckage and loss, that the moments when those capabilities were needed were the moments that determined whether everyone on the aircraft survived.

That cycle is playing out again with AI in place of autopilot, but now compressed from decades into years. The judgment formation pipeline can dry up in a few years. The consequences take longer to surface. By the time they do, the gap between what the organization needs and what it can produce is already measured in generations of people who were never developed. Aviation had the luxury of discovering the atrophy problem gradually, studying it, iterating on training regimes, adjusting after catastrophe. Organizations deploying AI at scale do not have that timeline. That time compression is the danger.


Aviation is the clearest documented example of an industry that confronted this exact automation sequence and emerged stronger, not by removing the automation but by investing deliberately in maintaining the human capabilities it was degrading. Air traffic controllers, aviation safety engineers, simulator designers, route network analysts: none of these roles existed in their current form before autopilot and its companion technologies made the expanded system possible. The automation did not reduce the scope of human contribution. It moved that contribution to where it mattered most and created new domains where human judgment was essential at scales that manual operations could never have reached. The expansion and the judgment investment were the same bet.

The time compression for the AI era also means the opportunity is closer. The organizations that invest in maintaining judgment capacity will not just avoid the atrophy. They will build the capacity to operate at scales and complexities that would have been impossible without both the automation and the humans capable of governing it.

The AI era is not the end of human judgment. It is the beginning of a period where judgment must be deliberately cultivated rather than accidentally produced. The accidental formation pipeline, the one that ran as a byproduct of slow execution, co-located work, and apprenticeship relationships that formed naturally in the course of building things together, is gone. What replaces it must be designed.

Making judgment visible so people still building their judgment can see how experienced people think. Designing roles where real consequence builds the world model that no training program can substitute for. Building decision processes that can receive what experienced judgment knows, even when it resists quantification. And contesting established judgment so it stays current rather than calcifying into the instincts of the last era. The mechanism categories are universal. Every organization will build them differently. The implementation is what makes them yours.

The VP of Engineering who left the building carried something the system could not replace. The architect whose exit conversation called out what nobody else would dare say left a gap no documentation could fill. The junior engineers who reported shipping faster than ever and learning less than they had in years were describing the early symptoms of a judgment formation pipeline that had gone dry.

None of it was anyone’s fault. The mechanisms that used to produce that judgment were never designed. They were inherited. And that inheritance is gone.

The organizations that build the judgment engine will compound their advantage with every person they develop, every cycle of observation and consequence they design, every decision process they redesign to receive what experienced people know. The organizations that do not will converge toward the same plausible center as every competitor drawing from the same well of AI-generated options.

The autopilot handled cruise. The humans handled the moments that mattered most. The question for every organization in the AI era is the same one aviation answered a generation ago: will you invest in maintaining the capacity for the moments that matter, or will you let it atrophy and hope those moments never arrive?

They will arrive. They always do.

Footnotes

  1. The structural case for decision infrastructure, and the three mechanisms that compose it, is developed fully in “Engineering Coherence: Building Decision Architectures for the AI Era”.

  2. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.

  3. Meadows, D. (2008). Thinking in Systems: A Primer. Chelsea Green Publishing.


This is the kind of work I help leaders through directly. If what I've described is what's happening in your organization, reach out and we can talk it through.

Work with me →