For sixty years, software has chased one goal: reliably delivering products that customers actually want. The history of our methodologies is a history of getting closer — and still falling short.
Waterfall gave us structure, but it was blind to the customer until it was too late to change course. Agile gave us speed and learning cycles, but a fast team can build the wrong thing more efficiently than a slow one ever could. Design Thinking finally centered the customer’s problem — and paired with Agile, it was supposed to be the answer. Yet rigorous customer research stayed slow, inconsistent, and dependent on the rare skill of the individual practitioner.
Each method fixed the last one’s flaw and left a new one. But underneath all of them sat the same unsolved bottleneck: understanding the customer accurately, fast enough to act on, and consistently enough to trust.
There is a deeper reason this bottleneck never broke. To see it, we have to borrow a lens from systems thinking.
Why we kept failing: the wrong tools for the domain
The Cynefin framework, developed by Dave Snowden, sorts problems into domains by the nature of their cause and effect. In the Complicated domain, cause and effect exist but require expertise to uncover — you analyze, you bring in experts, you find the good practice. In the Complex domain, cause and effect are only visible in hindsight. You cannot analyze your way to the answer in advance. You can only probe, sense, and respond — run small experiments, watch what emerges, and amplify what works.
Customer needs and markets live in the Complex domain.
This is the quiet error at the heart of sixty years of methodology: we kept attacking a Complex-domain problem with Complicated-domain tools. More analysis. More expert intuition. More upfront planning. More research rigor. Every one of those is the right response to a Complicated problem — and the wrong response to a Complex one. You cannot out-analyze a system whose future only becomes clear after it has already happened.
What the Complex domain actually demands is continuous sensing — staying close to the system as it shifts, reading the weak signals at its edges, and responding to what emerges rather than to what you planned. The trouble is that continuous sensing at any real scale was, until very recently, humanly impossible. So we did the next best thing: we sent in our most gifted individuals — the Christensens, the rare product visionaries — to sense the system by intuition. Sometimes it was brilliant. Mostly it was unrepeatable.
Foresight, in other words, has always been an art. The thesis of this work is that agentic AI is the first technology capable of turning it into a discipline.
The foundation: a foresight process we already knew
Here is the part that should make us humble. The intellectual machinery for anticipating customer needs already exists. It has for decades. We have simply never been able to run it consistently.
Four ideas, taken together, describe a complete foresight process:
Clayton Christensen showed that customers struggle with durable jobs to be done — and that they are far better at describing their struggles than at naming the solutions. The job is more stable than any stated preference. This is why anticipation is even possible: the underlying need persists even when no one can yet articulate what would satisfy it.
Anthony Ulwick sharpened this into its most useful form: outcomes are stable while solutions churn. People wanted to “keep their finances in order” long before ledgers, spreadsheets, apps, or AI advisors — and they will want it long after. This is the keystone of anticipation. We do not predict features, which is impossible. We anchor on durable outcomes, which hold still, and ask how they will be served next.
Everett Rogers gave us diffusion: innovation adoption follows a structured curve, and the leading edge precedes the mainstream. The future is already present, just unevenly distributed. The signals of what is coming are detectable at the fringe — among innovators and early adopters — before they are obvious in the center.
Simon Wardley gave us evolution: capabilities move along a readable path, from genesis to custom-built to product to commodity. You can map where a component sits today and reason about where it is heading. When a component commoditizes, new higher-order outcomes suddenly become addressable.
Put them in sequence and you have a method: find the durable job (Christensen), anchor on its stable outcome (Ulwick), sense where it is being served at the edges (Rogers), and reason about how the solution space will evolve next (Wardley).
This is a genuine foresight process. And it has always required a rare expert to hold all four lenses at once, continuously, across an overwhelming volume of signal. That is exactly why anticipation has stayed a rare personal skill: extraordinary when one gifted mind does it, and unreliable everywhere else.
What is actually new
Notice what this thesis does not claim. It does not claim a new theory of innovation. Christensen, Ulwick, Rogers, and Wardley did that work. It does not claim that AI can predict the future — systems thinking is explicit that the Complex domain forbids it. A market can always surprise you, and any honest method must stay humble before that fact.
The claim is narrower, and stronger for it.
These four frameworks describe a process that was always meant to run continuously — and never could, because no human can sense a complex system at scale without rest, without bias, and without losing the thread. Agentic AI is the first technology that can run this existing process: continuously, at scale, and consistently. It can sense weak signals across far more sources than any analyst. It can hold all four lenses at once without fatigue. It can reason about plausible solution-evolutions and surface them as structured, testable hypotheses — not answers, not predictions, but well-formed bets that a team then validates with real customers.
That is the shift. Not from human foresight to machine prediction, but from foresight as a rare personal skill to foresight as a working system. From a craft practiced brilliantly by a few, to a discipline that can be run reliably by many.
We are not replacing the intuition of the great product minds. We are, for the first time, building a system that can do what they do — sense the Complex domain continuously — and make it repeatable.
The conceptual model: what anticipatory discovery looks like in practice
It is one thing to argue that agentic AI could systematize foresight. It is another to show what that system actually does. The model that follows is not speculative. It emerged from designing and building a working platform that carries customer understanding from discovery through to delivery — and the design decisions, including the ones that resisted easy answers, turned out to be the most revealing part. Each principle below is reported as a finding, because that is what it is.
Finding one: the missing layer is where solutions are re-anchored to outcomes
Most product processes move from what to build (a prioritized list of features) directly to how to build it (specifications, epics, tasks). There is a silent gap in that handoff. The moment features are decomposed into delivery units, the connection to why — the underserved customer outcome that justified them — tends to drop away. Teams end up building a correct backlog for a purpose no one is still looking at.
The first design finding was that a discovery-to-delivery process needs an explicit layer whose only job is to hold that connection: a place where solutions are re-anchored to the outcomes they serve, before they fragment into build specifications. We can call it the value flow. It does not ask “what order do we build these in” (that is project management) or “how does the user move through them” (that is experience design). It asks: how do these features chain together to deliver a durable customer outcome?
This matters to the larger argument because it locates, concretely, where the golden thread is most likely to snap — and therefore where a system must work hardest to keep insight present. Anticipation is not only a front-of-process activity. It has to be defended in the middle, at the seam between deciding and building.
Finding two: value accumulates along a chain, anchored to a stable outcome
The next question was structural. Do features serve an outcome as a bundle — a set of capabilities that together satisfy it — or as a sequence, where each builds on the last and value accumulates step by step?
The answer that held up was a blend, with sequence as the spine. Features chain into a directed path; value accumulates as the chain progresses; and the whole chain is anchored to the outcome it serves. The outcome is the fixed point — Ulwick’s stable target — while the chain of solutions is what moves.
This is more than an implementation detail. It encodes a dynamic claim that static capability-mapping cannot: that an outcome is served progressively, and that anticipation is therefore about identifying the sequence by which an emerging outcome gets served as solutions evolve — not merely cataloguing the capabilities it will eventually require. This is exactly where Wardley’s evolution lens earns its place: a chain is a trajectory, and trajectories can be reasoned about.
Finding three: insight can be reconstructed, but it cannot be skipped
Then came the finding that did the most to sharpen the thesis, and it arrived as an objection. The platform allows a user to enter not from discovery but from a finished feature list — to skip the customer research entirely. Such a user has no outcomes. If the value flow must anchor on outcomes, what happens when there are none?
The tempting answer is to make the outcome anchor optional. That answer is wrong, and seeing why is the point. Drop the anchor and the value flow degrades into a plain dependency diagram — precisely the build-order artifact we had ruled out. The outcome is not decoration; it is the load-bearing element.
The resolution was to treat the absence of outcomes not as a gap to route around but as a trigger for reverse discovery: the system infers, from the features themselves, what underserved outcomes they appear to serve, and presents those as hypotheses for the user to confirm or correct. The user who skipped insight does not get to proceed without it. They get it reconstructed — and, crucially, labeled honestly as inferred rather than discovered, weaker evidence that ought to be validated against real customers.
The general principle this reveals is sharper than the tagline that inspired it. There is no foresight without insight — but the deeper finding is that insight can be reconstructed backward from solutions, and yet it cannot be skipped. A system can recover the missing why; it cannot proceed coherently without one. Anticipation has a non-negotiable precondition, even when that precondition has to be supplied after the fact.
Finding four: the division of labor is dictated by the domain, not chosen for comfort
The final and most consequential finding concerned the relationship between the human and the AI inside the flow. The AI proposes — candidate chains, a first-pass sequence, the inferred outcomes, the trade-offs it can see. The human commits — the actual sequence, the priorities, the resolution of competing chains. So far this is unremarkable; many tools claim some version of “AI suggests, human decides.”
The revealing question was narrower: when the human reorders the flow, what is the AI permitted to do? Only flag mechanical problems, like a linter? Or argue back with reasoning, like a partner?
Neither extreme survives contact with the thesis. An AI that only flags mechanical issues underclaims what we have argued AI can do — it reduces a foresight engine to a syntax checker. An AI that argues with every human judgment overclaims — it asserts an authority the Complex domain does not grant to anyone, human or machine. The answer that held was tiered: mechanical issues — dependency conflicts, cycles, structurally broken sequences — are flagged always, because they have knowable, analyzable answers. Reasoned judgment about which sequence best delivers value is offered only when the human asks for it.
What makes this more than a usability compromise is where the two tiers come from. They are not arbitrary product settings. They map exactly onto the two domains the thesis already rests upon. Mechanical validation lives in Cynefin’s Complicated domain — cause and effect are knowable, so the AI acts with authority and resolves the issue definitively. Reasoned sequencing lives in the Complex domain — there is no analyzable right answer, so the AI advises and the human commits. The interaction design was not imposed on the theory; it was predicted by it.
This is the strongest evidence the thesis has that its foundation is load-bearing rather than decorative. A framework that merely decorates an argument can be swapped out without consequence. A framework that predicts how the tool must behave — and would make it incoherent to behave otherwise — is doing real work. The Cynefin distinction told us, before we had built anything, where the machine may decide and where it must defer. That the build then confirmed it is the whole point of grounding a thesis in practice rather than assertion.
What the model adds up to
Put the four findings together and the conceptual model is complete. In practice, anticipatory discovery is a process that does four things. It keeps an explicit layer where solutions are re-anchored to outcomes. It sequences solutions into value chains built on stable outcomes. It treats insight as a precondition — reconstructing it when skipped, but never skipping it. And it divides the work between machine and human along one line: the knowable versus the genuinely uncertain.
None of these is a claim about predicting the future. Each is a claim about how to structure the present so that foresight becomes a discipline rather than a gift. And each emerged not from theory alone but from the friction of building the thing — which is the only kind of evidence that earns the word finding.
The honest limits: what building it taught me about what it cannot yet do
A thesis grounded in practice owes its reader the failures as well as the findings. Three limits emerged from the work — and the most important one emerged only when I turned the thesis’s own skepticism against the platform that embodies it.
The system is episodic; the domain demands continuous
The Complex domain, as this paper has argued from the start, demands continuous sensing — an ongoing probe-sense-respond posture, because the system shifts and its cause-and-effect is only ever visible in hindsight. Hold the working platform up against that standard honestly, and a gap appears.
What the platform clearly achieves is making discovery systematic. The Christensen-to-Ulwick chain — durable jobs, underserved outcomes, cross-persona synthesis — runs as a repeatable pipeline. It finishes in hours rather than weeks, with a consistency no individual can sustain. For the discovery half of the process, foresight as a discipline rather than an art has genuinely been demonstrated.
But it is demonstrated episodically. A team runs a project: discover, analyze, plan, build. However rigorous, that is a snapshot. As this paper was being finalized, the first sensing capability shipped: on-demand edge research that runs the Rogers lens (how is this problem being solved at the edges right now?) and the Wardley lens (where is the underlying capability in its evolution?) with live web grounding — each signal carrying its source and a timestamp, badged as a third evidence class, sensed, alongside discovered and inferred. Its first live run told the author something he did not know about his own market — the confrontation the practice section describes, performed on the person who wrote it. But it runs when a human clicks. Nothing in the current system watches the market between sessions. Nothing re-reads the edges on its own. Nothing raises its hand when the ground has shifted underneath a roadmap that was sound on the day it was written.
This is the difference between compressing discovery and making it continuous. They are related achievements, but not the same one. The honest statement of what has been shown is this: the foresight process that lived for decades in rare heads can now run as software, consistently, for anyone — and its sensing lenses now run live, on demand. The remaining frontier is making them run always. That means an agentic sensing loop — one that watches edge adoption and solution drift against anchored outcomes between sessions, and surfaces emerging patterns as testable hypotheses without being asked. That loop does not yet exist. It is named here as a limit precisely so it cannot be mistaken for a claim. But it is also the next question this work must answer — and the platform was shaped to ask it.
The conceptual model constrains the process, not the product
The second limit is subtler, and it cost real effort to learn. Finding One argues that a discovery-to-delivery process needs an explicit layer where solutions are re-anchored to outcomes. I built that layer as a product module — a visual value-flow canvas between planning and execution. It worked. And then, looking at it honestly, I parked it: as a standalone screen it added little that the platform’s existing evidence-thread — outcomes flowing into features flowing into epics — did not already carry.
The lesson draws a boundary around the model’s authority. The model tells you what a process must do. It does not tell you what a tool must show. The re-anchoring function was real and necessary — but this platform was already performing it quietly, through its data spine. Turning it into a dedicated screen added a module without adding a decision. Practitioners applying this model should hold the same discipline: implement the functions, and let each screen earn its place by the decisions it enables — not by its loyalty to the theory. A model that predicted its own module was unnecessary, and was believed, is worth more than a model defended past the evidence.
The acceleration compounds locally and erodes globally
The third limit concerns the AI-augmented building itself. The platform was built at a pace no conventional team could match — stage by stage, each stage verified, each locally correct. A disciplined review at the end told the other half of the story: the same acceleration that made each stage fast had quietly produced global duplication. The same save pattern reimplemented in two dozen files. The same endpoint skeleton repeated six times. Two registries describing the same modules, hand-synchronized. Every piece was right; the whole was quietly piling up debt.
This is not a side complaint about engineering — it is a limit of the way of working this paper advocates. AI acceleration improves each piece by default, not the whole. Nothing in the loop is responsible for noticing repetition, because each generation solves the problem in front of it, correctly, in isolation. The judgment that says these six things are really one thing stayed entirely human. And it had to be deliberately scheduled, because no failing test ever demands it. Teams adopting AI-accelerated delivery should treat periodic human-led consolidation not as cleanup but as a standing role: the system will not ask for it, and the debt is invisible until someone looks.
What the limits share
All three limits have the same shape, and it is the shape of the thesis itself. In each case, the machine did the knowable part: running the pipeline, drawing the module, generating the stages. And in each case, the genuinely uncertain judgment stayed human — seeing that episodic is not continuous, that a working module is not the same as a needed one, that individually correct pieces were adding up to duplication. The boundary between the Complicated and the Complex did not just predict the tool’s interaction design, as Finding Four showed. It predicted where the tool would stop. And it predicted where the human must stay.
What this means for practice
A conceptual model earns its keep on Monday morning. Three questions decide whether anticipatory discovery becomes practice or stays a paper: where a team starts, who runs it, and how it will most likely be gotten wrong. Years of coaching Agile transformations have taught me that the third question quietly answers the first — so I will take them together.
Start with the gap, not the lecture
The textbook entry point is clean: begin with discovery, state customer jobs as hypotheses, validate before building. Some teams will do exactly that, and the process serves them directly.
But that is not the team you will usually meet. The team you will usually meet is confident. They carry deep industry knowledge and years of customer proximity, and they will tell you — sincerely, and often with some justification — that they already know the problems and the features that address them. The worst move available is to argue with them. You are asking experts to concede that their expertise is insufficient, and no amount of methodology evangelism wins that argument. I have watched the frontal assault fail too many times to recommend it.
The move that works is to stop arguing and run the test. Let the team produce their feature list exactly as they would have — their expertise, fully expressed, nothing withheld. In parallel, run the disciplined discovery process: personas, jobs, underserved outcomes, the full evidence chain. Then sit down together and perform a gap analysis between the two. Where the team’s instincts converge with the evidence, their expertise is confirmed — and they now hold proof rather than assumption. Where the two diverge, the confrontation is not between the team and a consultant; it is between the team and evidence they watched being produced. Nobody loses face. The gap does the teaching.
This is not a workaround bolted onto the model — it is built into it. Finding Three established that a features-first entry must trigger reverse discovery: the system infers the outcomes a feature list presumes, and labels them as hypotheses. The gap analysis is that finding, turned into an adoption pattern. The team’s features imply a set of assumed outcomes; disciplined discovery produces the evidenced ones; the delta between them is the deliverable. The same architecture that keeps the golden thread intact for a rigorous team becomes the challenge mechanism for an overconfident one.
Note what this pattern respects: expert intuition is not dismissed — it is tested. Deep customer knowledge is real, and sometimes the gap analysis will vindicate it completely. That is a success, not a failure of the method. What the pattern refuses is the untested shortcut — expertise, a Complicated-domain asset, applied unexamined to a Complex-domain question. The gap analysis is the probe the Complex domain demands, run at the cost of one parallel discovery cycle.
Who has the standing to run it
This practice needs sponsors with a particular kind of freedom: the organizational standing to let a team’s assumptions be tested without it reading as an attack. In my experience that circle includes the Chief Product Officer, senior product management, and — less obviously but just as importantly — solution and technology architects. The architects matter because they sit at the seam this paper has been describing: the point where decided features become built systems, where the golden thread is most likely to snap. They feel the cost of building the wrong thing earlier than anyone, and they often hold cross-team credibility that a single product owner does not.
What unites these roles is not seniority for its own sake — it is air cover. The gap analysis only functions as a teaching instrument if a candid divergence between instinct and evidence is treated as a discovery, not a performance failure. Someone must hold that space open. That is a leadership act, and it cannot be delegated to the tooling.
The failure mode is the one we started with
Asked how organizations will most likely get this wrong, I find myself giving the same answer as the entry-point question — and that repetition is the point. The dominant failure mode is untested confidence: the belief that industry knowledge can substitute for discovery. It is now turbo-charged by a second temptation this paper has warned against throughout — treating AI-generated analysis as answers rather than hypotheses. A team that skips discovery because it trusts its gut, and a team that skips validation because it trusts the machine, are making the same error in different costumes. Both mistake confidence for evidence.
The entry pattern and the failure mode are therefore one design: start with the gap analysis because untested confidence is the failure mode. The practice does not assume teams will humbly follow a discovery process — it assumes they won’t, and builds the challenge into the workflow itself. That is what it means to design practice for the Complex domain: not a process that works when people behave ideally, but one that produces evidence even when they don’t.
References
Beck, K., Beedle, M., van Bennekum, A., Cockburn, A., Cunningham, W., Fowler, M., et al. (2001). Manifesto for Agile Software Development. agilemanifesto.org
Christensen, C. M. (1997). The Innovator’s Dilemma: When New Technologies Cause Great Firms to Fail. Harvard Business School Press.
Christensen, C. M., Hall, T., Dillon, K., & Duncan, D. S. (2016). Know your customers’ “jobs to be done.” Harvard Business Review, 94(9), 54–62.
Christensen, C. M., Hall, T., Dillon, K., & Duncan, D. S. (2016). Competing Against Luck: The Story of Innovation and Customer Choice. HarperBusiness.
Rogers, E. M. (2003). Diffusion of Innovations (5th ed.). Free Press. (Original work published 1962.)
Snowden, D. J., & Boone, M. E. (2007). A leader’s framework for decision making. Harvard Business Review, 85(11), 68–76.
Ulwick, A. W. (2005). What Customers Want: Using Outcome-Driven Innovation to Create Breakthrough Products and Services. McGraw-Hill.
Ulwick, A. W. (2016). Jobs to Be Done: Theory to Practice. Idea Bite Press.
Wardley, S. (2016–2020). Wardley Maps: Topographical Intelligence in Business. Published serially; available at learnwardleymapping.com and medium.com/wardleymaps.