The model demo is not the strategy
An AI demo can produce an impressive memo, table, or piece of code in minutes. That does not mean an organization will make a good decision with it. The real questions are more operational: Which workflow will absorb the output? Who will review it? What evidence is required before it is used? Who can stop the process when it fails?
AI strategy is therefore not just a model-selection exercise. A model sits inside a larger system that includes the task definition, data, decision rights, incentives, controls, appeals, and feedback.
That is the core error in the Frankenstein image. Assembling the parts is easy. The hard part is deciding when the resulting system may act, when it must stop, and who remains accountable for what it does.
Start with the workflow, not the tool
Before comparing models, a company should answer five questions:
- Which decision or customer outcome are we trying to improve?
- What does it cost to make that decision incorrectly?
- On which tasks has the model been tested reliably?
- What exactly will the human review, and when?
- What happens when the system produces an unexpected result?
Choosing a larger, faster, or cheaper model while these questions remain unanswered can enlarge the problem instead of solving it. Technology does not turn an ambiguous task into a well-defined one by itself.
A 2026 study in Organization Science makes the boundary visible. The preregistered experiment involved 758 Boston Consulting Group consultants performing 18 realistic knowledge tasks. On tasks within the observed AI capability frontier, AI users completed 12.2% more tasks, finished them 25.1% faster on average, and produced higher-quality solutions. On one complex managerial task selected outside that frontier, AI users were 19% less likely to produce a correct solution. The sample came from one consulting firm and largely involved early-career employees, so the rates should not be generalized to every job or organization.
The result does not say that AI is good or bad. It says something more useful: the same tool can behave differently when the task changes. Evaluation should therefore begin with the task boundary, the cost of error, and the control point in the workflow, not with a general claim about the model's intelligence.
Who owns the decision?
Saying that a system has human oversight is not enough. What does the human review? Is there enough time? Does the reviewer have the domain knowledge to challenge the output? Can the reviewer stop the system? Will a performance target or manager penalize that person for stopping it?
The decision chain should be separated into at least these rights:
- Recommend: What option, draft, or action does the system suggest?
- Decide: Who makes the final decision?
- Approve: Who releases the result to a customer, colleague, or public audience?
- Challenge: Who can contest the result, and with what evidence?
- Stop: Which conditions halt the workflow automatically?
- Publish or execute: Who moves the output into the world?
- Retire: Who shuts the system down when the case for it disappears?
The framework by Shrestha, Ben-Menahem, and von Krogh in California Management Review shows that there is no single way to combine human and AI decisions. Full delegation, AI filtering followed by human judgment, human selection followed by AI evaluation, and aggregated human-AI decisions are different structures. The right choice depends on the size of the search space, interpretability, the need for speed, and the degree of repeatability.
This distinction is more valuable than simply saying that a company will use AI. It designs an authority architecture, not just an adoption plan.
A control gate needs an evidence threshold
A control gate is not a checkbox. It is a decision about what evidence is required before an output can be used.
NIST's voluntary AI Risk Management Framework approaches this through four functions: Govern, Map, Measure, and Manage. In plain language, an organization sets governance, maps the use and its risks, measures the system, and manages what it finds. NIST's 2024 Generative AI Profile connects that approach to governance, content provenance, pre-deployment testing, and incident disclosure.
ISO/IEC 42001:2023 also addresses the management system around AI rather than a model in isolation. ISO describes requirements for establishing, implementing, maintaining, and continually improving an AI Management System. Whatever a company calls its internal process, the point is the same: the problem does not end at model parameters.
The OECD AI Principles, updated in 2024, emphasize human agency and oversight, transparency, safety and security, accountability, and traceability. The OECD's Responsible AI Due Diligence Guidance, published on February 19, 2026, gives enterprises six steps: embed responsible conduct in management systems, identify and assess impacts, prevent or mitigate them, track implementation and results, communicate actions, and provide for remediation when appropriate. The guidance is voluntary, but it offers a useful lifecycle for thinking about AI from purchase through retirement.
Law is moving in the same direction, but not on one universal calendar. Under the EU AI Act, implementation, supervision, and enforcement responsibilities sit with the AI Office and Member State authorities from August 2, 2026. Selected high-risk areas, including employment and critical infrastructure, are scheduled to apply from December 2, 2027, while high-risk AI embedded in regulated products is scheduled for August 2, 2028. Certain Article 50 transparency obligations begin on August 2, 2026, with a limited December 2, 2026 grace period for some marking obligations concerning systems already on the market.
The lesson is not that every company must install the same control on the same day. It is that a claim of compliance is empty until the use case, jurisdiction, sector, and risk level have been classified.
Incentives quietly rewrite the system
Even a well-designed formal process can be replaced by a different workflow in daily practice. If a team is measured only on speed and volume, shortening quality review, hiding exceptions, or accepting an AI recommendation without challenge may look rational. When speed and volume are the only measures, these behaviors become a predictable risk of the system.
AI adoption is therefore not only a training problem. It matters what the employee signs, which error the employee is allowed to report, and what happens after a challenge. A good system rewards making weak output visible, not only producing the right output quickly.
The Quarterly Journal of Economics field study by Brynjolfsson, Li, and Raymond offers a useful, bounded piece of evidence. Across 5,172 customer-support agents, access to an AI assistant increased customer issues resolved per hour by 15% on average. The effects varied across workers. The tool augmented agents, who remained responsible for the conversation and could ignore or edit its suggestions. This is a result from one company's deployment, not a productivity promise for every organization.
Raisch and Krakowski's management argument matters here as well: automation and augmentation are not simple substitutes. Each changes the conditions for the other. An exclusive focus on automation can weaken human knowledge and challenge capacity. An exclusive focus on human review can constrain speed and scale unnecessarily.
The human accountability chain must be visible
In the small operating example I use on my writing site, efetanyer.com, AI is not the final owner of a research-to-publication decision. The workflow separates question framing, source mapping, claim checking, drafting, language and presentation review, preview, and explicit approval. Moving material to the site or publishing it is not treated as an automatic consequence of a model output.
This is a one-person operating example, not evidence of enterprise transformation. I am not disclosing private prompts, personal records, client information, or unpublished research. The value of the example is not that a model can write. It is that decision rights and evidence gates can be made visible even in a small system.
That does not create a universal rule that AI must never publish. In a low-risk, reversible content workflow, the organization may choose a different design. It should still be clear who made the final decision, what evidence was retained, and who owns the correction when something goes wrong.
The strongest counterargument
It would be another mistake to explain every failed deployment as a governance failure. At least six alternatives should be tested:
- Model capability: The model may simply be unable to perform the task reliably. Use an evaluation set that includes both ordinary and out-of-distribution cases.
- Data and context: The data may be stale, incomplete, unrepresentative, or unauthorized. Audit lineage, freshness, coverage, and permissions first.
- Infrastructure: Latency, outages, weak retrieval, or a poor interface can erase theoretical model gains. Measure end-to-end failure rates.
- Economics: Licensing, compute, review time, and remediation costs may exceed the value created. Compare the full cost with the existing process, not just the model price.
- Skills and trust: Users may under-trust or over-trust the tool. Track training, challenge quality, error discovery, and concentration of use.
- Regulation and rights: A use may be restricted in a jurisdiction or sector, or require transparency, impact assessment, or recourse. Classify the legal scope before choosing a deployment pattern.
If these tests are reasonably satisfied and it is still unclear who decides, what evidence is enough, and which incentives shape behavior, the operating model becomes a strong candidate explanation. That is an inference that must be tested in each case.
Three operating scenarios
Assistant scenario: For a low-risk, reversible task, AI prepares a first draft or a set of options. A named human edits, checks the source, and approves the output. This is not enough when the cost of error is material.
Filter scenario: AI narrows a large set of options. A domain owner reviews edge cases and makes the final decision. This can work when human review of the full set is expensive but the final decision requires context.
Constrained automation scenario: A high-volume, low-risk action runs automatically with rollback, sampled review, monitoring, and a stop condition. This pattern cannot be carried automatically into a high-risk domain. The organization must decide in advance who can stop the system and how affected people receive a correction.
Practical operating-model checklist
- What decision or outcome are we improving? What is explicitly out of scope?
- Who owns the business outcome, the system, the data, and the risk decision?
- Who may recommend, decide, approve, challenge, stop, execute, and retire?
- What accuracy, quality, fairness, privacy, security, latency, and cost thresholds apply before use?
- Which tasks have been tested? Which similar-looking tasks sit outside the validated boundary?
- Does the human reviewer have enough time, authority, context, and domain knowledge to intervene?
- Can the organization trace the model, data, sources, version, material change, reviewer, and final decision?
- Do speed and volume targets reward people for skipping controls?
- Which error, challenge, complaint, incident, or data shift triggers a review?
- Who decides on correction, rollback, and retirement?
This is an editorial synthesis, not a certification or legal safe harbor. Its purpose is to make the decision architecture visible before a model is selected or scaled.
Conclusion
Changing the model is sometimes the right answer. But the strategy question often comes first: What work are we doing, what risk are we accepting, what evidence will justify scale, and what condition will make us stop?
Durable value depends less on acquiring a powerful tool than on tying it to a defined decision, assigning accountability, and making failure reversible. My judgment is simple: technology can be bought; value emerges only when authority, evidence, and correction are designed around its use.
Sources
- NIST AI Risk Management Framework
- NIST AI RMF Generative AI Profile
- ISO/IEC 42001:2023
- OECD AI Principles
- OECD Due Diligence Guidance for Responsible AI
- UNESCO Recommendation on the Ethics of Artificial Intelligence
- European Commission AI Act overview
- European Commission Article 50 FAQ
- EUR-Lex Regulation (EU) 2024/1689
- Dell'Acqua et al., Navigating the Jagged Technological Frontier, Organization Science
- Brynjolfsson, Li and Raymond, Generative AI at Work, The Quarterly Journal of Economics
- Shrestha, Ben-Menahem and von Krogh, Organizational Decision-Making Structures in the Age of Artificial Intelligence
- Raisch and Krakowski, The Automation-Augmentation Paradox
- Green, The Flaws of Policies Requiring Human Oversight of Government Algorithms
- Ghasemaghaei and Kordzadeh, Understanding how algorithmic injustice leads to making discriminatory decisions





