Iryna Tkachuk, Enterprise AI Advisor at AdvantageWorks Iryna Tkachuk 16 min read AI-assisted

Stop Writing Better Prompts and Start Killing Worse Ideas

A cork board of pinned paper idea cards, five crossed out in black ink and one circled in red with an arrow drawn to it

The fourth revision of the prompt is better than the third. The role framing is tighter, the output schema is explicit, the few-shot examples have been swapped for ones that actually match the edge cases. Somebody spent a good afternoon on it.

Nobody in that thread has asked whether the thing being prompted should exist.

Of the AI ideas that reach us, roughly nine in ten should never have been built. Not because they were badly executed. Because nobody applied a test that would have taken an hour, and the cheapest way to avoid that test is to spend the afternoon on the prompt instead.

That is the shape of most AI failure inside companies right now. Not a model that underperforms, not a data pipeline that breaks, not a security review that drags. An idea that was never examined, wrapped in enough craft to look like progress, moving steadily toward a build that will not change anything. The published failure rates for AI projects run somewhere between 70% and 95% depending on who is counting and what they count, and the interesting part is not the number. It is where in the timeline the failure actually happened.

If you own a roadmap or a budget, the expensive part of a bad AI idea was never the API bill. It was the quarter your team spent on it, and the three better ideas that sat in the queue while they did.

The prompt is the easiest thing to improve, so that is what gets improved

Prompt work has a property almost nothing else in this process has. It gives you feedback in seconds, it is visibly skilled, and it never requires telling anyone their idea is not worth doing.

Every other move at that stage is slower and more political. Checking whether the workflow has an owner means finding the owner. Establishing a before-state means somebody has to measure something nobody currently measures . Asking whether the business has already decided to tolerate this problem means asking a question with an uncomfortable answer. Rewriting the system prompt means opening a file.

So the file gets opened. The team iterates, the outputs improve, the demo gets better, and everyone involved can point at real work. The work is real. It is just not the work that decides the outcome.

The mechanism here is not laziness. It is that specification feels like commitment without being one. A detailed prompt reads as a considered decision about what the system should do, when what it actually encodes is an assumption about what the system should be for. Refine that assumption for long enough and it stops looking like an assumption.

A better prompt on a bad idea gets you a bad answer with more confidence and a nicer schema.

What prompt craft can and cannot fix

None of this makes prompt engineering unserious. Anthropic, OpenAI, Google, and Microsoft all publish good guidance on it, and the techniques work. Clear instruction, structured output, worked examples, and decomposition measurably change what a model produces.

What they cannot change is whether producing it matters. That question is upstream, it is answered once, and almost nobody schedules a meeting for it.

What "90% of AI projects fail" actually measures

Four widely quoted failure statistics are measuring four different things, and none of them means what the headline says.

The most-cited recent number comes from MIT's NANDA initiative, whose 2025 report on the state of AI in business found that roughly 95% of enterprise generative AI pilots produced no measurable return in the profit and loss statement. Gartner has published a separate forecast that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, weak risk controls, escalating costs, and unclear business value. RAND's 2024 analysis of AI project failure put the rate above 80%, roughly double the failure rate of comparable non-AI IT projects. And the figure that shows up most often in blog posts, the claim that 85% of machine learning projects fail, traces to a Gartner (2018) forecast that 85% of AI projects would deliver erroneous outcomes through 2022 because of bias, which is a claim about error rather than failure.

Why the numbers disagree

They are not measuring the same thing.

One study counts pilots, another counts projects, a third counts products that reached users. One defines failure as "no measurable P&L impact," another as "abandoned before production," another as "did not meet its stated objective." A pilot that was always meant to be a learning exercise counts as a failure under the first definition and a success under the third. A project that shipped on time and changed nothing counts as a success under the second definition and a failure under the first.

Treat the range as the real finding. Somewhere between two thirds and nineteen twentieths of enterprise AI work does not produce what it was funded to produce, under most reasonable definitions, across multiple independent measurements.

The more useful observation sits underneath the percentages, and it is the reason this article exists. When these studies describe where the work broke down, they overwhelmingly describe conditions that were already true before the project started. The data was not available. The workflow had no owner. The business case did not survive contact with finance . Those are not things that went wrong during delivery. They are things that were already wrong at selection and stayed wrong.

Failure looks like an execution problem because that is where it becomes visible

Organizations are heavily instrumented for the decision to stop and almost entirely uninstrumented for the decision to start. Every cause in the postmortem literature sits downstream of that one gap.

Read enough of those postmortems and the causes converge: data readiness, executive sponsorship, no production pathway , treating AI as an IT project rather than a business change. Each of those is real. Each is also a symptom with an earlier parent.

Walk one backward. A stalled pilot gets reviewed in month five. The review finds that the data needed to make the feature useful lives in a system the team does not have access to, and that the access request is sitting with a team that has no reason to prioritize it. That gets written down as a data readiness failure.

Back up one step. The access problem was knowable in week one. Nobody checked, because checking was not anybody's job at that point.

Back up again. The idea entered the backlog after a demo that impressed a senior person. Between that demo and the first sprint, no one was accountable for deciding whether to start. There was a process for stopping the work, several in fact, and every one of them required somebody to volunteer a negative opinion about a project that already had a sponsor.

So the failure that originated at selection gets discovered at delivery, and the postmortem names the place where it surfaced instead of the place where it happened. Which raises the practical question: what does an idea that will fail look like on the day it enters the backlog?

The four shapes of an idea that will not survive

Bad AI ideas are not random. They arrive in four recognizable forms, each with a one-line tell, and once you can name them you will find several sitting in your own backlog right now.

Four paper idea cards laid in a row overhead on a wood desk, each with a differently coloured tape tab and a different layout

The demo with no owner

Somebody built something impressive over a weekend. It is impressive. Ask who would use it on a Tuesday and the answer is a department rather than a person, or a persona rather than a name.

The tell: no individual's week gets worse if this never ships.

The problem the business already decided to tolerate

The pain is real and everyone agrees it is real. It has also been real for six years, nobody is funded to fix it, and the workarounds are mature. An AI solution to a tolerated problem still has to displace a workaround people are comfortable with and a budget line that does not exist.

The tell: there was no attempt to fix this before AI made it feel newly fixable.

The task where being wrong is expensive and review is not free

The model gets it right most of the time. The cases where it is wrong carry real cost, so a human checks the output. Now measure the checking. If verifying the answer takes a meaningful fraction of the time it would have taken to produce the answer, the economics break at the review step, not the model step, and no amount of accuracy improvement fixes that.

The tell: the business case assumes review is instant and free.

The workflow that is already broken

The process has undocumented exceptions, three people who know the real rules, and a handoff that only works because someone chases it. Automating it does not remove the mess. It makes the mess run faster, at higher volume, with less visibility into where it went wrong.

The tell: nobody can describe the current process end to end without saying "it depends."

The triage that takes an hour and saves a quarter

Eight questions, one room, one hour. The filter is short on purpose, because a filter that takes a week is a project, and a filter that becomes a project does not get run.

Macro close-up of a printed checklist pinned to cork, with a pencil tick beside one line and an ink cross beside another

Each question is answerable with evidence rather than opinion. The point is not to score the idea. The point is to find out whether anyone can answer at all.

Question

What a good answer looks like

What a red answer means

Who owns this workflow?

A named person whose week measurably improves

If it is a department, nobody will adopt it

What number moves?

An existing metric with a current value

If the metric has to be invented, there is no before-state

What is the current value of that number?

Measured, not estimated, and available today

Without a baseline you cannot prove the outcome later

What does being wrong cost?

A bounded, stated consequence

If nobody has thought about it, the review burden is unpriced

What does reviewing the output cost?

Minutes per item, from someone who would do the reviewing

If review is assumed free, the economics are untested

Is the data available today?

Accessible now, by this team, in this quarter

"After the migration" means this is not the idea to start

What happens if we do nothing?

A specific consequence with a timeframe

If nothing happens, the idea is optional and will be deprioritized anyway

Who can kill this, and on what evidence?

A named person and a stated condition

Without kill criteria, the project ends by exhaustion instead of decision

Run it with the people who would actually do the work, not with the people who would approve it. An hour is usually enough, and most of that hour is spent discovering which questions nobody can answer.

A failing score has two different meanings

This is the part most kill processes get wrong. A bad idea and a good idea with a missing precondition look identical on the scorecard and should be handled completely differently.

If the workflow has no owner and no number moves, that is a bad idea. Bin it and say so.

If everything holds except that the data lands after a platform migration in two quarters, that is not a bad idea. That is a good idea with a date. Write the re-entry condition down, name who checks it, and put it back in the queue. Ideas that get killed without a re-entry condition come back anyway, six months later, with the same gaps and a new sponsor.

The filter has limits. It is a filter, not a forecast. Ideas that pass can still fail on execution, and occasionally an idea that fails the filter turns out to be right for reasons the questions do not capture. What it catches reliably is the idea that nobody has examined, which is the overwhelming majority of what it will see.

What the ideas that survive have in common

The surviving tenth is never the impressive one. It is usually a specific, unfashionable task with a number already attached to it.

The common characteristics are consistent. Scope is bounded to something a small team can ship inside a quarter. The owner feels the pain personally rather than representing a group that feels it. A before-state was measured before anyone wrote code. The cost of a wrong answer is tolerable or cheaply caught. And the path to production exists on day one, meaning the systems it needs to touch are systems this team can already touch.

Two illustrative shapes, both common:

A support team with a measured median first-response time and a queue where most tickets fall into a small set of categories the team already recognizes. The number exists, the owner is the support lead, the failure mode is a wrong draft that a human deletes in seconds, and the tooling already has an API.

A quality assurance process where writing test cases for a release consumes a known number of hours, tracked because someone has to staff it. Same structure. Real baseline, named owner, cheap review, low blast radius, and the systems involved are already accessible.

Neither of those would survive a pitch meeting against something more ambitious. That is part of why they work. Both are the kind of thing that ships and then gets extended, which is how AI capability actually accumulates in a company.

The same mistake, one technology cycle earlier

None of this is new, which should be slightly alarming. What is new is that the thing which used to stop bad ideas has been removed.

The machine learning era produced the same failure profile for the same structural reason. The postmortems from that period name data quality, deployment gaps, missing MLOps, and organizational resistance. All downstream. The projects that failed generally failed because somebody picked a problem the business did not need solved, and everything after that was consequence.

Building used to be expensive, and that expense was itself the filter. When a machine learning project meant hiring data scientists, labeling a dataset, and spending six months before anyone saw a result, the cost of starting forced a conversation about whether to start. The conversation was not always a good one, but it happened, because somebody had to sign for it.

That filter is gone. A convincing prototype now takes an afternoon. The cost of starting has collapsed to nearly nothing, and the cost of finishing has not moved at all. Production still needs evaluation, monitoring, error handling, security review, change management, and someone to own it at 3am. What collapsed was only the part that used to trigger scrutiny.

So companies now generate far more AI ideas, validate them far less, and discover the cost much later. The pre-build filter used to be enforced by economics. Now it has to be enforced deliberately, by a person, in a meeting nobody wants to call.

Saying no has to be a decision the organization makes

The reason bad ideas survive reviews is rarely that nobody noticed. Usually three people noticed and none of them said it out loud, because individual refusal is expensive in a way individual enthusiasm is not.

Say yes to a bad idea and you share the failure with everyone else who said yes. Say no to an idea someone senior likes and you own the objection personally, immediately, and alone. If it later succeeds elsewhere, you own that too. The asymmetry is obvious to everyone in the room, which is why the room stays quiet.

That is an organizational design problem, not a courage problem, and it has organizational fixes.

Define kill criteria before work starts , in writing, agreed by the sponsor. "If we cannot access the source data by week three, we stop" is a decision the group made in advance, and executing it later is administration rather than confrontation.

Give the decision to start a named owner, the same way the decision to ship has one. Somebody signs for entry into the backlog.

Keep a written record of what was declined and why. This is the piece most teams skip, and it is the one that compounds. A visible list of rejected ideas with reasons turns saying no from an event into a practice, and it stops the same idea returning every two quarters with fresh packaging.

Review on a standing cadence rather than by exception. An idea that has to be actively killed will not be. An idea that has to be actively renewed usually gets an honest assessment.

Where prompt engineering earns its keep

Prompt craft has real value. Doing it before the decision is what strips the value out.

Once an idea has survived triage, prompt and context design becomes real engineering work and deserves real investment. Not the afternoon-of-fiddling version. The version with an evaluation set built from actual failure cases, versioned prompts, regression tests that run before a change ships, and measurement of the specific behavior the business cares about rather than a general sense that outputs look better.

That work is hard, and the vendor guidance on it is good. The difference is what it is attached to. Applied after selection, careful specification is how a validated idea becomes a reliable system. Applied before selection, it is a way of feeling productive about a question that has not been asked.

Decide, then specify. Doing it in that order costs nothing. Doing it in the other order costs a quarter.

What the ideas you did not kill actually cost

Most accounting of failed AI projects stops at spend. Tokens, licenses, contractor days, some fraction of salary. That number is usually small enough to absorb, which is exactly why the failures continue.

The real cost is the part nobody invoices.

Four engineers on a doomed pilot for a quarter is a quarter of that team's output, permanently spent. The roadmap slot that idea occupied was not available to anything else, and roadmap slots are the scarce resource in most organizations. The measurement infrastructure the team built for a workflow that did not matter is infrastructure that was not built for one that did.

Then there is the credibility. Every AI project that demos well and changes nothing raises the evidentiary bar for the next one. The finance partner who funded the last three asks harder questions about the fourth, which is rational. But the fourth one might be the good one, and it now has to clear a bar set by the failures of the first three. Organizations that have burned their credibility on unexamined ideas end up unable to fund the examined ones.

Killing nine ideas is what makes the tenth affordable and staffable, and what makes the room believe you when you bring it.

Key takeaways

  • The published failure rates for AI projects run from roughly 70% to 95% depending on the definition, and the causes they name are almost always conditions that were already true before the project started.
  • Refining a prompt is the cheapest available substitute for deciding whether the idea deserves a build, which is why it is the thing that gets done first.
  • Four shapes predict failure reliably: no named owner, a tolerated problem, an unpriced review burden, and a workflow that was already broken.
  • An hour of structured triage with the people who would do the work catches the overwhelming majority of ideas that should not proceed.
  • A bad idea and a good idea with a missing precondition need different outcomes. One gets binned, the other gets a written re-entry condition.
  • The cost of the ideas you do not kill is measured in roadmap slots and credibility, not in API spend.

If your backlog is longer than your capacity and you are not sure which parts of it deserve a build, that is the problem worth solving first. The AI Readiness Snapshot is a free 30-minute call that maps where AI will actually move a number in your operation, and which of your current ideas are waiting on a precondition rather than a decision.

Frequently asked questions

Most AI projects fail because the idea was never validated before the build started, not because the model or the engineering underperformed.

The causes named in the postmortem research are consistent across sources: the problem was misunderstood or optimized for the wrong metric, the data was not available, no one owned the workflow, and there was no path to production. RAND's 2024 study on AI project failure found more than 80% of AI projects fail, roughly twice the rate of non-AI IT projects, and its first named root cause is stakeholders misunderstanding or miscommunicating the problem to be solved. Every one of those conditions is knowable before a line of code is written. They surface during delivery because that is when someone finally checks, which is why the failure looks like an execution problem when it originated at selection.

That figure comes from one specific study with one specific definition, and it is not a general statement that 95% of AI work fails.

MIT's NANDA initiative published The GenAI Divide: State of AI in Business 2025, based on 150 interviews, 350 employee surveys, and 300 public implementations. It found roughly 95% of enterprise generative AI pilots produced no measurable return in the profit and loss statement. That is a narrow claim about pilots and P&L impact, not about projects being abandoned.

Other measurements use other definitions. Gartner forecast in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. RAND put the AI project failure rate above 80%. The widely quoted "85% of machine learning projects fail" traces back to a 2018 Gartner forecast that 85% of AI projects would deliver erroneous outcomes through 2022, which is a different claim from failure. Treat the honest range as roughly 70% to 95% depending on what is being counted.

Test the idea against evidence someone can produce today, and treat any question nobody can answer as the finding rather than a gap to fill later.

Six checks do most of the work:

  • Does a named person own the workflow, or is the owner a department?
  • Does an existing metric move, and what is its current measured value?
  • What does a wrong answer cost?
  • What does reviewing each output cost, in minutes, for the person who would review it?
  • Is the data accessible to this team this quarter, rather than after a migration?
  • Who can kill the project, and on what stated evidence?

An idea where review time approaches the time to do the task manually fails on economics no matter how accurate the model gets. The whole exercise takes about an hour with the people who would do the work, and most of that hour is spent discovering which questions have no answer.

Yes, but its value is conditional on the idea being worth doing, and it is the wrong first move.

Once an idea has passed triage, prompt and context design is real engineering: evaluation sets built from actual failure cases, versioned prompts, regression tests before changes ship, and measurement of the specific behavior the business cares about. The vendor guidance from Anthropic, OpenAI, Google, and Microsoft on techniques like clear instruction, structured output, worked examples, and task decomposition is good, and the techniques work. The field has also moved past prompt wording alone toward context engineering, which governs what a system knows at the moment it answers.

None of that changes whether producing the output matters. Refining a prompt is fast, visibly skilled, and never requires telling anyone their idea is not worth doing, which is exactly why it gets done in place of the decision. Decide, then specify.

A bad idea has no owner and no measurable outcome. A good idea at the wrong time has both, and is blocked by a single named precondition with a date.

The distinction matters because the two need opposite handling. An idea where the workflow belongs to no one and no existing metric moves should be declined and recorded as declined. An idea that passes every check except that its source data lands after a platform migration two quarters out should go back in the queue with a written re-entry condition and a named person who checks it.

Skipping that write-up is why the same idea returns six months later with the same gaps and a new sponsor, and gets evaluated from scratch by people who do not know it was already examined.

Agree the kill criteria in writing before work starts, so stopping is the execution of a rule the group already accepted rather than someone's personal objection.

Individual refusal is expensive in a way individual enthusiasm is not. Agreeing to a bad idea distributes the failure across everyone who agreed, while objecting to an idea a senior person likes concentrates the risk on one person. That asymmetry, not a shortage of courage, is why weak projects survive reviews.

Stage-gate practice fixes it structurally: criteria are defined before the stage begins so the team knows on day one what it must prove, the decision is made against those criteria rather than ones invented at the meeting, and the reasoning is recorded. A standing review cadence matters too, because an idea that must be actively killed rarely is, while an idea that must be actively renewed gets an honest assessment.

Editorial statement

This article was produced with AI assistance and has undergone human review and editorial control. Iryna Tkachuk holds editorial responsibility for its content within the meaning of Article 50(4) of Regulation (EU) 2024/1689.