Most businesses do not need another list of AI tools.
They need one workflow that is repetitive, costly enough to matter, supported by usable context, and safe enough to test with human approval.
This is the method I use to decide whether an AI idea earns a 14-day test or gets killed before we spend months building the wrong thing.
Running an AI-first company with 100+ AI agents made the pattern obvious: the first useful question is rarely about the model.
On LinkedIn, I summarised the method in four filters. Below is the full reasoning, the edge cases, and the questions to ask before approving a test.
The free 8-Filter AI Workflow Scorecard at the end takes the four public filters and adds the operational checks needed to make a real decision.
Start with the loss, not the model
Most AI conversations start too far downstream:
- Which model should we use?
- Which agent should we build?
- Which platform should we buy?
Those questions matter after the business problem is clear.
The better first question is:
What repetitive process is already costing the business time, money, delay, errors, lost leads, or customers?
A useful first AI workflow already exists. People are doing it today, its consequences are visible, and the company has enough context to test a narrower version.
If the process is vague, rare, unowned, or impossible to measure, adding AI will not make it a better first project.
The four filters
1. Repetition
A workflow needs enough real volume to generate evidence.
Daily and weekly processes are usually stronger candidates than quarterly tasks because the team can compare enough cases during a short test.
Look for:
- repeated reviews, classifications, drafts, checks, handoffs, and follow-ups;
- the same inputs arriving in a recognisable format;
- a clear trigger that starts the work;
- a person or team that performs it regularly.
Red flags:
- the process happens only a few times per year;
- every case is completely different;
- nobody can say when the workflow starts or ends;
- the workflow exists mainly as an exception.
2. Visible business loss
The workflow needs a baseline.
The loss does not have to be direct cash. It may be:
- hours spent per item;
- response delay;
- avoidable rework;
- missed follow-ups;
- error rate;
- lost conversion;
- a queue that keeps growing.
If the team cannot describe the current state, it will not be able to make an honest ROI claim later.
The purpose of the baseline is not to manufacture a large number. It is to decide what the test must improve.
3. Usable context
AI needs the same operating context a capable employee would need.
That context may live in:
- emails;
- CRM records;
- call transcripts;
- documents;
- tickets;
- policies;
- product data;
- examples of accepted work.
Having a lot of data is not the same as having usable context.
The important questions are:
- Where is the source of truth?
- Is the information current?
- Can the workflow access it safely?
- Are the rules documented?
- Can the team recognise a correct output?
AI cannot rescue a process that exists only in someone’s head.
4. A safe first version
The first version should help a person make or execute a decision. It should not silently take over the decision.
Good first actions include:
- draft;
- classify;
- summarise;
- compare;
- check;
- recommend;
- prepare a record for review.
Keep human approval before actions that are:
- customer-facing;
- financial;
- regulated;
- difficult to reverse;
- based on incomplete context.
The goal is not maximum autonomy. The goal is a test narrow enough to produce trustworthy evidence.
What a 14-day AI test should answer
Do not turn a promising workflow into a six-month transformation proposal.
Give it 14 days to answer four questions:
- Does it reduce time or delay?
- Does the output meet an agreed quality bar?
- Can people use it without creating new operational risk?
- Is the result strong enough to scale, or should the idea change or die?
A practical sequence:
- Days 1–2: measure the current baseline.
- Days 3–5: build the narrowest useful version.
- Days 6–12: run it on real eligible work with human approval.
- Days 13–14: compare time, cost, quality, risk, and adoption.
Then choose one:
SCALECHANGEKILL
Killing a weak AI idea after a narrow test is a useful business result. It is cheaper than funding the wrong transformation.
Example: lead qualification
Imagine a sales team manually reviewing every inbound lead.
The workflow may be worth testing when:
- new leads arrive every day;
- slow triage creates visible response delay;
- CRM records and qualification rules already exist;
- AI can prepare a recommendation while a salesperson approves the action.
The narrow first version is not “an autonomous AI salesperson.”
It may be:
Read the form and CRM context, identify missing information, propose a qualification category, and prepare the next-step note for human approval.
The test can then compare handling time, response delay, returned recommendations, and sales-team adoption.
This example passes the method only if the company can supply the actual rules, context, owner, baseline, and reviewer.
FAQ
What counts as an AI workflow?
A workflow is a repeatable sequence with a trigger, inputs, an owner, an output, and somebody who uses that output. “Use AI in sales” is not a workflow. “Prepare an inbound lead qualification recommendation from the form and CRM record” is.
Should the first AI workflow have the biggest possible ROI?
Not necessarily. The best first workflow combines a meaningful business consequence with enough volume, usable context, and a safe test boundary. A theoretically large opportunity that cannot be measured or reviewed is a weak first pilot.
Do we need perfectly clean data before testing?
No. You need enough reliable context to test a narrow workflow and identify what is missing. If nobody knows which source is authoritative, the first project may need to fix the process or context before adding AI.
Can a high-stakes process be the first AI project?
Usually not as an autonomous workflow. A narrow assistive version may still be testable if AI prepares an output and a qualified person reviews every important decision.
What if we do not have a baseline?
Measure it before building. Record volume, handling time, delay, errors, and exceptions for a small but representative sample. Without a baseline, the team may produce an impressive demo without learning whether the workflow improved.
Why use a 14-day test?
Fourteen days is long enough for a recurring workflow to produce real cases and short enough to prevent a weak idea from becoming a large implementation by inertia. Rare or highly regulated workflows may require a different test window.
Does a high score guarantee ROI?
No. The scorecard measures whether the workflow is testable enough to earn evidence. The 14-day result determines whether to scale, change, or kill it.
Does this replace an AI strategy?
No. It provides a disciplined next step. A wider strategy still needs business priorities, architecture, governance, security, ownership, and a portfolio of measured opportunities.
Score one real workflow
The free 8-Filter AI Workflow Scorecard converts the method into a 24-point decision.
It gives you:
- the four public filters plus four deeper operating filters;
- hard gates that override a misleadingly high total;
- fields for the current baseline and missing evidence;
- a human-approval boundary;
- an economic-headroom check;
- a ready 14-day test plan;
- a final
SCALE,CHANGE, orKILLdecision.