A hand switches off one workflow on a control panel while neighbouring channels remain active or under adjustment.

AI Maturity Is Not a Race to Level 7

I shut down one of my most important AI workflows after a week dominated by attempts to make it reliable.

It had connected sources, a detailed runbook, persistent records and acceptance checks. The final run still failed. Instead of reducing my work, it kept giving me more audits, reruns and repairs.

That changed how I think about AI maturity. Being able to build more does not mean you should run more.

The ladder: a seven-level AI capability model

In my development notes, I use a seven-level AI capability model. It is a working framework, not an accredited ranking.

Level 1: Casual User. Uses AI for simple answers and rewriting. Think: rewrite this email, summarise this document or explain this topic.

Level 2: Prompt User. Writes decent prompts but works ad hoc. Think: give AI a good brief for a presentation or campaign, but start again next time.

Level 3: Structured User. Uses templates, roles, examples and constraints. Think: give AI the same briefing structure, approved examples and rules each time.

Level 4: Workflow User. Turns recurring tasks into repeatable AI workflows. Think: run the same research, drafting and review sequence every week instead of reinventing the process.

Level 5: Advanced Operator. Uses AI across real workstreams with gates, quality checks and source material. Think: research a campaign, challenge unsupported claims, check the sources and approve only what meets the standard.

Level 6: Operating-Model Designer. Builds scalable AI-enabled ways of working with rules, decision gates and measurement. An AI operating model defines how people, AI and systems share work, decisions and accountability. Think: define the sources, acceptance criteria, owner, exception path and metrics so the workflow can be managed consistently.

Level 7: AI Platform Implementer. Connects AI to systems, data, automation, evaluation, governance and measurable outcomes. Think: retrieve approved product data, create the output, route it through the right approval, update the next system, record what happened and measure whether it helped.

The progression is useful. It moves from using AI, to organising the work, to designing how the work operates, to connecting that design across systems.

But the ladder can create the wrong instinct: if Level 6 is good, Level 7 must be better.

My experience taught me otherwise. The model tells me what capability has been built. It does not tell me whether every workflow should use all of it.

The repair loop: more control was not yet proof

A separate weekly campaign-research workflow made that distinction concrete.

A rerun surfaced 15 campaign entries, but that was not a verified count of qualifying campaigns. The audit still failed. Two campaigns were outside the pre-defined launch window. Two others lacked sufficient evidence that they had launched within it. The required video-verification process had not been demonstrated for every counted campaign.

The visible count problem had been fixed. The qualification problem had not.

Because campaigns could enter the counted total before their evidence was cleared, a report could meet its numerical target without meeting the brief.

For brand, content and commerce teams, that is the difference between a full research pack and a sound basis for the next campaign brief.

The same type of failure had appeared before, so another rerun was not enough. The rules themselves needed repair. Launch evidence, completion of the video-verification process and reconciliation of the totals now had to pass before a campaign could enter the count.

That improved the design. It did not yet prove better delivery.

The next run still had to earn that proof.

What worked: measurement without a new tactic

Another workflow showed the useful side of automation.

My LinkedIn post-performance process uses the same checkpoints for every tracked post: 24 hours, 48 hours and seven days. It preserves LinkedIn’s own analytics evidence, records the metrics and keeps missing or unattributable results visible instead of filling the gaps with assumptions.

After seven posts had completed that measurement cycle, the evidence did not justify changing the publishing strategy.

So I did not change it.

That is less exciting than announcing a new formula for follower and impression growth. It is also what a useful measurement system should do: change the decision only when the evidence supports it.

The workflow therefore produced something useful: a defensible decision.

That still does not prove that every part of the automation pays for itself. Useful evidence, reliable execution and positive ROI remain separate claims.

The test: three answers, not one score

I now read the maturity ladder alongside three questions: Can it run? Is the work accepted? Is it worth running?

Can it run? Prove that the workflow can reach the required sources, perform the permitted actions and produce an output.

Is the work accepted? Prove that repeated outputs meet the agreed standard. Keep failed runs, corrections and missing results visible.

Is it worth running? Compare the useful result with the existing process or a simpler alternative. Include setup, required review, maintenance and repair.

Having someone approve the work is different from needing them to keep fixing it. Approval is a control. Repeatedly fixing the same failure is repair. Both cost time, but only one signals a reliability problem.

Capability maturity describes what the AI-enabled system can reliably do. Outcome maturity asks whether that capability repeatedly creates worthwhile improvement. In my framework, full Level 7 requires both.

The real question is not how much AI I can get to operate. It is what is worth keeping in operation.

The decision: fund the outcome, not the level

Stopping is not automatically mature. Some workflows deserve a bounded repair. Others genuinely need deeper integration. I would judge maturity by whether the next capability improves the outcome, not by whether it moves the workflow up a level.

At the next AI rollout review, put three things side by side: implementation evidence, acceptance history and the business benefit the workflow was meant to create. Expand when the evidence supports it. Repair when the failure is specific and fixable. Simplify or stop when the value does not justify the effort. Fund the outcome, not the next level.


A few fast answers before you act

What are the seven AI capability levels?

The working model moves from Casual User, Prompt User, Structured User and Workflow User to Advanced Operator, Operating-Model Designer and AI Platform Implementer. It describes increasing capability, not a requirement that every workflow reach Level 7.

Can a Level 6 workflow be the better choice?

Yes. A governed Level 6 workflow may already solve the problem. Moving to connected Level 7 infrastructure makes sense when integration creates measurable additional value, not simply because the capability exists.

When should an AI workflow be stopped?

When repeated failures, repair effort or operating cost outweigh the useful result and another repair does not have a credible value case. Keep the learning even when you stop the workflow.

Does mature AI require less human involvement?

Not necessarily. Human judgment and approval may be necessary controls. The problem is when people have to keep fixing or rerunning work that should already be dependable.

How should leaders test an AI maturity claim?

Ask three separate questions: can it run, is the work accepted, and is it worth running? Success on one does not prove the next.

Published by

Sunil Bahl

Sunil Bahl

SunMatrix Ramble is an independent publication on AI, MarTech, advertising, and consumer experience, published since 2009. Sunil Bahl is a global transformation leader in consumer experience platforms and MarTech, with 27+ years of experience translating digital change into scalable platforms, operating models, and commercially useful outcomes.

Latest on Ramble