{"id":17511,"date":"2026-09-14T16:18:32","date_gmt":"2026-09-14T14:18:32","guid":{"rendered":"https:\/\/www.sunmatrix.com\/ramble\/?p=17511"},"modified":"2026-09-14T16:35:49","modified_gmt":"2026-09-14T14:35:49","slug":"my-ai-finished-i-wasnt-done","status":"publish","type":"post","link":"https:\/\/www.sunmatrix.com\/ramble\/my-ai-finished-i-wasnt-done\/","title":{"rendered":"My AI Finished. I Wasn&#8217;t Done."},"content":{"rendered":"<p>An AI workflow updated 125 files in my knowledge system. Then the chat broke.<\/p>\n<p>The files had changed, but I had no completion message. Fortunately, I was already using two chats: one to develop the strategy and prompt, the other to execute with a more capable model. While the execution chat was unavailable, I used the first to inspect the saved results without repeating the update.<\/p>\n<p>For leaders putting AI into recurring work, the practical question is how much work a team can safely hand over. My recent experiments show why the answer depends on verified delivery, clear ownership and the effort that comes back when a run fails.<\/p>\n<h2>The handover: 125 files and a broken chat<\/h2>\n<p>This was a controlled update to <a href=\"https:\/\/www.sunmatrix.com\/ramble\/when-a-second-brain-becomes-an-operating-model\/\" target=\"_blank\" rel=\"noopener noreferrer\">my Second Brain<\/a>. My Second Brain is a structured collection of knowledge files that AI can read and update within agreed limits. I used ChatGPT Work with Astra at Extra High to add four YAML metadata fields to each of 125 Markdown files. These are structured properties at the top of a file.<\/p>\n<p>A streaming error prevented a usable completion message, and the Work chat became inaccessible on my phone. Some hours later, back at my desk, I opened a support ticket with OpenAI. Independent of the open ticket, after a six-hour wait, the chat became accessible through my desktop browser, but its visible history was blank. When I asked, the AI said it could still access the history and confirmed completion.<\/p>\n<p>Because the updated files were saved outside the failed conversation, the other chat could inspect them. It confirmed that all 125 contained the four requested fields.<\/p>\n<p>For teams maintaining shared content and reporting, the equivalent requirement is that completed work remains verifiable when the conversation that produced it becomes unavailable.<\/p>\n<p>Consider a regional content update handed from one team to another. The recipient needs to know which records changed, which checks passed and what remains unresolved before deciding whether anything needs to be repeated. Otherwise, even a successful update can leave the next team investigating instead of using it.<\/p>\n<p>The update succeeded. The handover failed. I still had to spend time checking that the work was complete.<\/p>\n<h2>The report: plenty of output, more auditing<\/h2>\n<p>My weekly campaign-intelligence task exposed a different problem.<\/p>\n<p>One report had 16 campaign cards and six AI items. The audit still rejected its coverage claim. Some video links were marked as usable without confirming that readers could actually watch the campaign videos.<\/p>\n<p>I went through three audit passes on the campaign workflow. A later weekly run still needed more corrections.<\/p>\n<p>The brief sounded straightforward: find relevant campaigns, establish what launched and when, and provide working links to the sources and relevant campaign videos. The report arrived. Then I had to check whether it had met that brief.<\/p>\n<p>An unchecked campaign claim can become an agency briefing error. Catch it after work has started, and brand and agency teams may have to redo the research, recommendation and brief.<\/p>\n<p>Updating the runbook gave the next run better instructions. But I still had to check subsequent reports to see whether the same failures returned.<\/p>\n<p>The real question is how much work remains with the person who supposedly delegated the task to the AI.<\/p>\n<h2>The fallback: ten URLs, back on my desk<\/h2>\n<p>My website-indexing workflow made that burden concrete.<\/p>\n<p>The intended daily routine was to inspect eligible Ramble pages in Google Search Console, request indexing for up to ten URLs and record the outcomes. This meant submitting requests for Google to consider.<\/p>\n<p>In a batch of ten selected URLs, ChatGPT Work successfully submitted one before Google Search Console returned an error. That left nine unsubmitted. I submitted those manually.<\/p>\n<p>I subsequently defined a recovery path. On the next daily occurrence, the ChatGPT Work cloud browser could not be opened or run for the task. ChatGPT advised me to open a support ticket. But repeated checking and manual execution left too much work with me. So I stopped the workflow without opening a support ticket for this failure.<\/p>\n<p>The technical cause remains unresolved. The operating outcome was clear: this workflow had failed to take the recurring task off my hands.<\/p>\n<p>In a content operation, someone still has to notice the missed work, clear the backlog and decide whether the workflow should continue. A manual fallback can be sensible. It still needs an owner, time and a place in the business case.<\/p>\n<h2>The infrastructure: delivering a checked result<\/h2>\n<p>My LinkedIn post-performance workflow showed what a more complete delivery could look like. It used authenticated browser access to open the analytics, export the native spreadsheet and inspect its contents. It then recorded the metrics in my maintained Second Brain source, read them back to check that everything had been saved, deleted the exact temporary export and verified its removal.<\/p>\n<p>That sequence of connected access, execution and verification completed successfully. The result was available in the record for subsequent analysis, with the temporary file removed after the checks.<\/p>\n<p>For shared marketing reporting, the same approach would give the next analyst a saved result they could inspect and reuse. Source retrieval, recording and verification would be part of delivery, reducing the work each recipient has to reconstruct.<\/p>\n<p>This chain produced useful work. Reliable repetition and net time saved still need demonstrating.<\/p>\n<h2>The model: one part of the service<\/h2>\n<p>Comparing ChatGPT Plus and Pro pushed me to separate three questions:<\/p>\n<ul>\n<li>Can the model do the thinking the task requires?<\/li>\n<li>Can the environment complete the required actions?<\/li>\n<li>Does the workflow deliver usable work with an acceptable review burden?<\/li>\n<\/ul>\n<p>A more capable model may help with difficult reasoning. It does not establish that a browser session will work, a saved result will be checked or a missed delivery will reach the right person.<\/p>\n<p>The same separation belongs in an enterprise buying decision. Model quality, tool access, capacity and operating effort each need assessment. A successful run is evidence for that task and setup; attributing the benefit to a higher subscription tier requires a separate comparison.<\/p>\n<h2>The economics: count the work that comes back<\/h2>\n<p>I would judge an AI workflow by accepted output and the total human effort needed to obtain it. By accepted output, I mean a result that meets the agreed requirements and is usable for its intended purpose.<\/p>\n<p>If a report needs two repairs before acceptance, it is one delivered report with three attempts behind it. A scheduled run that produces nothing usable still belongs in the record. So does time spent maintaining instructions, recovering access and finishing the task manually.<\/p>\n<p>For a team pilot, three measures would make the business case more useful: acceptance on first delivery, human review and recovery effort, and whether the output was actually used. Missing and abandoned runs must stay visible. Before claiming time savings, count the time spent running, checking and fixing the AI workflow, then compare it with the time the same task took before.<\/p>\n<p>Ownership matters just as much. Someone needs responsibility for missed runs and recovery. Someone needs authority to accept the result. In a small team, those may be the same person. Both responsibilities still need to be explicit.<\/p>\n<p>Controls also take time. If every harmless step needs approval, the task keeps returning to its owner. Checks should sit at consequential decisions, while already-authorised work continues. The result must be worth the money and staff time spent producing, checking and correcting it.<\/p>\n<h2>The next step: test the AI handover<\/h2>\n<p>The next test is whether a colleague can use the result and handle a failed run without calling the person who built the workflow. That shows whether the team can operate it when its creator is unavailable.<\/p>\n<p>Choose one recurring workflow and define what its recipient needs to accept the result. Name who handles failures, run it through its normal schedule and test a handover to another team member. Record missing and repaired deliveries, actual use and total human effort against the existing process. Expand when the evidence supports the extra scale; narrow or stop when the work keeps coming back.<\/p>\n<hr>\n<h2>A few fast answers before you act<\/h2>\n<h3>Is a scheduled task the same as an AI agent?<\/h3>\n<p>No. A schedule determines when work starts. An AI agent uses a model, instructions and tools to pursue a goal and choose its next steps. In ChatGPT Work, a scheduled task can run this kind of workflow. Setting the schedule alone does not make it reliable.<\/p>\n<h3>What should count as successful AI delivery?<\/h3>\n<p>Successful delivery means the result meets its agreed requirements and is available where the recipient needs it. Record any review, correction or recovery required. A useful result can still have failed its first delivery.<\/p>\n<h3>Why keep AI workflow records outside the chat?<\/h3>\n<p>Saved results and decision records support continuity when a conversation becomes unavailable. Recovery depends on what was preserved and can be inspected. Another AI\u2019s assurance of completion should be checked against the saved result.<\/p>\n<h3>Will a more expensive AI plan make workflows reliable?<\/h3>\n<p>A higher price does not establish reliable execution. Evaluate model quality, tools, access, capacity and recovery separately, using comparable work. Attribute a subscription benefit only when the evidence supports it.<\/p>\n<h3>How much governance does an AI workflow need?<\/h3>\n<p>Match controls to the consequences. Name who accepts the result and who handles failures. Use mechanical checks where possible, reserve approval for decisions that require it and let already-authorised work continue. Measure the review burden alongside errors caught.<\/p>\n<h3>What should an enterprise AI pilot measure?<\/h3>\n<p>Measure acceptance on first delivery, human review and recovery effort, and practical use against the existing process. Include missing runs, abandoned work and instruction maintenance. Test whether another team member can operate the workflow.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>An AI workflow updated 125 files in my knowledge system. Then the chat broke. The files had changed, but I had no completion message. Fortunately, I was already using two chats: one to develop the strategy and prompt, the other to execute with a more capable model. While the execution chat was unavailable, I used &hellip; <a href=\"https:\/\/www.sunmatrix.com\/ramble\/my-ai-finished-i-wasnt-done\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">My AI Finished. I Wasn&#8217;t Done.<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":17515,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_seopress_titles_title":"","_seopress_titles_desc":"What makes an AI workflow reliable? Lessons from failed handoffs, manual recovery, human review and verified delivery.","_seopress_robots_index":"","_seopress_robots_follow":"","_seopress_robots_imageindex":"","_seopress_robots_snippet":"","_seopress_robots_primary_cat":"","_seopress_robots_breadcrumbs":"","_seopress_robots_freeze_modified_date":"","_seopress_robots_custom_modified_date":"","_seopress_robots_canonical":"","_seopress_social_fb_title":"","_seopress_social_fb_desc":"","_seopress_social_fb_img":"","_seopress_social_fb_img_attachment_id":0,"_seopress_social_fb_img_width":0,"_seopress_social_fb_img_height":0,"_seopress_social_twitter_title":"","_seopress_social_twitter_desc":"","_seopress_social_twitter_img":"","_seopress_social_twitter_img_attachment_id":0,"_seopress_social_twitter_img_width":0,"_seopress_social_twitter_img_height":0,"_seopress_redirections_value":"","_seopress_redirections_enabled":"","_seopress_redirections_enabled_regex":"","_seopress_redirections_logged_status":"","_seopress_redirections_param":"","_seopress_redirections_type":0,"_seopress_analysis_target_kw":"","_seopress_news_disabled":"","_seopress_video_disabled":"","_seopress_video":[],"_seopress_pro_schemas_manual":[],"_seopress_pro_rich_snippets_disable_all":"","_seopress_pro_rich_snippets_disable":[],"_seopress_pro_schemas":[],"iawp_total_views":36,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[6418,32],"tags":[6725,6696,6721,6720,9948,9947,9943,6730,9944,9945,9941,6689,9946,6727,6691,672,9934,9740,9942,9949,9940,6726],"class_list":["post-17511","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","category-emerging-technology","tag-agentic-workflows","tag-ai-agents","tag-ai-governance","tag-ai-operating-model","tag-ai-productivity","tag-ai-reliability","tag-astra","tag-chatgpt","tag-chatgpt-plus","tag-chatgpt-pro","tag-chatgpt-work","tag-enterprise-ai","tag-google-search-console","tag-human-in-the-loop","tag-knowledge-management","tag-linkedin","tag-marketing-measurement","tag-martech","tag-openai","tag-scheduled-tasks","tag-second-brain","tag-workflow-automation"],"jetpack_shortlink":"https:\/\/wp.me\/pgYpE1-4yr","jetpack_featured_media_url":"https:\/\/www.sunmatrix.com\/ramble\/wp-content\/uploads\/connection_lost.jpg","_links":{"self":[{"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/posts\/17511","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/comments?post=17511"}],"version-history":[{"count":4,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/posts\/17511\/revisions"}],"predecessor-version":[{"id":17518,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/posts\/17511\/revisions\/17518"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/media\/17515"}],"wp:attachment":[{"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/media?parent=17511"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/categories?post=17511"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.sunmatrix.com\/ramble\/wp-json\/wp\/v2\/tags?post=17511"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}