Practical limits
An answer is not an outcome: what has to be true before AI can finish work for you
Assistants now answer well and act inside limits. The distance between a good answer and finished work is made of records, authority, checks and follow-up, and most of that is on the business, not the model.
- Published
- September 9, 2026
- Last checked
- September 9, 2026
- Written by
- Dryvn AI editorial
- Edited by
- Dryvn AI editorial
- Reviewed by
- Not yet reviewed by a named person
- Read as
- Editorial (Dryvn's practical method)
The moment the answer is right and nothing has happened
Ask an assistant how to chase an unpaid invoice and it will give you a decent answer. Ask it to draft the message and it will draft a good one. Ask it to send the message, log that it was sent, notice when the customer does not reply, send the next one, stop when payment lands, and tell you only if something goes wrong, and you have described a job. The answer was the easy part.
We think this is the most common misunderstanding about AI in business right now. People see the quality of the answers and assume finished work is close behind. Sometimes it is. Usually the gap is not the model. It is everything around it.
What is genuinely true today
We want to be fair to the tools, because they have moved. OpenAI's agents documentation describes agents that plan, call tools, hand off between specialists and keep enough state to complete multi-step work, and it documents guardrails and approval flows that pause before risky steps [S29]. Anthropic's computer-use documentation describes a model that can take screenshots and operate a mouse and keyboard, with the advice to run it in an isolated environment and have a person confirm decisions that carry real consequences [S30]. Those are vendor claims from official pages, checked on 9 September 2026, and they describe capability that exists.
So an assistant can act. What the same documentation says, in every case, is that it should act inside limits with a person confirming what matters. That is not a disclaimer. It is a description of how finished work actually gets finished.
The four things an answer does not have
A record to act on
To chase the invoice, the system needs to know which invoice, for which customer, at which amount, sent on which date, with which terms. If that lives in three places or in someone's head, the assistant either guesses or asks. Guessing is how you get the wrong customer chased. Asking is fine, but then the person is still doing the work.
Authority to act
Sending a reminder is one thing. Offering a payment plan, waiving a fee or threatening to pause work is another. Somebody in the business decides which of those an assistant may do on its own, which need a yes, and which it may never do. If nobody has written that down, one of two things happens: the assistant does too little and people stop trusting it, or it does too much and people stop trusting it. The approval boundaries worksheet exists because this decision is almost never made before the tool is bought.
A check that the outcome happened
"Sent" is not "paid". A message going out is evidence that a message went out. Finished work is the money in the account, the appointment confirmed, the part delivered. Somebody has to decide what counts as evidence for each piece of work and make sure the system looks for it. Without that, the assistant reports success on the easy part and the business finds out about the hard part from the customer.
Follow-up when it did not
Most work does not complete on the first try. The customer does not reply, the supplier is closed, the tenant is not home. A system that finishes work needs an exception route: a place the item goes when the normal path fails, and a person who looks there. This is the loop, and it is the difference between a workflow that runs steps and a system that gets things done. The systems and loops guide explains that distinction without any product in it.
Illustrative example · fictional
The same task, with and without the four things
A fictional cleaning company asks an assistant to handle overdue invoices. Version one: the owner forwards an invoice and says chase this. The assistant drafts a polite reminder, the owner sends it, and that is the end. Two weeks later the owner remembers and does it again. The assistant did its job. The work is not done.
Version two: invoices live in one place with a customer, an amount and a due date. The owner has written that the assistant may send two reminders on days 7 and 14 in the company's own words, may not offer discounts, and must hand anything unanswered at day 21 to her. Payment landing in the account is the evidence of completion; the assistant checks for it before each reminder and stops when it sees it. Unanswered items appear in one list she reviews on Fridays. Now the assistant is doing the work, and she is doing the deciding.
The model is identical in both versions. Everything that changed is a decision the business made.
Why this is good news
It means the useful work is available to you now and does not depend on which model wins next year. Records, authority, evidence and exception routes are things you can put in place this month. They make work finish more reliably with or without AI, and they are exactly what any capable system needs before it can carry more. Anthropic's own guidance on building agents says to add complexity only when it demonstrably improves outcomes [S11]. We would put it more bluntly: fix the handoff before you buy anything that promises to automate it.
A checklist before you hand work to any system
- Can you point to the record the work acts on, in one place, with the fields the work needs?
- Is it written down what the system may do alone, what needs approval, and who can revoke it?
- Have you defined what counts as evidence that the outcome happened, not just that a step ran?
- Is there a named place unfinished items go, and a person who looks there on a fixed day?
- Have you run three real items through it with a person watching before you let it run unattended?
If any answer is no, the work is not ready to hand over, and that is true whether the system is an assistant, an automation or a new hire. The handoff assessment walks one piece of work through those questions and gives you a one-page action sheet.
The maintained resources behind this article
- AI assistant, automation, agent or business software: which do you need?Tell the four categories apart by the job each one does, and pick the right one for a specific task in your business, including the case where what you already own is enough.
- Where is work getting stuck between your tools?Find the handoffs in your business that still depend on copying information, remembering follow-ups or chasing people, and get one next action for each gap.
- What should AI be allowed to do without asking?Decide, action by action, what an AI assistant or automation may do on its own, what it may do inside an approved scope, and what always waits for a person.
- What is an AI system, a workflow and a loop?Learn the difference between a system, a workflow, an automation, an agent and a feedback loop using two everyday business examples, and map one of your own on a worksheet.
Sources
Dates are when each source was last checked by the editor. Sources support specific claims; they are not endorsements.
- S29Agents SDK · OpenAI developer documentation · checked September 9, 2026Describes agents that plan, call tools, hand off between specialists and keep state, with guardrails and resumable approval flows that pause before risky work. Vendor documentation.
- S30Computer use tool · Claude Platform documentation (Anthropic) · checked September 9, 2026Describes desktop control through screenshots, mouse and keyboard, and advises human confirmation for decisions with meaningful real-world consequences. Vendor documentation.
- S11Building effective agents · Anthropic · checked September 9, 2026Published 19 December 2024. Distinguishes workflows from agents and advises adding complexity only when it demonstrably improves outcomes.
B01 · Published September 9, 2026 · Opinions are the author's. Vendor facts are dated and sourced above; scenarios are labelled as scenarios.
