AI for Small Business Admin
10 min read · updated August 4, 2026
The question is not which tasks a model can do. It is which tasks are cheap to check and cheap to get wrong. That test rules out most of the paperwork people try first, and it rules in a handful that genuinely save hours.
The test: which tasks qualify
Three questions, in order. A task has to pass all three, and most candidates fail the second.
- Can I check the output faster than I could produce it? Drafting a supplier email: yes, reading takes twenty seconds and writing takes five minutes. Reconciling a bank statement: no, checking is the whole job.
- If a mistake gets through, what does it cost and can I undo it? A clumsy sentence in a newsletter is embarrassing and reversible. A wrong figure in a quotation is a binding number. A wrong figure in a VAT return is a matter for the tax authority.
- Does it need information only I have? If so, the task is not “write this”, it is “write this from what I supply”, and the supplying is most of the work. Sometimes that is still worth it. Often it is not.
The pattern that emerges: transformation of text you already have is reliably worth it, and generation of facts you do not have is reliably not. Everything in the next section is a transformation.
Five tasks and how to do each
1. First drafts of routine correspondence
Chasing invoices, declining a request, confirming an order, explaining a delay. These are high-volume, low-stakes and formulaic, which is the ideal profile. The method that makes them yours rather than generic is in writing emails that still sound like you. Build four or five reusable prompts rather than starting fresh each time; the second use is where the saving is.
2. Turning a mess of notes into a document
Job notes into a quotation, scribbles into a work order, a call into a follow-up email. The input is yours, so nothing is invented, and the output structure is fixed.
Turn the notes below into a customer-facing quotation. Use only the items and prices in the notes. Where a price is missing, write [PRICE?] rather than estimating. Where the scope is ambiguous, list it at the end under ASSUMPTIONS I MADE. Do not total the quotation. I will do that in the spreadsheet. Format: intro sentence, itemised list, assumptions, next step. [notes]
The instruction not to total is the important one and it recurs throughout this page. Arithmetic goes in a spreadsheet, where it is reproducible and auditable, not in prose produced by a language model.
3. Reading things you would otherwise skim
A supplier’s terms change, a long tender document arrives, a council letter runs to nine pages. The extraction method — pull out obligations, dates and figures with quoted sentences and page numbers, then compose from the extract — is in summarising a long report. For anything with legal effect this produces a map of where to read, not a substitute for reading.
4. Standardising your own records
Inconsistent product names across three years of invoices, addresses in four formats, a customer list with duplicates. Text normalisation against a rule you specify is a good fit, provided the output is a proposal rather than an edit.
Here are 200 supplier names as typed by different people. Group them into what you believe are the same supplier. For each group, propose one canonical name. Output three columns: original | proposed canonical | confidence (high/low). Put anything you are unsure about in the low bucket. Do not merge anything. This is a proposal I will review.
Then you review the low-confidence rows and spot-check ten of the high ones. Never let a bulk edit run against your live records from an unreviewed list.
5. Drafting the internal documents you keep postponing
A holiday policy, an onboarding checklist, a returns procedure, a method statement. A generated draft is generic and that is fine: it is faster to correct a generic document than to face a blank page. Anything with legal effect — contracts, employment terms, safety documentation — gets a professional review before it is used, and the draft is a way to arrive at that review with the questions already identified.
Two tasks that look ideal and are not
Both of these are high-volume, repetitive and text-shaped, which is why they are attempted first, and both fail the second question in the test.
Pricing anything. Quotes, estimates, discounts. The task looks like drafting and is actually a decision with a number in it, and a quoted price is a commitment. The safe version is the one above: the model assembles the document, you supply every price, and your spreadsheet does the arithmetic. The moment a figure in the output did not come from you, you are sending commitments you have not made.
An unsupervised chat assistant on your website. This is the small-business project that most reliably goes wrong, because it removes the human from between the draft and the customer, which is the only control the rest of this page relies on. It will answer questions about your returns policy, your lead times and your prices confidently, and it has no access to any of them unless somebody built that. A statement made to a customer by something on your website is, in most consumer protection regimes, a statement made by your business.
Where you do want one, the constraint that makes it survivable is that it answers only from a document you wrote, quotes it, and says it does not know otherwise — the same rule as asking questions of your own documents. That is a real project rather than an afternoon, and it is worth knowing that before starting rather than after.
Is it worth it? Do the arithmetic
Claims about hours saved are meaningless without your own numbers, so here is the calculation with the assumptions labelled. Fill in your own.
FOR ONE TASK
A times per week you do it e.g. 12
B minutes it takes you now e.g. 6
C minutes with a draft (writing the prompt,
reading, correcting) e.g. 2.5
D minutes per week checking, on top e.g. 4
(be honest: this is where estimates fail)
Weekly saving = A x (B - C) - D
= 12 x 3.5 - 4
= 38 minutes
SET-UP COST, PAID ONCE
E minutes to build and test the reusable prompt e.g. 45
Break-even = E / weekly saving = 45 / 38 = about 1.2 weeks
THE TASKS THAT FAIL THIS
A one-off task where E is paid and never amortised.
A task where B is small: shaving 40 seconds off something you do
twice a week is 80 seconds.
A task where D is large, which is every task where checking is
the actual work.Run it on three candidate tasks before changing anything. Usually one is clearly worth it, one is marginal and one is a bad idea that felt compelling, and the arithmetic is what tells them apart. If you want a view of the wider picture first, a realistic starting point for small businesses covers workflow choice at the level above this.
The review step nobody skips twice
Every task above ends in a review, and the review is specific to the task rather than a general instruction to be careful.
| Output | Description |
|---|---|
| customer email | Read for anything you have promised. A drafted email will cheerfully commit you to a delivery date or a discount you did not authorise. |
| quotation | Check every price against your price list and total it yourself. Never send a total the model produced. |
| extracted terms | Find two quoted sentences in the original document. A quote you cannot find invalidates the extract. |
| normalised records | Review every low-confidence row and ten random high-confidence ones before applying anything. |
| policy draft | Assume every legal or numeric claim in it is wrong until checked. Employment law in particular is jurisdiction-specific and a draft will confidently mix jurisdictions. |
Records, tax and anything a regulator reads
Draw one hard line here and it removes a whole category of risk: a language model may help you organise and describe your records. It may not compute a figure that goes on a return, and it may not tell you what your obligations are.
Tax rules are jurisdiction-specific, they change every year, and they are precisely the material where a confident wrong answer is indistinguishable from a right one. The consequences do not fall on the tool. Categorising your own transactions is admin and is fine; the same boundary applies to household finances and is set out in AI for personal finance admin.
Where this goes wrong
The plausible number. The characteristic small-business failure is a figure that appears in a document because the surrounding sentence needed one. It will be the right order of magnitude and it will be wrong. This is why totals live in spreadsheets.
The confident answer about your obligations. Asked about notice periods, VAT thresholds, licensing or employment rights, you will get an answer. It will be fluent, jurisdiction-blind and possibly two years out of date. Use it to know what to look up.
Automation before verification. Connecting a drafting tool directly to an outbox removes the only step that was making the arrangement safe. Keep a human between the draft and the customer until you have a hundred drafts of evidence that the class of task is reliable — and even then, not for anything that quotes a price.