Back to Home
Automation

Human Approval Before an Automated Send: The Legal and Engineering Case

Why anything your automation sends a customer needs a human checkpoint in Australia, the three approval patterns, and how to stop one becoming a rubber stamp.

13Labs Team29 July 202613 min read
human in the loopapproval workflowsSpam Actautomation compliancebuildAutomation

Contents

Why Every Automated Send Needs a Human Checkpoint

Anything an automation sends to a customer needs a person approving it first, because the business that owns the automation carries the legal liability for every word in the message. The software vendor does not. This is not a statement about how good the models are. It is a statement about who a regulator writes to when a message is wrong. Under Australian law the sender of a commercial electronic message, the maker of a price representation, and the entity holding personal information are all the business. An automation is a tool the business used, in the same way a mail merge is a tool the business used. So the engineering question is not whether to have a checkpoint. It is how to build one a human actually reads. A queue that shows a hundred finished drafts and an approve button teaches the reviewer to click approve, which is worse than no gate at all because it produces an audit log that says a person checked when nobody did. What follows is the error-rate case for a gate, the four Australian legal regimes that make it non-optional for customer messages, the three approval patterns worth knowing, and the design details that separate a real checkpoint from a formality.

Error Rate Is a Function of Grounding, Not a Fixed Property of the Model

How often a language model produces something false depends mostly on how tightly the output is constrained to verified data you supplied, not on which model you picked. On grounded tasks, where the model is summarising or reformatting source material placed directly in front of it, top-tier models report hallucination rates of roughly 0.7% to 1.5% across 2025 and 2026 model-evaluation reporting. On high-complexity reasoning tasks, where the model has to work something out rather than restate it, reported rates exceed 33%. Same models, same week, different job. That spread is the entire design brief. An automation that drafts an email from three verified comparable sales and a contact record sits near the low end. An automation asked to decide what a property is worth sits near the high end, and no amount of prompt rewriting moves it down. If you want a lower error rate, you change what you are asking for and what you feed it, not which vendor you buy from. Ignore the figure circulating in 2026 that language models hallucinate 50% to 82% of responses. It comes from statistics roundups that blend incompatible task types into one average, and it cannot tell you anything useful about where a gate belongs. A number that does not distinguish grounded summarisation from open-ended reasoning is not a measurement of your system. The organisational picture is not improving on its own. McKinsey's Global Survey on AI found the share of organisations reporting at least one negative consequence from generative AI rose from 44% in 2024 to 51% in 2025, with inaccuracy the most commonly reported risk of all.

The Spam Act Makes You the Sender, Whatever Pressed Send

Every commercial email or SMS your automation sends is a commercial electronic message under the Spam Act 2003 (Cth), enforced by ACMA, and the obligations attach to your business regardless of whether a person or a scheduled job triggered it. There are three, and they are short. You need consent, express or inferred, before you send. The message must clearly identify the sender and carry accurate contact details. And it must include a working unsubscribe facility, with requests honoured within 5 business days. Penalties are calculated by reference to how many messages went out in a day. A business that sends more than 50 commercial electronic messages without consent on a single day is exposed to up to 1,000 penalty units. We are quoting penalty units rather than dollars on purpose. The Commonwealth penalty unit value moves with indexation, and dollar conversions published a couple of years ago are already stale. Look up the current value in the Crimes Act 1914 before you convert it. Automation changes the shape of this risk rather than the size of it. A person sending fifty emails by hand notices around email thirty that the list looks wrong. A scheduled job does not notice, and it does not stop at fifty. ACMA's recent penalties are not small, and they are not confined to obvious spammers. | Business | Penalty | Date | |---|---|---| | Southern Phone | A$2.5 million, plus about A$1.2 million in customer refunds | September 2025 | | Betfair | A$871,000 | June 2025 | | Optus | A$826,000 | September 2025 | | Exetel | A$694,000 | June 2025 | ACMA's stated enforcement priorities for 2025 and 2026 include persistent unwanted spam and telemarketing, with an explicit commitment to escalate against businesses that do not act on compliance alerts and early warnings.

The Privacy Act Does Not Care That the AI Is Someone Else's Product

An organisation using a commercial AI product remains responsible under the Australian Privacy Principles for the personal information it puts into that product. The OAIC said so directly in guidance published on 21 October 2024, covering chatbots, content-generation tools and productivity assistants. The OAIC's best-practice position goes further: organisations should not enter personal information, and particularly sensitive information, into publicly available generative AI tools at all. A human checkpoint helps here in a way that is easy to overlook. The gate is where somebody notices the automation pushed an entire contact record into a prompt when the draft only needed a first name and a suburb. The date to plan around is 10 December 2026. APP 1.7, inserted by the Privacy and Other Legislation Amendment Act 2024 (Cth), commences then. Where an entity arranges for a computer program to use personal information to make, or to directly support the making of, a decision that could reasonably be expected to significantly affect an individual's rights or interests, the privacy policy must disclose the kinds of personal information used and the kinds of decisions made. That provision is deliberately technology-neutral. It catches rule-based scoring tools and automated assessment systems as squarely as it catches language models, so "we only use templates" is not an exit. The OAIC will have infringement notice and compliance notice powers over failures, and civil penalties apply. If you are building anything that sends an automated assessment to a consumer, an appraisal, a quote, an eligibility answer, the disclosure needs writing well before December. See Privacy Act and Consumer Law for Australian software for the wider obligations.

Australian Consumer Law Applies to What the Machine Said

Section 18 of the Australian Consumer Law prohibits conduct in trade or commerce that is misleading or deceptive, or likely to mislead or deceive, and it applies regardless of intent. An automation that meant no harm can still put a business in breach. Section 18 carries no civil pecuniary penalty on its own. That reads like relief and is not, because conduct that breaches s 18 very often also falls within s 29 or s 30, which do carry penalties. Section 30 is the one to know if you touch property: it prohibits false or misleading representations in connection with the sale, possible sale or promotion of an interest in land, including representations about price. The exposure grew this year. The Treasury Laws Amendment (Doubling Penalties for ACCC Enforcement) Act 2026 commenced on 28 March 2026, doubling the maximum penalties available under the ACL. The maximum corporate penalty is now reported at A$100 million, up from A$50 million, with alternative measures based on the benefit obtained or on turnover. There is precedent for the property case specifically. In ACCC v Gary Peer & Associates Pty Ltd [2005] FCA 404, a real estate agent breached s 30 by advertising a house for auction with a price guide substantially below the vendors' actual asking price. Nothing about that reasoning changes if a model generated the guide. The practical read for a builder: any automated message containing a number a customer might rely on is a representation, and representations are the highest-risk thing your automation can emit.

Victorian Real Estate: Where the Checkpoint Is Effectively Mandatory

In Victoria, an automated price estimate sent to a residential buyer is a price representation regulated under the Estate Agents Act 1980 (Vic) and enforced by Consumer Affairs Victoria, which makes a human approval gate the practical minimum rather than a nice-to-have. The Statement of Information regime applies to every residential property offered for sale in the state. It must carry an indicative selling price, either a single figure or a range no wider than 10%. It must list the three most comparable properties sold, with address, date of sale and price. It must give the median house or unit price for the suburb, no more than 6 months old and based on a period of 3 to 12 months. It has to be displayed at every open for inspection, included with online advertising, and given to a prospective buyer within 2 business days of a request. Consumer Affairs Victoria publishes a penalty of more than A$48,842, being 240 penalty units, for non-compliance. An agent who sets an unreasonable estimated selling price, or advertises below it, may also forfeit their commission on the sale. The regime covers residential sales only; rural, commercial and industrial property are exempt. Enforcement is current, not theoretical. The Victorian Government's underquoting taskforce has monitored over 700 auctions and issued more than A$1.1 million in fines to estate agencies, and Consumer Affairs Victoria has laid criminal charges, including against a Melbourne agency over the estimated selling price on an Ivanhoe townhouse. We are deliberately not quoting section numbers for the Estate Agents Act. The underquoting provisions are widely cited in secondary commentary and we have not confirmed them against the legislation itself, so check legislation.vic.gov.au rather than a blog if you need one. The Act, the regulator and the penalty are the parts that drive the design decision anyway.

The Business Owns What Its Automation Says

In Moffatt v Air Canada, decided by the British Columbia Civil Resolution Tribunal on 14 February 2024, the airline argued its chatbot was a separate entity responsible for its own actions. The tribunal rejected that and found Air Canada liable for negligent misrepresentation. This is a Canadian small-claims decision. It is not binding in Australia and no Australian court is required to follow it. It is worth knowing because it states plainly the principle Australian regulators already apply: the business is the speaker, and the tool it chose to speak through is its own problem. Every Australian regime above lands the same way. ACMA writes to the sender. The OAIC writes to the APP entity. The ACCC and Consumer Affairs Victoria write to the trader who made the representation. In each case that is the business running the automation, not the vendor who supplied the model. Read your vendor terms with that in mind. Most AI platform agreements disclaim responsibility for output accuracy and place it on the customer, which is not a loophole being exploited; it matches where the legal liability was always going to sit. Deciding who owns this automation internally is part of the same question.

The Three Approval Patterns, and When Each One Fits

There are three patterns for putting a human in the loop, and the right one depends almost entirely on whether the action can be undone. | Pattern | How it works | Fits | |---|---|---| | Approval queue | The automation drafts the action, pushes it to a queue and waits for a decision | Anything irreversible or externally visible, including every customer email | | Confidence-based routing | A calibrated score sends high-confidence actions straight through and low-confidence ones to the queue | High-volume work where a genuine calibration signal exists | | Post-action approval | The automation acts, then surfaces the completed action for review inside a set window, with a compensating action if rejected | Reversible internal actions such as a CRM field update or a draft record | The word carrying the weight in confidence-based routing is calibrated. A score is calibrated when outputs it rates at 90% confidence are correct about 90% of the time. Most confidence numbers coming out of an AI system are not calibrated, they are a model's self-report, and a self-report is exactly wrong in the case you care about. An uncalibrated score routes confidently wrong output straight past the human and into a customer's inbox. If you have not measured your score against real outcomes on a held-out sample, you do not have a routing signal, you have a placebo. Post-action approval fails on customer messages for a simpler reason. It depends on a compensating action existing, and an email cannot be un-sent. A follow-up correction is a second message about a mistake, not an undo. Anything that leaves your systems and reaches a person belongs in pattern one.

How to Build a Gate That Does Not Become a Rubber Stamp

A checkpoint only works if the reviewer is shown enough to disagree with the draft. Five design choices decide whether that happens. Show the reasoning and the source data, not just the finished text. If the automation drafted a price estimate, put the three comparable sales, their dates and their prices next to the draft. A reviewer reading only prose checks grammar. A reviewer reading prose alongside the data that produced it checks the number, which is the thing a regulator would ask about. Design for batch review from day one. An automation that generates a hundred actions overnight needs grouping, filtering and bulk decisions, not a hundred separate popups. Reviewers who face a hundred popups start approving to clear the backlog, and you have built the rubber stamp yourself. Keep queues specialised. A single generic queue mixing pricing drafts, appointment changes and refund approvals forces every reviewer to context-switch and pushes them toward the fastest safe-looking answer. Route by expertise instead. Never auto-approve irreversible actions, whatever the confidence score says. Sending a message, posting publicly, moving money, hard-deleting a record. The category is defined by reversibility, not by risk appetite. Keep an audit log of who approved what and when, with the draft and the source data as they stood at approval time. This is the artefact you produce if Consumer Affairs Victoria asks an agent to justify how an estimate was determined, and it only exists if you built it before you needed it. One closing point worth stating plainly. A human checkpoint is not a sign the automation is immature or that the technology is not ready. In a regulated communication it is the design that makes the automation legally deployable at all. The gate is not the training wheels. It is the brakes, and nothing ships without them.

Frequently Asked Questions

Does a human checkpoint defeat the point of automating? No, because the expensive part was never the clicking. The automation still assembles the data, drafts the message, checks it against your rules and queues it in seconds. The reviewer spends a few seconds deciding rather than several minutes researching and writing. On a batch of fifty drafts with a shared review screen, the saving is most of the work, and you keep the legal position that a person approved every send. Can I skip the gate if the model is only summarising verified data? Grounding lowers the error rate substantially, with top-tier models reporting roughly 0.7% to 1.5% on grounded tasks, but it does not change who is liable. The Spam Act, the Australian Consumer Law and the Australian Privacy Principles all attach to the business sending the message regardless of how the message was produced. A low error rate is a good reason to review faster, not a reason to stop reviewing. What actually happens on 10 December 2026? APP 1.7 commences, inserted by the Privacy and Other Legislation Amendment Act 2024 (Cth). From that date, an entity whose computer program uses personal information to make, or directly support, a decision that could reasonably be expected to significantly affect someone's rights or interests must disclose in its privacy policy the kinds of personal information used and the kinds of decisions made. It is technology-neutral, so rule-based tools are captured too. Is confidence-based routing safe for customer emails? Only if the score is genuinely calibrated against measured outcomes, and in most builds it is not. An uncalibrated score routes confidently wrong drafts past the reviewer, which is the exact failure mode a gate exists to prevent. For outbound customer messages we would default to a hard approval queue and use routing for internal, reversible actions instead. Is the Air Canada chatbot case binding in Australia? No. Moffatt v Air Canada was decided by the British Columbia Civil Resolution Tribunal on 14 February 2024 and has no authority in an Australian court. It is quoted because the tribunal rejected the argument that a chatbot is a separate entity responsible for its own actions, and that is the same position Australian regulators take when they write to the business rather than the software vendor.

Build the approval gate before you build the send

buildAutomation works with your team on the real systems, designing the review queue, the audit log and the escalation rules alongside the automation itself, so the thing you ship is one a regulator and your staff can both stand behind.

See buildAutomation