Most businesses that "try AI" do the same thing: someone opens a chat window, pastes in a problem, gets a plausible answer, and nothing about the business changes. An AI agent is different. It has a job, it knows how your business works, it can open your files and use your tools, and it hands anything risky back to you for approval. Set up properly, it does work; set up badly, it does damage.
This is the method we use to set up AI agents in a business. It is our own delivery playbook, the one we follow on client engagements and the one we used on ourselves first, written out for an owner or manager to follow. It is the same method behind our Claude consulting and implementation service. It has ten steps. Most of them are not about AI at all.
The value is not "we set up an AI". It is that the business now has a written description of how it runs, who does what, and what must never happen, that an AI can act inside. The AI part is comparatively easy.
The model in one page
An AI operating layer in a business has four parts. Every step below builds one of them.
Two principles hold it together. Single source of truth: every fact is written in exactly one place and everything else points at it, because the moment a fact exists in two documents they will disagree and nobody will know which is right. Structure before automation: never automate a process nobody has written down.
Step 1 — Qualify: is this a real problem, and is AI the answer?
Before any set-up, establish four things: what the business actually does, in the owner's words rather than the website's; where the pain is, meaning the task that eats hours, gets forgotten, or costs money when it goes wrong; whether AI is genuinely the right answer; and who owns the decision and who will actually use the thing once it is built.
Sometimes the honest answer is "your real problem is that three systems don't talk to each other, and that's an integration job, not an AI one." Say so. The questions that separate a real project from a curiosity:
- What happens today if this task doesn't get done?
- Who does it now, and how long does it take them?
- Where does the information live: email, a spreadsheet, someone's head?
- What would you have to see in 90 days to call this a success?
- What must never happen? (This becomes a hard rule in Step 8.)
Step 2 — Discovery: map how the business runs
This is the step clients underestimate and benefit from most, even before an agent runs. You are producing the raw material for the instruction layer. Capture:
- The business. Legal entity, trading name, ABN, where it operates, who the customers are, what it sells, and how a job flows from enquiry to invoice.
- The people. Who does what, by function rather than by name (functions outlive staff). Who approves spend, who signs, who talks to customers. Where the bottleneck sits, which is usually one person who is the only one who knows something.
- The systems. Every tool in daily use, which one is authoritative for which data (only one can be), and which integrations are really someone re-keying data by hand.
- The rules. Regulatory obligations, licence conditions, professional standards, internal policy on what needs approval and what must never be automated, confidentiality and privacy constraints.
- The language. Their jargon, abbreviations and job names. Get this wrong and every output reads as written by an outsider.
The most valuable hour of discovery is usually watching someone do the task rather than asking them to describe it. People describe the process they think they follow, not the one they actually follow.
Step 3 — Build the workspace
A single root folder is the business's world. Everything the agent needs to know is inside it or reachable from it. For most businesses that is the cloud drive they already use, so it syncs to everyone's machines. Inside it: the instruction document at the root, a Projects folder with one sub-folder per job or client, a Templates folder for the documents the business reuses, and a Reference folder for the standards, price lists and policies the agent should cite.
Then the single highest-leverage habit: one session per project, launched from that project's folder. A session started at the business root is head office. A session started inside a client's folder is that engagement and only that engagement. Context stays clean, work runs in parallel, and confidentiality becomes structural rather than a matter of remembering, because one client's material simply is not in scope when you are working on another's.
Decide early, and write down, what the agent may read directly, what must be exported to it, and what it must never touch. "It can see everything" is not an answer; it is a liability.
Step 4 — Write the instruction layer
This one document is the engagement; everything else is scaffolding around it. It loads automatically at the start of every session, so it must be complete, current and short enough to be read. The structure that works: one paragraph on what the workspace is and the agent's role; the business (entity, customers, lines of work); the team of co-workers and the routing rules for who handles what; how work is organised; a map of the folders; the infrastructure facts that are tedious to look up and expensive to get wrong; and the hard rules.
Writing standards that matter more than they sound:
- Write decisions, not descriptions. "The invoice must be approved by the director before it is sent" beats ten sentences of background. State the rule and the consequence.
- Mark the deliberate things as deliberate. Where a choice looks like an inconsistency but is intended, say so, or a well-meaning agent will "fix" it.
- Include the failure mode. Say what breaks when it is done wrong, and how to recover.
- Be specific about identity. Legal entity versus trading name versus the name on the sign. Getting this wrong puts the wrong company on a contract.
- Keep it current or delete it. A stale instruction is worse than a missing one, because it is followed with confidence.
Step 5 — Design the co-workers
Create a co-worker when a body of work has its own standards, its own judgement and its own hand-off boundaries: a distinct professional discipline such as legal, design or accounting; work that needs a different tone or standard of care; a domain where the wrong answer has a different kind of consequence. Do not create one for every task. If it differs from an existing role only by subject matter, or you cannot state what it must hand off and to whom, it is a task, not a co-worker.
Each co-worker is one short file: who it is and what it owns, the hard limits at the top if the discipline has any (legal is not a solicitor; finance does not move money), enough context about the business to work without asking, the specific things it does, the standards it works within, the house rules restated in its own terms, and who it hands off to.
Give them names. A co-worker with a first name gets addressed like a colleague, "Jason, draft the NDA", and that changes how people use the system. It stops being a tool you operate and becomes a team you delegate to. It also makes hand-offs legible. Two sentences like "Stacey owns how it looks. Cameron owns getting it live" prevent months of ambiguity, and that scope boundary is the real design work.
Step 6 — Write the prompts like you would brief a new hire
The instruction layer and the co-worker files are the prompts. Lead with the role and what it owns. State the non-negotiables early and plainly; constraints buried in paragraph nine get followed less reliably than constraints under the first heading. Prefer rules for behaviour and examples for judgement: "never quote a price" is a rule, while knowing when a design is finished needs an example. Write the standard, not the aspiration: "high quality" means nothing, but "a tidy centred layout with a stock gradient is a fail, not a baseline" is actionable. Say what to do when uncertain: ask, assume and flag, or stop, and pick one per role. Then test it by handing it to someone new. If a competent person could not act on it, the agent cannot either.
Step 7 — Add memory
Sessions end; the business continues. Memory is what stops every session starting from zero. Keep one fact per file, not one file per topic, so each can be updated or deleted without collateral damage. What belongs: preferences and working style that took effort to learn, decisions and their reasoning (especially ones that look odd without context), project state that is not derivable from the files, and corrections, so the same mistake is not made twice. What does not: anything the files already record, anything that only matters to today's conversation, and anything that will be false next week unless it is dated. Convert "last Tuesday" to an actual date. Keep a one-line-per-fact index that loads every session, and never put the facts themselves in the index.
Step 8 — Track the work in one place
The agent's task list and the business's task list must be the same list. Two lists means work gets lost in the gap. Give each project a short code and number the tickets so work is referenceable in conversation, the way people actually talk about jobs. If the business already has a system, integrate rather than replace; a parallel system nobody updates is worse than none.
Step 9 — Guardrails (non-negotiable)
These go into every instruction document we write, adapted to the business. They are the reason an owner can let an agent near their operations at all.
- Never handle credentials. No passwords, API keys, tokens or card numbers entered into any field, and no account creation. The agent sets everything else up and leaves the one credential step to a human. A client will test this, and passing that test is worth more than any feature.
- Nothing irreversible without a human's OK. Publishing, sending client email, spending money, deleting data, going live. Prepare it, then ask. Approval for one action is not approval for the next.
- Honesty in anything customer-facing. No invented statistics, testimonials, case studies or logos. Illustrative examples labelled as examples.
- Confidentiality between clients. One client's material never enters another's work. One session per project makes this structural.
- Licensed assets only. Imagery, music and fonts licensed for commercial use or owned by the business. Releases for identifiable people; always parental consent for children.
- Local context. AUD, DD/MM/YYYY, Australian spelling. Small, and immediately obvious when wrong.
- Sector rules on top. Insurance, finance, medical, legal and construction each add obligations. Capture them in discovery and write them in.
Then test each rule deliberately before handover: ask the agent for a password and watch it refuse; ask it to send an email and watch it come back for approval; work two projects and check nothing crosses over.
Step 10 — Hand over, then operate and improve
An implementation nobody uses is a failed implementation. Train the habit, not the tool. The three things the team must leave knowing: launch the session from the folder the work is in; address the co-worker by name for the job you want; and anything irreversible will come back to you for approval, which is the design, not a limitation. Show them a real job end to end on their own live work, because a demo on invented data teaches nothing and they can tell. Hand over the documents, name an owner inside the business who keeps the instruction layer current, and book the 30-day and 90-day reviews before you leave.
At each review ask: what did it actually get used for (often not what was scoped; follow the real usage), where did it get things wrong and is that an instruction gap or a scope gap, what has changed in the business that the instructions do not know yet, which manual step is now the bottleneck, and is anything in memory now false. Instruction gaps are the common finding. The agent did what it was told; it was told the wrong thing, or nothing. Fix the file, not the prompt of the day.
The one lesson under all ten steps
When we reviewed our own set-up against this playbook, every problem we found was the same problem: a fact written down more than once. A team roster in three documents that disagreed. A list of names that had to be updated by hand in two places. Build the structure so each fact has exactly one home and everything else references it. That is the difference between a system that stays true and one that quietly rots.
If the "task" you are trying to automate is a specific recurring job rather than a whole operating layer — quoting, invoice chasing, follow-ups — our AI automation service is usually the faster starting point.