The first fortnight of an AI pilot is when you can still change course cheaply. After that, people start defending the work they have already shown upstairs.
We typically recommend a focused four to six week proof of value: one workflow, one team, measurable outcomes. The first two weeks should already show whether that shape is real.
Signs the pilot can graduate
Someone can name the workflow in a sentence. Ticket triage, first-pass reports, reconciling two records. If the answer is still "we are exploring AI", you do not have a pilot yet.
There is a baseline written down before anyone prompts a model. Hours per week on the current process, error rate, time to first response. We hold AI work to numbers like these. If they do not move, the pilot does not graduate.
A human reviewer is designed in from day one. Where the system pauses, what the reviewer needs to see, and how those decisions feed back. Build the agent around that path. Patching it in after a security review is late.
There is an evaluation set of real tasks, even a small one. Without it you cannot tell if a model change made things worse. We build that harness before an agent touches production.
Someone owns the output. A named person in the operation who will live with the result. We have seen plenty of AI work fail because the sponsor changed roles, the people using the tool did not trust it (often correctly), the training never happened, or nobody was accountable for what the system produced.
Data residency and privacy are answered in week one. Australian businesses have Privacy Act duties and, in some sectors, extra rules such as APRA CPS 230. If nobody can say where the data goes, the pilot is not heading for production.
Signs it is becoming a slide deck
The work started with a model, then looked for a home. Start with the workflow. Ask where intelligence changes the economics of a process.
The only users are the project team. Operators have not sat with it. We start by shadowing the people who do the work today.
There is no escalation path. The demo answers every question. A production system needs a place to stop and a person who can override it.
Spend, latency, accuracy, and drift are not being watched. Token cost belongs in the operating cost of the AI.
The engagement will end with a strategy document and no offer to build the first thing, or to brief whoever will. A plan still has to name the first production workflow.
What to do in week two if it is drifting
Narrow the scope to one workflow and one team. Write the baseline. Put a reviewer in the loop. If those cannot be done in a few days, pause the pilot.
A 20-minute call is enough to walk through where you sit. Book a discovery call if you want that conversation.