codexier.

Pricing & Buying

Budgeting for a First AI Project

By CodexierPublished 6 min read

Most first AI projects in small companies fail in a predictable way: the budget goes into building, nothing is set aside for checking the output, the running cost was never estimated, and nobody decided what result would count as failure. This guide splits the budget into four lines, discovery, pilot, running and review, sized so that the pilot proves or kills the idea for a modest sum, and gives the rule for when to stop funding.

Discovery before building

Discovery answers four questions before any money goes into building. Which task is being automated, precisely, with its inputs and outputs. What data the company actually has, where it lives and whether it can legally be used. What a correct result looks like, so the pilot can be judged. And what the task costs today in hours and errors, which is the number every later decision is compared against. This is a short, fixed-price piece of work; our AI integration audit is exactly this step, priced on the pricing page, and it often ends by recommending against building, which is the cheapest possible outcome.

A pilot sized to learn

A pilot is not a small production system. It is a controlled test on one task with real inputs, run long enough to see the failure modes, with every output reviewed by a person. Sizing it means deciding those four things and nothing else.

Pilot decisionRight-sizedToo big
ScopeOne task, one team, one channelEvery enquiry type across the company
DurationFour to eight weeksOpen-ended until it works
DataA sample of real cases from the last monthsA new data platform built first
Success measureSet before start: accuracy, time saved, cost per caseDecided after seeing the results
IntegrationManual or lightweight; copy and paste is fineFull integration with every system
ReviewA named person checks every outputTrust the system from day one

Examples of well-sized pilots: answering one category of email, drafting quotes from a form, classifying incoming invoices, handling bookings by phone for one location as in our restaurant phone answering guide.

Running costs: usage and upkeep

AI systems have two running costs that a normal software budget does not. Usage is the metered cost of the model calls, which scales with volume and with how much text or speech each case involves. Upkeep is the human work of keeping the system accurate: updating the documents it answers from, adjusting prompts when the business changes, and re-checking output after model updates. Both continue for as long as the system runs.

  • Estimate usage from the pilot: cost per case multiplied by monthly volume, with a margin for growth and longer cases.
  • Prefer products with a fixed monthly fee where volume is predictable; metered pricing rewards low usage and punishes success.
  • Budget upkeep as hours per month from someone who knows the business, not as an afterthought for IT.
  • Include the ordinary software costs too: hosting, integration, logging, and the licence for any platform the system runs on.

Human review time

Review is the line most first budgets forget and the one that decides whether the project is safe. During the pilot, someone checks every output against the success measure and records the misses. After launch, review shrinks to a sample plus every case the system flags as uncertain, but it never reaches zero. The person doing it must know the task well enough to spot a plausible wrong answer, which usually means an experienced employee, not an intern.

During the pilot

Full review of every case. Budget the reviewer's hours as a project cost; it is the same money as the build.

After launch

Sampled review on a schedule, plus escalations. Tie the sample size to the risk of a wrong answer reaching a customer.

Escalation design

The system must be able to say it does not know and hand over. Cases it hands over are not failures; they are the review budget working.

Records

Keep the review log. It is the evidence for the stop-or-continue decision and, for anything touching personal data, part of your GDPR documentation.

When to stop funding

Write the stop rule before the pilot starts, when nobody has anything to defend. It has three parts: the measure, the threshold and the date. For example: the system must handle a stated share of cases without human correction, at a cost per case below the manual cost, by the end of week eight. If the threshold is not met by the date, the project stops and the discovery findings are kept for another attempt later. If it is met, the running budget is approved and the pilot becomes production, with review continuing. Companies that skip this step end up funding a system that is almost good enough for years.

When you do not need an AI project: if the task is rare, the manual cost is small, or the data does not exist in usable form, the discovery step will say so and the budget stops there, which is the point of doing discovery first. If you have a task in mind and a rough idea of what it costs today, book a free 15-minute call and we will tell you whether it is a candidate for a pilot and what the four-line budget would look like; for a support-desk use case our chatbot setup is the usual next step.

Frequently asked questions

How big should a first AI budget be?

Small enough that a negative result is affordable and large enough to run a real pilot with review. The discovery step is a fixed price, the pilot is a few weeks of work plus reviewer time, and the running cost is estimated from the pilot before you commit to it.

Should we build with an off-the-shelf tool or custom?

For a pilot, use the lightest tool that lets you test the task with real data. Custom work is justified after the pilot, when you know the task is worth it and what the integration must do.

What about GDPR in the pilot?

Use real data only where you have a legal basis and a processor agreement with the AI provider, prefer EU-hosted services, and keep personal data out of prompts where the task allows. Discovery should include this check so the pilot does not start with a compliance problem.

Who should own the project internally?

The person who owns the task being automated, not IT. They know what a correct result looks like, can do the review and will notice when the business changes in a way that breaks the system.

Have a task you think AI could take?

Describe the task and what it costs you today. In fifteen minutes we tell you whether it is a pilot candidate, what discovery would cover and what the stop rule should be.

Book a free 15-minute call