How do you run an AI proof of concept in 4 weeks?

You run an AI proof of concept in 4 weeks by narrowing it to one question, agreeing in advance what result counts as success, building an evaluation set from your real data in week 1, measuring two or three approaches against it in week 2, testing a small pilot with real users in weeks 3 and 4, and ending with a written go or no-go. A PoC is a measurement, not a demo.

Most AI pilots that stall do so for predictable reasons: nobody agreed what success meant, the test used clean examples instead of real ones, or the running cost was discovered after launch. Each of those is cheap to prevent in week 0 and expensive to fix in month 6. This guide is the playbook we follow; if you want us to run it with you, see our AI proof of concept service.

When Goal Output
Week 0 One question, success threshold, data access A one-page PoC brief
Week 1 Evaluation set and first measurements 30-100 real examples, baseline results for 2-3 models
Week 2 Tune and analyze failures Accuracy, latency, cost per request; preliminary go or no-go
Weeks 3-4 Pilot with real users Usage feedback, final report, decision

What question should an AI PoC answer?

An AI PoC should answer one question in this form: can this AI approach do this specific task on our real data, accurately enough and cheaply enough to be worth building? "Can AI help our support team?" is too broad to test. "Can a model draft first replies to billing tickets that a support lead accepts without edits at least 8 times out of 10?" can be tested in four weeks.

Good PoC candidates share a few traits. The task is done by people today, so there are real examples and someone who can judge outputs. Accuracy matters enough that you need a number before investing. The data is messy or unusual, so it is not obvious that it will work. Typical examples: extracting fields from scanned documents, classifying and routing inbound requests, answering questions over internal documents, drafting replies, or an agent that completes one workflow through your tools.

You probably do not need a PoC when the task is well understood and low risk, such as summarizing support threads. Then build it directly. If you have several ideas and no clear first one, rank them first; our AI audit does that for an existing product.

How do you set success criteria before you start?

Write the success criteria down before any code, with a number for quality, a limit for cost and, where users wait for the answer, a limit for latency. Get the business owner of the process to agree. A threshold set after seeing results is not a threshold.

Task type Quality criterion Cost and speed criteria
Document extraction At least 90-95% of fields correct on the eval set Under a set cost per document; batch processing acceptable
Classification and routing At least 90% correct category; known errors are cheap Under a set cost per item
Drafting replies 8 of 10 drafts accepted without edits by a domain expert Answer in under 10 seconds
Questions over documents 85-90% answers correct and supported by a cited source Under 5 seconds to first words
Agent for one workflow Completes 8 of 10 test cases with no unsafe action Under a set cost per completed task

The numbers above are typical starting points, not rules. Use what the business case needs: if a wrong extraction costs a refund, aim higher; if a human reviews every output anyway, the bar can be lower because the value is time saved. Also define what "correct" means for each output type, ideally with two or three graded examples, so different reviewers judge the same way.

How do you build the evaluation set?

Build the evaluation set from 30 to 100 real inputs, each paired with the output you would expect or with grading criteria. This set is the most valuable thing the PoC produces: it keeps measuring quality through every later change of prompt, model or data pipeline.

How to collect it:

  1. Sample, do not hand-pick. Take recent real cases at random, then add known hard ones: scans, missing fields, mixed languages, ambiguous requests.
  2. Write the expected output with the person who does the task today. For open text, write grading criteria instead of a single right answer.
  3. Split it. Keep about a fifth aside and do not look at it while tuning; score it only at the end, so you know the result is not overfitted.
  4. Mask sensitive data if needed. Anonymized data is fine for a PoC, as long as the structure and difficulty stay real.
  5. Write the scoring script in week 1: exact match for structured fields, a rubric graded by a second model with human spot checks for open text.

What happens in each of the four weeks?

Week 1: data and baseline

Get data access, build the evaluation set and run a simple baseline with two or three models, for example one from OpenAI, one from Anthropic and an open model if data must stay on your servers. Log every call with token counts from day one. By the end of the week you know roughly how far each model is from the threshold. Our comparison of the Claude and OpenAI APIs explains the main differences between providers.

Week 2: tune and analyze failures

Improve the best one or two approaches against the evaluation set: better prompts and examples, structured output, retrieval if answers depend on documents (see our RAG implementation guide), or a different pipeline. Then read every failure and group them by cause. By the end of week 2 you have accuracy, latency and cost per request, and usually a preliminary go or no-go.

Weeks 3 and 4: pilot

If the numbers are promising, wrap the best approach in a minimal interface or API with logging and cost limits, and let a few real users try it on real work. A pilot answers what an evaluation set cannot: do people use it, do they trust it, and does it save time. Finish with the final scoring on the held-out examples and the report.

How do you make the go or no-go decision?

Make the decision against the criteria from week 0, with the failure analysis next to the numbers. There are more than two possible outcomes, and naming them avoids stretching a weak result into a project.

Result Decision Next step
Meets quality, cost and speed thresholds; pilot users keep using it Go Production build with permissions, monitoring and rollout
Close to threshold; failures cluster in one fixable cause Go with conditions Fix the cause (data, narrower scope, human review), then build
Works only on a subset of cases Narrow go Build for the subset, route the rest to people
Far from threshold or too expensive per request No-go Document why; consider a non-AI solution or revisit later

A no-go is a valid, useful result. In the AI Grief Companion we built for a US startup, prompting alone could not reproduce one specific person's voice, which led to LoRA fine-tuning on Qwen2.5-7B. That is exactly the kind of question a PoC settles early, before the full build is planned around the wrong approach.

What does an AI proof of concept cost?

Cost depends on what the PoC must prove and whether it runs on real data and real systems. With us, an AI PoC runs as one AI Sprint: $10,000 fixed for up to 4 weeks, including scoping, the evaluation set, the prototype, model comparison, the pilot, the report and a handover call. Model and hosting usage during the PoC is paid directly by you and is usually small; we estimate it up front.

If the answer is go, a working AI MVP typically takes 8 to 14 weeks more. When comparing quotes from other vendors, check what "proof" means: a demo on hand-picked examples and a deployed prototype measured on real cases are both called a PoC, and they are not the same purchase. Our guide to fixed-price AI development explains how to compare them.

What should you get at the end of an AI PoC?

At the end of a well-run AI PoC you should own five things, whether or not you continue with the same team:

  • Prototype code in your repository, written so the core can carry into production.
  • The evaluation set and the script that scores any future version against it.
  • A results report: accuracy on the held-out set, failure patterns, latency, cost per request and estimated monthly cost at your volume, for each model tested.
  • An architecture note for the production version: data flow, guardrails, integration points and open risks.
  • A written go or no-go with the reasoning and, if go, the scope of the next stage.

Our verdict

Run an AI PoC when the outcome is genuinely uncertain and a wrong bet would be expensive. Keep it to one question, set the threshold before you start, use real data including the ugly cases, measure cost from the first call, and write the decision down. Four weeks is enough for most single-task PoCs, with the answer usually visible by the end of week 2.

Next step

Describe the task and send a few real examples. In a 30-minute call we will tell you whether a PoC is the right step, what the threshold should be and what data we need. See our AI proof of concept service and the AI Sprint, or contact us.

Case studies

Frequently asked questions

For a 4-week PoC, 30 to 100 real examples with the outputs you would expect are usually enough to see whether the approach works and where it fails. Include hard and ugly cases, not only clean ones. Before a production launch, grow the set to 100-200 examples and keep adding failures you find in real use.

A good criterion is a number on your own data that a business owner agrees with before work starts, such as at least 90 percent of extracted fields correct, 8 of 10 drafted replies accepted without edits, or answers under 5 seconds at under a set cost per request. Include a quality threshold, a cost limit and, where relevant, a latency limit.

A proof of concept answers whether the AI approach works on your data at acceptable accuracy and cost, measured on an evaluation set. A pilot puts a working version in front of real users to see whether they use it and whether it saves time. A 4-week plan can include both: the PoC answer by about week 2 and a small pilot in weeks 3 and 4.

Then it did its job at a fraction of the cost of a failed build. A good failure report says how far results were from the threshold, which inputs broke the approach, and what would need to change: more or cleaner data, a narrower task, human review of every output, or a solution without AI. Sometimes a narrower version of the idea passes and is worth building.

With us, an AI PoC runs as one AI Sprint: $10,000 fixed for up to 4 weeks, including the prototype code, the evaluation set, a results report and a go or no-go recommendation. Model and hosting usage during the PoC is paid directly by you and is usually small. When comparing quotes, check whether the PoC runs on your real data with measured results or is a demo on hand-picked examples.

Let’s start your project
Book a call