03 Method
How we measure whether it works
With a test set and a number, which is the whole difference between this and a slide deck. The test set is the artefact the engagement rests on, and it remains yours.
Step 1: Define the correct answer
For each example, what your team would say is right. If two of your own people disagree, that disagreement is the first finding and it matters more than the model.
Days 1 to 3Step 2: Build the test set
Between 50 and 200 real documents or messages, covering the awkward cases rather than the tidy ones. A test set of clean examples measures nothing useful.
Days 3 to 7Step 3: Run and count
Accuracy per field, not overall. An invoice extractor at 95% overall can be 99% on the amount and 70% on the vendor name, and only the per-field view shows it.
Week 2Step 4: Read the failures
What kind of document it fails on, and whether prompting, retrieval or a validation rule fixes it. Most of the improvement comes from here rather than from a bigger model.
Week 2 to 3Step 5: Recommend, including not building
A written view with the accuracy, the cost per document and the review load it implies. We have recommended against building often enough for that to be a real outcome.
Week 3 to 4
04 Design
The human review threshold
The single most consequential setting in the whole system, and a business decision rather than a technical one.
| Threshold | Goes straight through | Reaches a person | Errors that escape |
|---|---|---|---|
| None | 100% | 0% | All of them, into your process |
| Low confidence only | 92% | 8% | The rare confident mistake |
| Conservative | 70% | 30% | Almost none, at a real review cost |
- The threshold is set from what an error costs, not from what looks good in a report. A misread invoice amount and a misrouted support ticket do not deserve the same setting.
- Everything below the threshold goes to a queue with a named owner and a service level, exactly as an approval workflow would.
- The review rate is monitored, because a rising share below the threshold is the earliest signal that the incoming documents have changed or the provider has updated the model.
- Every output is validated against a schema before it is accepted at all, so a malformed answer fails loudly rather than entering the process quietly.
The useful part was being told the vendor name field was only 70% accurate. We set that one to always review and let the amounts through, and the process worked from week one.
05 Governance
Where does your data go?
You get this in writing before a pilot runs, not after a procurement team asks. Five questions, five answers.
- Which provider, and which region. Named, with the processing location stated. Where data must remain in a named jurisdiction, that constrains the model choice and we say which options are then available.
- Whether your inputs are used for training. Answered per provider and per plan, because the default differs between consumer and business terms and the difference matters.
- What retention applies. How long the provider holds a request, and what we hold in logs. Both are configurable and both are decided before launch rather than discovered later.
- What personal data is involved. Sending personal data to a model is processing it, so the General Data Protection Regulation and its equivalents apply. We map the fields and remove what the task does not need.
- What never leaves. Often more than clients expect. Redaction before the request, retrieval over an index you host, and processing only the extract rather than the whole document are all ways of shrinking what is sent at all.
06 The people
Who do you actually work with?
Four founders, named. The hard part of an AI integration is not the model, it is deciding what accuracy is good enough and who reads the output when it is not.
The evaluation set is built before any model is chosen, and the people who build it are the people who report the numbers.
-
Dhanush Prabha
Co-Founder, CTO and CMO
Grounds the model in your data, wires it into the system of record, and scores it against a labelled set.
-
Sriram Ravichandran
Founder and CEO
Picks the one use case worth measuring and decides what accuracy has to be worth in money.
-
-
07 Cross-border
How do we work with clients in another country?
None of this needs you to be in any particular country. The work is remote either way, so these are the answers a buyer asks for before signing, and they are the same on every engagement we run.
Working hours
Our day runs on UTC+5:30. The overlap window with your team is written into the scope rather than assumed, and everything outside it runs asynchronously.
How we communicate
One written update a day on the channel you already use, a standing weekly call inside the overlap window, and a named person to escalate to. Nothing important is agreed only on a call.
Who you contract with
Synerdyn Private Limited, the company behind IncorpX, named in the agreement with its registration number. The governing law and the forum are agreed before you sign, not after a dispute.
Currency and payment
Invoiced in your currency or in ours, your choice, and settled by bank transfer. Milestones are tied to deliverables you can see, never to elapsed time.
What you own
Copyright in everything built for you is assigned on final payment: source code, design files, prompts, configuration and documentation. Third-party licences are listed by name so nothing is a surprise later.
Where data sits
You choose the region your data is stored and processed in, and the answer is written down before the build starts, with who at IncorpX can reach it and for how long.
Offset from our working day standard time
- London-5:30
- Dubai-1:30
- Singapore+2:30
- Sydney+4:30
- New York-10:30
- San Francisco-13:30
IncorpX is a brand of Synerdyn Private Limited. The contracting entity, the governing law and the invoicing currency are all named in the proposal before you sign anything.
08 Reference
Terms used on this page
- Large language model
- Written LLM: a model that produces plausible text continuations. Excellent at drafting and extraction, never a source of guaranteed correctness.
- Hallucination
- A confident output the source material does not support. Reduced by retrieval, thresholds and validation, and never eliminated.
- Retrieval augmented generation
- Written RAG: answering from your own indexed documents, with the passage shown so the answer can be checked.
- Test set
- A fixed collection of real examples with known correct answers, used to measure accuracy before and after any change.
- Human review threshold
- The confidence level below which output is routed to a person rather than accepted automatically.
- Token
- The unit a model provider bills by, roughly a fragment of a word. It is why AI cost scales with volume rather than with user count.
- Schema validation
- Checking an output has the expected shape and types before it is accepted, so a malformed answer fails loudly.
- GDPR
- The EU and UK data protection regime. Article 83 sets administrative fines of up to €20 million or 4% of worldwide annual turnover, whichever is higher, which is why consent, retention and access control are build decisions here rather than paperwork.
09 Questions
AI integration FAQs
Next step
Send the task and 50 real examples of it.
You get a measured accuracy figure per field, a cost per document, the review load it implies, and a written recommendation that is allowed to say do not build.
Read by Dhanush Prabha, our CTO, not a form queue. Reply usually within one working day, in your time zone.

