03 Method
How we measure whether it works
With a test set and a number, which is the whole difference between this and a slide deck. The test set is the artefact the engagement rests on, and it remains yours.
Step 1: Define the correct answer
For each example, what your team would say is right. If two of your own people disagree, that disagreement is the first finding and it matters more than the model.
Days 1 to 3Step 2: Build the test set
Between 50 and 200 real documents, covering the awkward cases rather than the tidy ones. A test set of clean examples measures nothing useful.
Days 3 to 7Step 3: Run and count
Accuracy per field, not overall. An extractor at 95% overall can be 99% on the amount and 70% on the vendor name, and only the per-field view shows it.
Week 2Step 4: Read the failures
What kind of document it fails on, and whether prompting, retrieval or a validation rule fixes it. Most improvement comes from here, not from a bigger model.
Week 2 to 3Step 5: Recommend, including not building
A written view with the accuracy, the cost per document and the review load implied. We have recommended against building often enough for that to be a real outcome.
Week 3 to 4
The threshold decides everything downstream
04 Governance
Where does your data go?
You get this in writing before a pilot runs, not after a procurement team asks. Five questions, five answers.
- Which provider, and which region. Named, with the processing location stated. Where data must remain in India, that constrains the model choice and we say which options are then available.
- Whether your inputs are used for training. Answered per provider and per plan, because the default differs between consumer and business terms.
- What retention applies. How long the provider holds a request and what we hold in logs. Both are configurable and both are decided before launch.
- What personal data is involved. Sending personal data to a model is processing it. The Digital Personal Data Protection Rules, 2025, notified on 13 November 2025 by the Ministry of Electronics and Information Technology, apply, with the main obligations commencing on 13 May 2027.
- What never leaves. Often more than clients expect: redaction before the request, retrieval over an index you host, and processing only the extract rather than the whole document.
05 Reference
Terms used on this page
- Large language model
- Written LLM: a model producing plausible text continuations. Excellent at drafting and extraction, never a source of guaranteed correctness.
- Hallucination
- A confident output the source material does not support. Reduced by retrieval, thresholds and validation, and never eliminated.
- Retrieval augmented generation
- Written RAG: answering from your own indexed documents, with the passage shown so the answer can be checked.
- Test set
- A fixed collection of real examples with known correct answers, used to measure accuracy before and after any change.
- Human review threshold
- The confidence level below which output is routed to a person rather than accepted automatically.
- Token
- The unit a model provider bills by, roughly a fragment of a word. It is why AI cost scales with volume rather than user count.
- Data Protection Board of India
- The adjudicating body under the Digital Personal Data Protection Act, 2023. The Schedule to that Act sets fixed rupee ceilings rather than a share of turnover: up to ₹250 crore for failing to take reasonable security safeguards against a personal data breach, up to ₹200 crore for failing to notify the Board and the people affected, and up to ₹50 crore for a breach of any other provision. That is why consent, retention and access control are build decisions here rather than paperwork.
06 Questions
AI integration in Dehradun: FAQs
Next step
Send the task and 50 real examples of it.
You get a measured accuracy figure per field, a cost per document, the review load it implies, and a written recommendation that is allowed to say do not build.
Read by Dhanush Prabha, our CTO, not a form queue. Reply usually within one working day, in your time zone.

