Skip to content
Praval Technologies

Case study

Every document said what it was. Nobody could read them all.

A mixed SharePoint estate held engineering drawings alongside leases, insurance certificates and vendor contracts, with no shared taxonomy. A classification service now reads each file the way a reviewer would and returns category, type, site and the evidence behind the call.

Document categories classified
7Document categories classifiedAcross 15 document types and 29 sub-document types
AI calls per file, adaptively
2→1AI calls per file, adaptivelyThe vision call is skipped when the text layer is enough
Files classified concurrently
8Files classified concurrently

Sample content: not published

This engagement is sample content. This page is excluded from the sitemap and search indexing until the details are confirmed.

Storage was never the problem. Knowing what was in it was

The documents already existed on SharePoint. What nobody could do was say what a given file was, which site it belonged to, or where it sat in a governance taxonomy, without opening it.

The estate was mixed in the way real ones are. Engineering drawings (electrical, mechanical, fire, plumbing, single-line diagrams, as-builts) sat in the same libraries as insurance certificates, leases, financial reports, warranties, vendor contracts and emergency procedures.

Manual tagging does not survive that, and the reasons are specific:

  • No consistent taxonomy. No shared classification standard was applied across files or reviewers, so the same document could be labelled three different ways.
  • Site names never matched. One facility might appear as a folder code, a building code, a street address or a facility name, with nothing tying any of them back to a region, market or site manager.
  • The evidence was in the image. Title blocks, revision stamps, As-Built marks and equipment tags are what actually identify a drawing. A text search will never surface them, because they are not in the text layer to begin with.
  • A confidence score is not enough to act on. A classification had to cite the evidence behind it, or the reviewer would re-open the file and check it anyway.

Reading a document is cheap. Reading every document is not

The obvious design sends every file to an image-understanding model. It works, and it is too expensive to run across a whole library; most documents are text-rich and need nothing of the sort.

So the classifier decides for itself whether looking at the page is worth the call. A lightweight check runs first: if the PDF has a real text layer, that text clears a 2,000-character floor, and neither the filename nor the folder path signals an engineering drawing, the vision call is skipped and classification runs on extracted text alone.

The conditions are conjunctive by design. Any one of them failing routes the file through vision, so scans, CAD files and anything drawing-like are always looked at. The gate is deliberately biased toward spending the call: a missed revision stamp costs more than a redundant API request.

One file path in, a structured and defensible answer out

The service signs in to SharePoint on its own, resolves the site, subsite and library from the path, and downloads the file through Microsoft Graph, with cached lookups for repeat sites and a fallback search across every drive when a direct path does not resolve.

Classification is not left to the model's judgement alone. A single master reference, kept in one module, defines 7 categories, 15 document types and 29 sub-document types, each with a written description of what counts as a match. A strict priority order sits on top: insurance and leases are checked first, because their legal language is unambiguous, and drawings last, where a second tiebreak puts single-line diagrams above As-Built status.

Site clues are resolved rather than recorded. Whatever is available (an address from a title block, a facility name, a folder path as a last resort) is matched against two governance workbooks by direct code first, then fuzzy address and city overlap. Anything below the threshold comes back as Unknown rather than guessed at, which keeps a wrong site code out of the index. That is the one error that would quietly corrupt every report built on top of it.

The response is flat JSON: category, document type, sub-document, site and a confidence score, with the reasoning that cites the evidence behind the call.