Legalnaut

EVIDENCE 2026-08-28 2 MIN READ

Why your served bundle is invisible to search

Scanned material has no text in it. Until it does, the strongest document in the file cannot be found by anyone preparing the case.

You search the file for a word you are certain appears in it, and get nothing back. The document is there. The word is on the page. What is missing is any text for the search to match, because the page is a photograph of a page.

A scan is a picture, not a document

A PDF can hold two very different things. One is text with layout information around it. The other is an image of a piece of paper, wrapped in a PDF container so it can be emailed.

Both open in the same viewer and look identical to a reader. Only one can be searched. Court files are full of the second kind: material photocopied, faxed, scanned by an office machine set to whatever it was set to, and served as a single large file.

The failure is silent, which is the problem

Nothing warns you. Search returns no results, which looks exactly like search returning no results because the word is not there. The document sits in the file, indexed by filename only, and whoever is preparing the case works around a hole they do not know exists.

This is the part worth internalising: an unsearchable document in a large file is functionally a missing document, and it goes missing without any error message.

What OCR actually gives you

Optical character recognition reads the image and produces text. On a clean scan of a typed page the result is close to perfect. Quality falls off with generation loss, so a fax of a photocopy of a fax produces fragments, and handwriting in the margin usually produces nothing.

Two practical consequences follow.

First, run it on import rather than on demand. OCR at the point a document enters the file means the file is searchable from the start. OCR when someone notices a gap means the gap has already cost something.

Second, keep the original. The OCR text is for finding the document. The document is what gets quoted, exhibited and put in front of the court.

Knowing what failed is half the value

A page of ligature soup is not text, and treating it as text is worse than treating it as nothing. It pollutes search results and it will eventually be quoted by something that cannot tell the difference.

What you want from the process is a clear answer in two parts: here is the text we could recover, and here is the list of documents where recovery failed. That second list is a work queue. Those are the pages a person has to open, and there are usually far fewer of them than the size of the bundle suggests.

Where this sits in the workflow

In Legalnaut, extraction runs by content rather than by file extension, because served material arrives named badly or not named at all. A local parser tries first. When what comes back is too short or too broken to be real text, the document goes to OCR, and audio goes to transcription with timestamps kept so a quotation can be checked against the recording.

Documents where nothing usable came back are listed as such rather than left to look like empty documents. That list is short and it is worth an hour.

Legalnaut does this to your own case file: a chronology built from the documents, every finding showing the source it came from. See the plans.