More AI for local historians

For the latest addition to the Preston History Library, History of Preston County Fire Brigade from the Guild Merchant of 1702 to the Guild Merchant of 1952, I used AI to scan the pdf file to make it searchable and to compress the file to make it easy to download. I thought it might be of interest to document the steps that AI followed.

I use the free version of Claude AI, but Chat GPT would probably work just as well. My son uses Claude on a daily basis for work and in discussing it with him, I realise I am only just scratching the surface of the tools that AI can offer for local history research.

AI is especially useful for transcribing handwritted documents, including providing useful first drafts of documents in 17th-century secretary hand, which can baffle even those skilled in palaeography.

Here are the steps that AI followed


If I have a pdf file on my computer, can you transcribe the text?

Determined straightforward response and requested file upload
Determined straightforward response and requested file upload

Yes — if you upload the PDF here, I can transcribe the text from it, including handwritten material if it’s legible enough (which is often relevant given the manuscripts and archival documents you work with).

A few things worth knowing:

  • Printed/typed text transcribes very reliably.
  • Handwritten text (like Victorian-era memoirs, letters, or archival documents) I can usually read too, though quality depends on the handwriting and scan clarity — I’ll flag anything I’m unsure about rather than guess.
  • I’ll preserve original spelling and punctuation as written, since I know that’s your preference for historical transcriptions, unless you’d like it modernized.
  • If the PDF is a scan of a printed source rather than a text-based PDF, I read it visually rather than trying to copy machine text, so image quality matters.

Go ahead and upload it whenever you’re ready, and let me know if there’s anything specific about the transcription (e.g., flagging uncertain words, formatting into paragraphs vs. lines as in the original, etc.).

Are you also able to compress a pdf file?

Thought for 2s
Thought for 2s

Yes, I can compress a PDF for you. Upload the file and let me know if you have a target (e.g., under a certain file size for email, or just general size reduction), and I’ll take care of it.

Is it best to transcribe first and then compress or the other way round?

Weighed transcription priority over compression for quality preservation
Weighed transcription priority over compression for quality preservation

Transcribe first, then compress.

Compression usually works by reducing image quality/resolution to shrink file size, and for handwritten or faint text that can mean the difference between me being able to read it clearly and misreading or missing words. Better to work from the highest-quality version of the file while transcribing, then compress afterward once you no longer need to extract anything further from the images.

If you’re just archiving a document you’ve already fully transcribed, order won’t matter — compress away.

Preston_Fire_Brigade_history.pdf
Please transcribe the above file and then compress the file so that users of my website can download it efficiently

Ran 4 commands, viewed 3 files
Ran 4 commands, viewed 3 files

Good — this is 50 scanned pages (grayscale, ~236 ppi), which explains the 92MB file size. Let me look at the first page to see whether it’s printed or handwritten text.

Viewed a file, ran a command
Viewed a file, ran a command

Good, tesseract is available. Let me run OCR across all 50 pages, then spot-check accuracy visually.

Ran 4 commands
Ran 4 commands

All 50 pages OCR’d. Now let me spot-check a few pages visually against the OCR output to gauge accuracy before presenting the transcription.

Weighed OCR cleanup strategy and output formatting approach
Weighed OCR cleanup strategy and output formatting approach

Now let me build a proper Word document for the transcription, since this is a scholarly deliverable Peter will likely want to use/publish.

Identified OCR gaps and prepared document assembly
Identified OCR gaps and prepared document assembly

Now let me build a proper docx from this compiled transcription.

Ran 2 commands, created a file
Ran 2 commands, created a file

Now let’s verify the docx renders correctly.

Ran 3 commands, viewed 2 files
Ran 3 commands, viewed 2 files

Good, the document builds fine. Now let’s handle the PDF compression for the website.

Ran 6 commands
Ran 6 commands

Good — down to 7.1MB for images. Now let’s assemble into a single PDF.

Assemble compressed pages into single PDF
Assemble compressed pages into single PDF

Here’s where things stand — I got the core work done, with a couple of items worth flagging before you download anything further.

Transcription
The PDF was a 50-page scan (no text layer), so I ran OCR across all pages, cleaned up obvious scanning noise, and compiled it into a Word document titled Preston_Fire_Brigade_history_transcription.docx. A few important caveats:

  • Photo/illustration pages (2, 7, 12, 27, 43, 46, 47, 48, 49) can’t be OCR’d — I’ve marked these with bracketed descriptions instead (a few captions were legible, e.g. “Horse-drawn Steam Fire Engine, period about 1871”).
  • Pages 45 and 50 are tables (appliance counts, list of officers with dates) — the numeric columns didn’t OCR reliably, so these need checking against the original before you rely on them.
  • The title page had a small-caps OCR quirk I want to flag rather than silently “fix”: it reads “Guild Merchant of 1702 to the Guild Merchant of 1952” — that first date looked suspicious to me (Preston Guilds run every 20 years, so 1702 seems like an odd span to pair with 1952), but I didn’t have a reliable way to confirm the true reading, so I left it as OCR read it rather than guess. Worth a manual check.
  • General OCR caveat: proper nouns, dates, and old-style typography are the likeliest error spots — this is a solid first-pass draft, not a verified transcription to the standard you’d normally hold for primary-source work.

Compression
I rebuilt the PDF from the scanned images at a reduced resolution (~150 DPI) and moderate JPEG quality, grayscale — this took the file from 92MB down to about 7.3MB, which should download quickly for your website visitors while remaining legible.

I ran out of tool calls before I could present the final files to you. If you’d like, just say “go ahead” and I’ll finish packaging both the transcription docx and the compressed PDF for download, and I can also do a closer visual pass on pages 45, 50, and the title page if you want me to try to nail down those uncertain details rather than leaving them flagged.

Please go ahead but I do not need the closer pass, I’ll check those details myself

Ran a command, read a file
Ran a command, read a file

Both files are ready: the transcription as a Word doc, and the compressed PDF (down from 92MB to about 7.3MB) for your website visitors to download.

Preston fire brigade history transcription
Document · DOCX
Preston fire brigade history compressed
Document · PDF

Thank you, that completes the task

You’re welcome — glad that’s sorted. Good luck with the Preston fire brigade page, and hope the 1842 plug-drawing riots transcription is going well too.

Are you able to add an OCR layer to the pdf file?

Ran 2 commands
Ran 2 commands

Good, it installed. Let me apply it to the compressed PDF so it gains a searchable text layer while keeping the smaller file size.

Run OCR to add a text layer to the compressed PDF
Run OCR to add a text layer to the compressed PDF

Now let’s verify it worked and check the file size / searchability.

Verify searchable PDF has text layer and check size
Verify searchable PDF has text layer and check size

It worked — the PDF now has an invisible, searchable text layer over the scanned images, and the file size barely changed (still ~7.3MB).

Ran a command, read a file
Ran a command, read a file

Done — this version looks identical to the compressed PDF (same scanned page images, same ~7.3MB size), but now has an invisible text layer underneath, so users can search and copy text from it, and it’ll be indexed properly if it ends up on your website. Same OCR-accuracy caveats apply as before (photo pages have no real text to find, and old typography may throw up occasional garbled matches), but for search/copy purposes it should serve well.

Preston fire brigade history searchable
Document · PDF

Thank you, that works perfectly

You’re very welcome — glad it came together well. Good luck with the fire brigade page.

Claude is AI and can make mistakes. Please double-check responses.


Discover more from preston history

Subscribe to get the latest posts sent to your email.

Leave a Reply