Google Stopped Reading The Inside Of Books
The Skills Pack — Ten bucks, all thirty-four of the agent skill folders I run this whole operation on, one file instead of you hunting them down one at a time like a sucker. Voice, fiction, store ops, promo, audio, the works. I use voice-prime every single day and I’d have paid the ten bucks for that folder by itself. Buy it once, no subscription creeping back on you in March.
A Runestone the Faroe Islands Weren’t Supposed to Have — A chapel floor on Sandoy, out in the Faroes, just gave up a runestone, which basically does not happen out there. It came up tangled with the busted pieces of a Virgin-and-child statue in what looks like an early Christian prayer house, the carving still clean enough that somebody flew a specialist in to read it. Nobody’s dating it. “Viking Age into early Medieval” is archaeology-speak for we have no damn idea, and the stone’s fragile enough they had it in the National Museum before anybody could argue about leaving it in the ground.
THE TIP. Last week I showed you how to force the Wayback Machine to save a page on command. Here’s the other half of archive.org almost nobody uses: it’ll also search the actual TEXT inside millions of scanned books, for free, no account, no bullshit paywall pretending to be a database.
Go to archive.org/search, type your term, then look for the “Search inside” filter, or just tack &sin=TXT onto the URL. That hunts the OCR layer of everything scanned in, not just the titles and descriptions people bothered to type up. I’ve found out-of-print field manuals this way, dead newsletters, whole decades of a magazine nobody digitized anywhere else, off a phrase I half-remembered from a paragraph I read once and lost.
Search a PHRASE, not a word. Three or four words in quotes, specific enough it can only live in the one document you’re actually chasing. A single word against millions of OCR’d pages returns garbage forever and you’ll quit before page four.
The OCR is rough. Old type, water damage, and whatever busted scanner somebody was running that day. So when your first phrase whiffs, run it again with the typo a scanner would make. An “rn” that reads back as “m” catches more than you’d think. Anything printed before about 1800, search the f. Books that old used the long s, which scanners read as an f, so “success” comes back “fuccefs.” The misspelling finds documents the real word never will.
I use this constantly, fiction research, the occult stuff, chasing citations nobody else bothered to preserve. Google quit reading the inside of books a long time ago and never sent out a notice. All of it’s still down there. They just stopped telling anybody where.
Somebody search something wild in there and reply and tell me what you dug up, fam. I’ll go look. ~ J.D.
End of brief · the tower is warm