Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
67% Positive
Analyzed from 612 words in the discussion.
Trending Topics
#page#text#point#extension#region#hotkey#auto#scans#ocr#chrome
Discussion Sentiment
Analyzed from 612 words in the discussion.
Trending Topics
Discussion (11 Comments)Read Original on HackerNews
HN isn’t a fan of the generated readmes though, though vibed software (thoroughly used) can be all good.
OK yeah seemed tedious so figured that must’ve not been the only way [probably if you handwrite you could clear that up beforehand]
Jury is still out on which is more trustworthy handling any personal data, Microsoft or Google. Neither.
OCR It is a Chrome extension for that gap. You drag out a capture region once — the text block of the reader, say. After that, one hotkey per page screenshots that exact rectangle, OCRs it, and appends the result to a running transcript. Or start an auto-run and it captures, turns the page, and repeats until the document ends. Then Copy all, or Download .txt, and you have a file to paste into Claude or drop into an agent's context.
Everything runs locally. Tesseract's wasm build and the language data (~10 MB) are committed into the extension, so there are no network requests at all, no API key, and no host permissions at install — single captures ride on activeTab. The irony of an AI-adjacent tool that never talks to a server was not lost on me, but the pages you're capturing are often exactly the ones you don't want to ship to a third party.
Three things turned out more interesting than expected:
- MV3 service workers have no DOM and no Worker, so cropping and OCR live in an offscreen document.
- The next-page control is stored as a point, not a CSS selector. A point survives DOM re-renders and reaches into cross-origin iframes and shadow roots, which nothing the top frame can express does. Routing it was the fiddly part: window.screenX inside an iframe reports the browser window, not the frame, so frames locate themselves by walking same-origin ancestors, and across an origin boundary the parent hands the offset down by postMessage.
- The auto-run waits for each page's OCR before turning. That's what makes end-of-document detection work; a timer-based loop sails past the last page and fills your transcript with copies of it.
Limitations: Chrome's own PDF viewer can't be auto-advanced (it's a plugin no extension can inject into, though capturing from it works fine); the region is a fixed rectangle on screen, so resizing or zooming mid-run breaks it; and accuracy tracks the source — crisp rendered text reads at 93-95% confidence, scans need cleanup before they're worth feeding to anything.
Tests drive a real headless Chrome over CDP, which had its own surprises: Chrome 137+ ignores --load-extension, and headless can't show the optional-permission prompt, so the suite installs a copy with the grant baked in plus a real toolbar click via Extensions.triggerAction to prove the ungranted path still works.
MIT, no build step: https://github.com/thiagotigaz/ocr-it