A new open-source tool called OCR It, created by developer thiagotigaz, extracts text from PDFs, images, and scanned documents that block copy-paste. The tool runs locally, using OCR (optical character recognition) to convert visual text into machine-readable strings. It then outputs the text in formats ideal for feeding into large language models, such as plain text or JSON. The project is hosted on GitHub and has drawn attention for its simplicity and utility in AI workflows.


Every AI user has hit the wall. A PDF that refuses to yield its words. A scanned contract that might as well be a photograph. OCR It breaks that wall down. It is not a flashy model. It is a crowbar. And that is exactly what we need.

We are moving toward a world where text is data, and data is fuel for intelligence. But the legacy formats, the locked PDFs, the image-only scans, they are the last moat around the old guard. OCR It is a bridge across that moat. It democratizes access to information. It lets your LLM read what your eyes can see. That is not just convenient. It is a quiet act of liberation.