Handwritten text recognition OCR, local on your host. Linux only so far. Local systray for recognizing text from clipboard or a file. (And a bonus plugin for xournal++)
  • JavaScript 38.2%
  • Shell 23.6%
  • Lua 21.8%
  • Python 14.7%
  • Dockerfile 1.7%
Find a file
Vincent Batts 7755115ad9 Reframe root README around ocr-daemon/ocr-tray, not the Xournal/Joplin path
ocr-daemon and ocr-tray are the general-purpose, working core; the
Xournal++ -> Joplin notes workflow is one use case built on top of them
(xournal-plugin), not the shape of the project. Renames the project title
to match the repo (ocr-and-pals), reorders the components list and
screenshots accordingly, and splits the "Why this shape" notes by which
component they actually belong to.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 09:43:05 -04:00
.forgejo/workflows Add Forgejo Actions workflows to build and publish .deb packages 2026-09-20 08:58:42 -04:00
.github/workflows Add Forgejo Actions workflows to build and publish .deb packages 2026-09-20 08:58:42 -04:00
docs/screenshots Add recognition-result-dialog screenshot to the README 2026-08-20 16:36:58 -04:00
ocr-daemon Force umask 022 in build-deb.sh to fix dpkg-deb permission errors 2026-09-20 09:03:01 -04:00
ocr-tray Force umask 022 in build-deb.sh to fix dpkg-deb permission errors 2026-09-20 09:03:01 -04:00
xournal-plugin Standardize ocr-daemon's socket and cache paths for packaging 2026-08-20 14:52:38 -04:00
LICENSE Package ocr-daemon for easy install: CLI command, systemd service, MIT license 2026-08-20 10:27:32 -04:00
README.md Reframe root README around ocr-daemon/ocr-tray, not the Xournal/Joplin path 2026-09-20 09:43:05 -04:00

ocr-and-pals

Local, offline handwriting OCR: a shared daemon plus the small apps ("pals") that put it to use. Runs TrOCR (via Transformers.js) entirely on-device, with its own line-segmentation step so a whole handwritten page can be recognized at once instead of one line at a time. Started from a narrower question — could libzinnia-style handwriting recognition work as a plugin for Xournal++ and Joplin? — and landed on a more general shape: a standalone local OCR service, usable from anything that can reach it.

Components

  • ocr-daemon — the core piece. A local HTTP (and Unix-socket) daemon that runs TrOCR for handwriting recognition. Everything else in this repo is a client of it. Also runs in Docker with optional Bearer-token-authenticated network listeners, for use beyond just "something else on this machine."
  • ocr-tray — a GTK system tray app for recognizing handwriting from any clipboard image (e.g. a screenshot) or file on disk. The general-purpose way to use ocr-daemon day to day, independent of any particular note-taking app.
  • xournal-plugin — a native Xournal++ Lua plugin, and the original motivating use case: exports the current page, posts it to ocr-daemon, and lets you insert the result as a real text box or save it to a .txt file for something like Joplin to import. See Use case: Xournal++ → Joplin below.
  • Joplin plugin — not built yet; the other half of that same use case.

Screenshots

ocr-tray: menu xournal-plugin: Plugins menu xournal-plugin: recognized text
The ocr-tray system tray menu, showing "Recognize from Clipboard", "Recognize from File...", and "Quit" Xournal++'s Plugins menu, with "Recognize Handwriting (current page)" selected, over a page of handwritten test text The recognition dialog showing recognized text — "this is a hand written note and testing OCR." — with Insert as text / Save to file / Discard buttons

Quick start

cd ocr-daemon
npm install
npm start

That runs ocr-daemon in the foreground for trying it out. For something that sticks around, see ocr-daemon/README.md for running it as a plain ocr-daemon command (npm link) or as a systemd user service (npm run install:systemd).

From there, pick a client:

  • General clipboard/file OCR — install ocr-tray per ocr-tray/README.md and use Recognize from Clipboard or Recognize from File... from the tray menu.
  • Xournal++ notes — install the plugin per xournal-plugin/README.md and try Recognize Handwriting (current page) from the Plugins menu.

ocr-daemon and ocr-tray also have .deb and AppImage builds — see ocr-daemon/README.md and ocr-tray/README.md. If you install ocr-tray's .deb after already using ./install-autostart.sh from a checkout, remove the per-checkout ~/.config/autostart/ocr-tray.desktop first — otherwise both autostart entries fire at login.

Both .debs are also built and published automatically on pushes to main by .forgejo/workflows/ocr-daemon-deb.yml and .forgejo/workflows/ocr-tray-deb.yml — see their "Published builds" sections linked above for the apt source to install from directly.

For running ocr-daemon somewhere other than "on the same machine as its callers" — a server, a container, exposed to a LAN — see ocr-daemon/README.md for the config-file-driven authenticated listeners, its Docker section, and .github/workflows/ocr-daemon-docker.yml for a working build/run/publish pipeline driven by GitHub Actions Variables and Secrets.

Use case: Xournal++ → Joplin

The project's original motivation, and still the most fully-built path through it: handwritten pages in Xournal++, recognized via ocr-daemon, landing as text in Joplin. Today that's xournal-plugin handling the Xournal++ side end to end (recognize → insert as text or save to file); the Joplin-side import piece doesn't exist yet, so getting text into Joplin is currently a manual last step. Neither Xournal++'s Lua plugin sandbox nor Joplin's sandboxed plugin runtime can load a native ONNX runtime directly, which is why this is a daemon both can call over HTTP rather than an in-process library either could use directly.

This is one use case built on ocr-daemon, not the shape of the project — ocr-tray covers the same recognition need for anything else (a screenshot, a scanned page, any app's clipboard) without caring about Xournal++ or Joplin at all.

Why this shape

A few things that weren't obvious going in, each covered in more depth in the relevant README.

ocr-daemon:

  • TrOCR only recognizes one line at a time. Feed it a whole page and it silently collapses everything into one short, usually-wrong guess — ocr-daemon does its own line segmentation so callers can just post a full page.
  • Ruled/dot-grid paper can get misread as text lines. ocr-daemon's segmenter is tuned against a real dot-grid export; xournal-plugin sidesteps the problem entirely by exporting ink-only (background="none").
  • Publishing ocr-daemon's default Docker port does nothing at all. It's bound to 127.0.0.1 inside the container's own network namespace, which -p mapping can't reach — confirmed unreachable even hitting the container's real internal IP directly. Exposing anything from a container requires an explicit, authenticated 0.0.0.0 listener via a config file; see ocr-daemon/README.md.

ocr-tray:

  • GTK3's clipboard API can't see images GNOME Shell's screenshot tool puts on the clipboard, at least on this Wayland setup — confirmed by comparing Gtk.Clipboard.wait_for_image() (nothing) against wl-paste (reads it fine) on the same clipboard state. ocr-tray uses wl-clipboard directly for both reading and writing instead; see ocr-tray/README.md.

xournal-plugin:

  • A flatpak-sandboxed Xournal++ has no network access at all — not even to 127.0.0.1 — but does have full host filesystem access. ocr-daemon listens on a Unix socket for exactly this reason; see ocr-daemon/README.md and xournal-plugin/README.md.
  • app.openDialog text can't be selected or copied, and there's no clipboard-set API. Confirmed against Xournal++'s own source, and clipboard-set-via-helper-process doesn't work around it under flatpak+Wayland either (tested three ways). The plugin instead offers inserting real (selectable) text or writing straight to a file.

Status

ocr-daemon and ocr-tray are the working core: local, offline handwriting recognition from any clipboard image or file, decent-but-not-perfect accuracy (small model, CPU inference), both with .deb and AppImage builds and CI-published packages. xournal-plugin covers the Xournal++ side of the original note-taking use case end to end; text placement in the document is approximate, and there's no Joplin-side piece yet.

License

MIT