- JavaScript 38.2%
- Shell 23.6%
- Lua 21.8%
- Python 14.7%
- Dockerfile 1.7%
ocr-daemon and ocr-tray are the general-purpose, working core; the Xournal++ -> Joplin notes workflow is one use case built on top of them (xournal-plugin), not the shape of the project. Renames the project title to match the repo (ocr-and-pals), reorders the components list and screenshots accordingly, and splits the "Why this shape" notes by which component they actually belong to. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|---|---|---|
| .forgejo/workflows | ||
| .github/workflows | ||
| docs/screenshots | ||
| ocr-daemon | ||
| ocr-tray | ||
| xournal-plugin | ||
| LICENSE | ||
| README.md | ||
ocr-and-pals
Local, offline handwriting OCR: a shared daemon plus the small apps ("pals") that put it to use. Runs TrOCR (via Transformers.js) entirely on-device, with its own line-segmentation step so a whole handwritten page can be recognized at once instead of one line at a time. Started from a narrower question — could libzinnia-style handwriting recognition work as a plugin for Xournal++ and Joplin? — and landed on a more general shape: a standalone local OCR service, usable from anything that can reach it.
Components
ocr-daemon— the core piece. A local HTTP (and Unix-socket) daemon that runs TrOCR for handwriting recognition. Everything else in this repo is a client of it. Also runs in Docker with optional Bearer-token-authenticated network listeners, for use beyond just "something else on this machine."ocr-tray— a GTK system tray app for recognizing handwriting from any clipboard image (e.g. a screenshot) or file on disk. The general-purpose way to useocr-daemonday to day, independent of any particular note-taking app.xournal-plugin— a native Xournal++ Lua plugin, and the original motivating use case: exports the current page, posts it toocr-daemon, and lets you insert the result as a real text box or save it to a.txtfile for something like Joplin to import. See Use case: Xournal++ → Joplin below.- Joplin plugin — not built yet; the other half of that same use case.
Screenshots
ocr-tray: menu |
xournal-plugin: Plugins menu |
xournal-plugin: recognized text |
|---|---|---|
![]() |
![]() |
![]() |
Quick start
cd ocr-daemon
npm install
npm start
That runs ocr-daemon in the foreground for trying it out. For something that sticks around, see
ocr-daemon/README.md for running it as a plain ocr-daemon command
(npm link) or as a systemd user service (npm run install:systemd).
From there, pick a client:
- General clipboard/file OCR — install
ocr-trayperocr-tray/README.mdand use Recognize from Clipboard or Recognize from File... from the tray menu. - Xournal++ notes — install the plugin per
xournal-plugin/README.mdand try Recognize Handwriting (current page) from the Plugins menu.
ocr-daemon and ocr-tray also have .deb and AppImage builds — see
ocr-daemon/README.md and
ocr-tray/README.md. If you install ocr-tray's .deb
after already using ./install-autostart.sh from a checkout, remove the per-checkout
~/.config/autostart/ocr-tray.desktop first — otherwise both autostart entries fire at login.
Both .debs are also built and published automatically on pushes to main by
.forgejo/workflows/ocr-daemon-deb.yml and
.forgejo/workflows/ocr-tray-deb.yml — see their "Published
builds" sections linked above for the apt source to install from directly.
For running ocr-daemon somewhere other than "on the same machine as its callers" — a server, a
container, exposed to a LAN — see ocr-daemon/README.md
for the config-file-driven authenticated listeners, its Docker section,
and .github/workflows/ocr-daemon-docker.yml for a working
build/run/publish pipeline driven by GitHub Actions Variables and Secrets.
Use case: Xournal++ → Joplin
The project's original motivation, and still the most fully-built path through it: handwritten pages
in Xournal++, recognized via ocr-daemon, landing as text in
Joplin. Today that's xournal-plugin handling the Xournal++ side end to end
(recognize → insert as text or save to file); the Joplin-side import piece doesn't exist yet, so
getting text into Joplin is currently a manual last step. Neither Xournal++'s Lua plugin sandbox nor
Joplin's sandboxed plugin runtime can load a native ONNX runtime directly, which is why this is a
daemon both can call over HTTP rather than an in-process library either could use directly.
This is one use case built on ocr-daemon, not the shape of the project — ocr-tray covers the same
recognition need for anything else (a screenshot, a scanned page, any app's clipboard) without caring
about Xournal++ or Joplin at all.
Why this shape
A few things that weren't obvious going in, each covered in more depth in the relevant README.
ocr-daemon:
- TrOCR only recognizes one line at a time. Feed it a whole page and it silently collapses
everything into one short, usually-wrong guess —
ocr-daemondoes its own line segmentation so callers can just post a full page. - Ruled/dot-grid paper can get misread as text lines.
ocr-daemon's segmenter is tuned against a real dot-grid export;xournal-pluginsidesteps the problem entirely by exporting ink-only (background="none"). - Publishing
ocr-daemon's default Docker port does nothing at all. It's bound to127.0.0.1inside the container's own network namespace, which-pmapping can't reach — confirmed unreachable even hitting the container's real internal IP directly. Exposing anything from a container requires an explicit, authenticated0.0.0.0listener via a config file; seeocr-daemon/README.md.
ocr-tray:
- GTK3's clipboard API can't see images GNOME Shell's screenshot tool puts on the clipboard, at
least on this Wayland setup — confirmed by comparing
Gtk.Clipboard.wait_for_image()(nothing) againstwl-paste(reads it fine) on the same clipboard state.ocr-trayuseswl-clipboarddirectly for both reading and writing instead; seeocr-tray/README.md.
xournal-plugin:
- A flatpak-sandboxed Xournal++ has no network access at all — not even to
127.0.0.1— but does have full host filesystem access.ocr-daemonlistens on a Unix socket for exactly this reason; seeocr-daemon/README.mdandxournal-plugin/README.md. app.openDialogtext can't be selected or copied, and there's no clipboard-set API. Confirmed against Xournal++'s own source, and clipboard-set-via-helper-process doesn't work around it under flatpak+Wayland either (tested three ways). The plugin instead offers inserting real (selectable) text or writing straight to a file.
Status
ocr-daemon and ocr-tray are the working core: local, offline handwriting recognition from any
clipboard image or file, decent-but-not-perfect accuracy (small model, CPU inference), both with .deb
and AppImage builds and CI-published packages. xournal-plugin covers the Xournal++ side of the
original note-taking use case end to end; text placement in the document is approximate, and there's no
Joplin-side piece yet.


