Skip to content

OCR engine benchmark — Greek + French corpus

A comparison of every OCR engine Aglaïa can drive, across input DPI, on a hard bilingual corpus. Goal: find the lowest DPI that still reads accurately, per engine.

Corpus

athanase-ocr-test.agl — 24 pages (a 12-scan two-page-spread extract, morales scans 62–73). Content stress-tests OCR: French prose + ancient polytonic Greek, titles, footnotes, block quotes, and reference apparatus. Each page is a processed chosen-layout image, native 300 dpi (~1700×2408).

Method

Per engine × DPI ∈ {100, 150, 200, 300}:

Terminal window
AGLAIA_OCR_DPI=<dpi> AGLAIA_OCR_COMPLEMENT=<none|surya|glm> \
uv run python -m aglaia ocr <copy>.agl --ocr <engine> \
[--ocr-lang fr-FR+el-GR] --export pdf:g4+md
  • Timing is split into model-load (one-off, VLM server spin-up) and page processing (steady-state), reported separately by the CLI.
  • Mistral is DPI-independent — it uploads the native G4 PDF in one whole-document request — so it runs once and serves as the accuracy grounding (mistral@300).
  • Accuracy is an order-insensitive word-overlap (Dice on word multisets) of the plain-text-normalised Markdown:
    1. vs each engine’s own native-300 output → DPI degradation;
    2. vs mistral@300 → coarse cross-engine accuracy (Mistral is a reference, not ground truth — it misses some punctuation / diacritics).

Timing — page processing (s/page)

engine100150200300notes
apple_vision0.480.600.620.70fastest; ~0 load
apple_docs0.880.961.011.00Vision + structured doc
mistral———6.64native only; 1 request, 159 s total, ~0 load, ~$0.02
apple_docs + glm12.35.35.66.4complement re-OCRs Greek blocks
apple_docs + surya11.015.016.412.6
glm15.314.821.046.1input-size sensitive
surya43.039.040.746.2slowest; output-token-bound (DPI-flat)
unlimited———10.2DPI-independent (whole-doc, raw blobs — like Mistral); per-page (window=1); ~2.6 s load

VLM model load ≈ 2–3 s (surya/glm/unlimited); 0 for Apple, Mistral.

unlimited is a whole-document engine (recognize_rows), so — like Mistral — it OCRs the raw stored page blobs and is DPI-independent: AGLAIA_OCR_DPI never resamples its input, and 100/150/200/300 produce byte-identical output. It runs per-page by default (window=1). Its fused multipage R-SWA path (window>1, AGLAIA_UNLIMITED_WINDOW) is numerically unstable in this mlx-vlm build — fusing 2+ pages triggers erratic repetition loops (a page balloons to ~30k words; higher repetition_penalty makes it worse), so it’s opt-in until fixed. At window=1: 10.2 s/page, 0.899 vs Mistral — fast and accurate, in the surya/glm band but ~4× faster than surya.

Accuracy — word-overlap vs mistral@300

engine100150200300
apple_vision0.8760.8810.8710.873
apple_docs0.8780.8830.8720.875
surya0.8380.9100.9250.928
glm0.8120.8910.9020.909
apple_docs + surya0.8870.8750.8730.871
apple_docs + glm0.7520.8630.8580.868
unlimited (native, window=1)0.8990.8990.8990.899

unlimited is DPI-independent (one native value repeated); 0.899 puts it in the top band with surya/glm — clean per-page transcription, no Cyrillic hallucination.

Even Apple agrees only ~0.87 with Mistral (same content, different formatting/footnote handling) — so ~0.87 is the “as good as Apple” band, ~0.90+ is closer to Mistral.

DPI degradation — word-overlap vs the engine’s own native-300

engine100150200
apple_vision0.9210.9300.933
apple_docs0.9210.9270.930
surya0.8640.9480.967
glm0.8210.9220.941
apple_docs + glm0.7780.9050.913

Recommendation — a single global default: 200 dpi

For now Aglaïa uses one global OCR DPI of 200 — the value already used in the GUI, and now the CLI default too (override with --ocr-dpi). Why a single default rather than per-engine tuning:

  • 200 is the accuracy ceiling — 300 buys essentially nothing (surya 0.925→0.928, glm 0.902→0.909) at ~2× the cost, so there’s no reason to go higher.
  • 200 is safe for the local VLMs, which clearly degrade below 150 and are at/near their peak by 200.
  • Apple loses nothing meaningful at 200 — 100→200 is only +0.14 s/page (0.48→0.62), so a lower Apple-only default isn’t worth the CLI/GUI split.

The per-engine “knee” (below, from the data) is where you could go lower to save time on a specific engine, but the small Apple delta and VLM floor make a uniform 200 the pragmatic choice:

engine classknee (lowest accurate)note
Apple (vision, docs)100accuracy flat 100→300 (but only 0.14 s/page saved)
Local VLMs (surya, glm)150100 clearly degrades
apple_docs + complement150100 collapses (0.78)

Cross-engine finding: Mistral is both the accuracy reference and faster than every local VLM (6.6 s/page vs surya’s 40+, glm’s 15–46) — only Apple beats it on speed. Trade-off: cloud/paid vs the local engines’ free/offline.

Engine notes

  • apple_vision / apple_docs — fast, DPI-robust, solid on both scripts. Best default when macOS Vision is available.
  • surya — tracks Mistral best (0.928) but slowest; output-token-bound, so DPI barely changes its time.
  • glm — good from 150 up; time grows sharply with DPI.
  • mistral — best accuracy/speed of the non-Apple engines; one billed request per document.
  • unlimited — now working via our in-process MLX port (../unlimited-ocr-mlx). Whole-doc, DPI-independent, per-page (window=1): 10.2 s/page, 0.899 vs Mistral — top-band accuracy at ~4× surya’s speed, no Cyrillic hallucination. The fused multipage R-SWA path is unstable (repetition loops) and opt-in. A strong local option once the q4 model is published.

Caveats

  • Mistral is a reference, not ground truth — scores are relative fidelity; final judgement on Greek/French still wants a human spot-read.
  • One book’s typography and one page style; other corpora may shift the knees.
  • Repetition-loop degeneration on low-context crops (Surya/GLM complement) is suppressed by repetition_penalty=1.15 on the served VLMs.