On a randomised 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library at ATR-1's release, Leo achieved roughly a 5% character error rate — 61% fewer errors than the next-best model (Transkribus Text Titan I ≈ 13%; Claude Opus ≈ 23.3%; Gemini 2.5 Pro ≈ 24.8%; GPT-4.1 ≈ 56.7%). Full study: https://docs.google.com/spreadsheets/d/1HnY1BNUiI2KwAglntTnO_3241DgBoIBgtfA9lDqU8zw/edit?gid=1799472259#gid=1799472259
