WhatFontIs-Bench: A Synthetic Benchmark for Font Family Identification in Real-Looking Images

Identifying the typeface used in a photograph is a common practical task, served by several commercial and open tools, yet there is no public test set on which such tools can be compared: photographs of text rarely come with a reliable font label, and near-identical typefaces published under different names make manual labelling error-prone. We introduce WhatFontIs-Bench, a synthetic benchmark in which the ground truth is known by construction. Version 1.0 contains 11,995 JPEG images of a single word set in one of 600 fonts (200 sans-serif, 200 serif, 100 slab serif, 100 monospaced), rendered onto CC0 photographs of real surfaces, inside real scenes and on printed objects, at three controlled difficulty levels. Every image records the exact font, the text, the word outline, the outline of every letter and all rendering parameters, in JSONL and COCO format. We define a family-level top-k evaluation protocol and report a first baseline: a production font-identification API that searches a catalogue of over 1.2 million fonts reaches 83.7% top-1, 93.3% top-5 and 96.5% top-20 accuracy, with sans-serif typefaces markedly harder (75.7% top-1) than slab serifs (95.0%). The dataset is released under CC BY 4.0.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22876579
Primary Topic
Handwritten Text Recognition Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

WhatFontIs-Bench: A Synthetic Benchmark for Font Family Identification in Real-Looking Images

Alexandru Cuibari
Zenodo (CERN European Organization for Nuclear Research)
Handwritten Text Recognition Techniques
preprint

WhatFontIs-Bench: A Synthetic Benchmark for Font Family Identification in Real-Looking Images

Alexandru Cuibari
preprint en

Abstract

Identifying the typeface used in a photograph is a common practical task, served by several commercial and open tools, yet there is no public test set on which such tools can be compared: photographs of text rarely come with a reliable font label, and near-identical typefaces published under different names make manual labelling error-prone. We introduce WhatFontIs-Bench, a synthetic benchmark in which the ground truth is known by construction. Version 1.0 contains 11,995 JPEG images of a single word set in one of 600 fonts (200 sans-serif, 200 serif, 100 slab serif, 100 monospaced), rendered onto CC0 photographs of real surfaces, inside real scenes and on printed objects, at three controlled difficulty levels. Every image records the exact font, the text, the word outline, the outline of every letter and all rendering parameters, in JSONL and COCO format. We define a family-level top-k evaluation protocol and report a first baseline: a production font-identification API that searches a catalogue of over 1.2 million fonts reaches 83.7% top-1, 93.3% top-5 and 96.5% top-20 accuracy, with sans-serif typefaces markedly harder (75.7% top-1) than slab serifs (95.0%). The dataset is released under CC BY 4.0.

Zenodo (CERN European Organization for Nuclear Research)
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

WhatFontIs-Bench: A Synthetic Benchmark for Font Family Identification in Real-Looking Images — Alexandru Cuibari · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS