开源了评估视觉大语言模型(VLLM)对古代汉字视觉感知能力的基准测试Chronicles-OCR。该数据集覆盖了从甲骨文到草书的3000年演变历程,包含7种历史书体与2800张均衡图像。评估涵盖字形定位、细粒度识别、古代文本解析和字体分类四项核心任务,旨在探究视觉分布随时间的变化如何影响模型感知。相关论文与代码已开源。
腾讯混元开源的视觉感知基准,专攻古汉字识别,覆盖从甲骨文到草书的三千年演变,做 OCR 和视觉模型的可以拿来测测自家模型在历史文本上的感知退化。
🎉 🎉 🎉 We're open-sourcing Chronicles-OCR, a visual perception benchmark evaluating VLLMs on ancient Chinese characters.
The dataset spans 3,000 years of evolution. It covers 7 historical scripts from Oracle Bone to Cursive, featuring 2,800 balanced images across highly diverse physical media.
We assess models on 4 core tasks:
- Character Spotting
- Fine-grained Recognition
- Ancient Text Parsing
- Script Classification
The evaluation reveals how visual distribution shifts affect model perception over time.
Explore the dataset and paper below. 👇
📄 Paper: https://arxiv.org/abs/2605.11960
🔗 GitHub: https://github.com/VirtualLUOUCAS/Chronicles-OCR
来源:Tencent Hy · x.com