一种用于图形手写组件和统计书写者分析的聚类方法。

A clustering method for graphical handwriting components and statistical writership analysis.

作者信息

Crawford Amy M, Berry Nicholas S, Carriquiry Alicia L

机构信息

Department of Statistics Iowa State University Ames Iowa USA.

Berry Consultants Austin Texas USA.

出版信息

Stat Anal Data Min. 2021 Feb;14(1):41-60. doi: 10.1002/sam.11488. Epub 2020 Nov 24.

DOI:10.1002/sam.11488

PMID:33664929

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC7894190/

Abstract

Handwritten documents can be characterized by their content or by the shape of the written characters. We focus on the problem of comparing a person's handwriting to a document of unknown provenance using the shape of the writing, as is done in forensic applications. To do so, we first propose a method for processing scanned handwritten documents to decompose the writing into small graphical structures, often corresponding to letters. We then introduce a measure of distance between two such structures that is inspired by the graph edit distance, and a measure of center for a collection of the graphs. These measurements are the basis for an outlier tolerant -means algorithm to cluster the graphs based on structural attributes, thus creating a template for sorting new documents. Finally, we present a Bayesian hierarchical model to capture the propensity of a writer for producing graphs that are assigned to certain clusters. We illustrate the methods using documents from the Computer Vision Lab dataset. We show results of the identification task under the cluster assignments and compare to the same modeling, but with a less flexible grouping method that is not tolerant of incidental strokes or outliers.

摘要

手写文档可以通过其内容或书写字符的形状来表征。我们关注的是在法医应用中所做的那样，利用书写形状将一个人的笔迹与来源不明的文档进行比较的问题。为此，我们首先提出一种处理扫描手写文档的方法，将书写分解为小的图形结构，这些结构通常对应于字母。然后，我们引入一种受图编辑距离启发的两个此类结构之间的距离度量，以及一组图形的中心度量。这些度量是一种抗离群值均值算法的基础，该算法基于结构属性对图形进行聚类，从而创建一个用于对新文档进行分类的模板。最后，我们提出一个贝叶斯层次模型，以捕捉作者生成分配到特定聚类的图形的倾向。我们使用计算机视觉实验室数据集中的文档来说明这些方法。我们展示了在聚类分配下识别任务的结果，并与相同建模但分组方法不太灵活且不能容忍偶然笔画或离群值的情况进行比较。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/7748/7894190/7b66643724b0/SAM-14-41-g001.jpg

相似文献

A clustering method for graphical handwriting components and statistical writership analysis.一种用于图形手写组件和统计书写者分析的聚类方法。

Stat Anal Data Min. 2021 Feb;14(1):41-60. doi: 10.1002/sam.11488. Epub 2020 Nov 24.

A statistical approach to aid examiners in the forensic analysis of handwriting.一种协助笔迹司法鉴定人员进行法医分析的统计方法。

J Forensic Sci. 2023 Sep;68(5):1768-1779. doi: 10.1111/1556-4029.15337. Epub 2023 Jul 14.

A Set of Handwriting Features for Use in Automated Writer Identification.一组用于自动书写者识别的笔迹特征。

J Forensic Sci. 2017 May;62(3):722-734. doi: 10.1111/1556-4029.13345. Epub 2017 Jan 5.

iVision HHID: Handwritten hyperspectral images dataset for benchmarking hyperspectral imaging-based document forensic analysis.iVision HHID：用于基于高光谱成像的文件司法鉴定分析基准测试的手写高光谱图像数据集。

Data Brief. 2022 Feb 16;41:107964. doi: 10.1016/j.dib.2022.107964. eCollection 2022 Apr.

Offline text-independent writer identification using a codebook with structural features.基于结构特征码本的离线文本无关手写体作者识别

PLoS One. 2023 Apr 25;18(4):e0284680. doi: 10.1371/journal.pone.0284680. eCollection 2023.

Size influence on shape of handwritten characters loops.大小对手写字符环形状的影响。

Forensic Sci Int. 2007 Oct 2;172(1):10-6. doi: 10.1016/j.forsciint.2006.11.005. Epub 2007 Jan 4.

Text-independent writer identification and verification using textural and allographic features.使用纹理特征和书写特征进行文本无关的作者识别与验证。

IEEE Trans Pattern Anal Mach Intell. 2007 Apr;29(4):701-17. doi: 10.1109/TPAMI.2007.1009.

Simulation Detection in Handwritten Documents by Forensic Document Examiners.法医文件检验人员对手写文件的模拟检测

J Forensic Sci. 2015 Jul;60(4):936-41. doi: 10.1111/1556-4029.12801. Epub 2015 Jul 19.

Writer verification of partially damaged handwritten Arabic documents based on individual character shapes.基于单个字符形状对部分受损的手写阿拉伯文文件进行书写者验证。

PeerJ Comput Sci. 2022 Apr 20;8:e955. doi: 10.7717/peerj-cs.955. eCollection 2022.

Writer identification using hand-printed and non-hand-printed questioned documents.使用手写和非手写可疑文件进行书写者识别。

J Forensic Sci. 2003 Nov;48(6):1391-5.

引用本文的文献

Interpol questioned documents review 2019-2022.国际刑警组织对2019年至2022年文件的审查

Forensic Sci Int Synerg. 2023 Feb 24;6:100300. doi: 10.1016/j.fsisyn.2022.100300. eCollection 2023.

本文引用的文献

A Set of Handwriting Features for Use in Automated Writer Identification.一组用于自动书写者识别的笔迹特征。

J Forensic Sci. 2017 May;62(3):722-734. doi: 10.1111/1556-4029.13345. Epub 2017 Jan 5.

Text-independent writer identification and verification using textural and allographic features.使用纹理特征和书写特征进行文本无关的作者识别与验证。

IEEE Trans Pattern Anal Mach Intell. 2007 Apr;29(4):701-17. doi: 10.1109/TPAMI.2007.1009.

Tight clustering: a resampling-based approach for identifying stable and tight patterns in data.紧密聚类：一种基于重采样的方法，用于识别数据中的稳定且紧密的模式。

Biometrics. 2005 Mar;61(1):10-6. doi: 10.1111/j.0006-341X.2005.031032.x.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

一种用于图形手写组件和统计书写者分析的聚类方法。

A clustering method for graphical handwriting components and statistical writership analysis.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献