• Suppr超能文献
  • 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
定价套餐&价格
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Win 客户端微信小程序
定价
会员套餐积分包API 积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026
  1. 首页
  2. 分享广场
  3. AI与数字人文方法在历史研究中的应用实践、风险与方法论反思

AI与数字人文方法在历史研究中的应用实践、风险与方法论反思

深度研究匿名用户发表于 2026年05月06日 22:2826阅读
发起深度研究
发起深度研究

1. AI与数字人文在历史研究领域的应用发展概述

1.1 技术介入历史研究的演进历程

历史研究的技术介入并非一蹴而就,而是伴随信息技术的发展呈现出阶段性演进特征。早期,技术主要扮演“史料载体数字化”的角色。这一阶段的核心是实现历史文献的电子化存储和初步检索,例如将纸质档案、古籍、报刊等扫描成图像,或通过人工录入转化为文本,极大地扩展了史料的可及性和保存性。数字图书馆、在线数据库的兴起便是这一时期的代表,使得学者能够跨地域、跨时间地查阅资料,提升了研究效率。

进入21世纪,随着互联网的普及和计算能力的提升,数字人文(Digital Humanities, DH)的概念逐渐成型并发展壮大。数字人文强调将计算方法应用于人文学科研究,超越了简单的数字化存储,开始注重对史料的“计算分析”。在这一阶段,历史研究开始运用基础的文本分析工具,例如关键词频率统计、共现分析等,以期从大规模文本中发现新的模式和趋势。这种方法被称为“远读”(distant reading),与传统的“细读”(close reading)形成对比,使得学者能够处理和分析传统方法难以企及的海量数据,从而对历史叙事形成宏观的认识 1。数字工具的普及,例如Google Books、JSTOR和数字化报纸数据库等,使得学者在追踪特定主题、人物、地点或时代信息时,能够使用搜索功能获取定性信息,这极大影响了历史学家研究方式 2。

近年来,人工智能(AI)技术的飞速发展,特别是机器学习(Machine Learning, ML)和自然语言处理(Natural Language Processing, NLP)技术的成熟,标志着技术介入历史研究进入了“AI深度参与分析”的新阶段。与数字人文阶段侧重于“辅助分析”不同,AI技术开始承担更复杂的认知任务,如自动化信息提取、模式识别、关系推断等。例如,光学字符识别(OCR)技术对手写文献和古籍的识别精度不断提高,为后续的文本挖掘奠定了基础。知识图谱的构建则能够将分散在不同史料中的人物、事件、地点等实体进行关联,自动揭示复杂的历史关系网络。此外,AI在处理多模态数据(如历史图像、地图)方面也展现出巨大潜力,通过图像识别、地理信息系统(GIS)与历史地图的结合,实现对历史地理变迁和空间模式的量化分析。

这一演进历程体现了技术在历史研究中从“工具”到“方法论革新者”的角色转变。从最初的档案数字化解决史料可及性问题,到数字人文阶段通过计算方法辅助发现宏观模式,再到AI阶段深度参与史料解析、知识构建乃至潜在的因果推断,技术对历史研究范式的影响愈发深刻 3。然而,伴随技术能力边界的拓展,其可能带来的偏差风险和方法论争议也日益凸显,促使研究者在拥抱新工具的同时,进行更为审慎的反思。

1.2 核心技术应用场景整体框架

当前AI与数字人文技术在历史研究中的应用已形成一个多维度、相互关联的框架,涵盖了从原始史料处理到高级历史分析的各个环节。这些核心技术及其应用场景主要包括以下几个方面:

首先是档案数字化与智能识别。这主要涉及光学字符识别(OCR)技术,特别是针对历史文献的优化应用。历史档案往往包含手写体、繁体字、异体字、模糊印刷或残损文本等复杂情况,传统的OCR技术难以有效处理。通过深度学习和专门训练的模型,研究者正努力提升对这些特殊史料的识别精度,将海量的非结构化图像数据转化为可计算的文本数据。例如,高精度的OCR能将数百万页的历史报纸、日记或官方文书数字化,为后续的文本挖掘和分析奠定基础。

其次是大规模文本挖掘与分析。一旦历史文献被数字化为可读文本,文本挖掘技术便能发挥其巨大潜力。这包括:

  • 主题建模(Topic Modeling):通过算法识别文本集合中反复出现的主题,帮助历史学家从海量文献中发现隐藏的宏观议题、社会关注点或思潮演变,例如分析近代报刊中关于社会改革、民族主义或经济危机的讨论强度和趋势。
  • 情感分析(Sentiment Analysis):评估历史文本中的情绪倾向,揭示特定历史时期或事件中人们的情感态度,例如研究不同群体对某项政策或战争的情绪反应,但需注意其在历史语境下的复杂性与局限性。
  • 命名实体识别(Named Entity Recognition, NER):自动识别并分类文本中的人名、地名、组织机构名、时间等关键实体,为构建结构化数据和知识图谱提供基础。
  • 关联性分析与网络构建:通过分析文本中实体的共现关系,揭示人物之间的社会网络、机构间的互动关系或事件的因果链条,辅助理解历史事件的复杂性。

再者是历史知识图谱的构建与应用。知识图谱是将历史实体(人物、事件、地点、时间等)及其相互关系以结构化的图形形式展现的技术。通过整合来自不同来源的数字化史料,知识图谱能够实现:

  • 实体消歧与对齐:解决历史文献中同名异人、异名同人等问题,确保知识图谱的准确性。
  • 复杂关系建模:构建人物之间的师生、亲属、政党关系,事件与参与者、发生地点、时间的关系,从而形成一个庞大而精细的历史知识网络。
  • 高级查询与推理:历史学家可以利用知识图谱进行复杂的查询,例如“找出所有参与过某次起义并在某个地区担任过官职的人物”,甚至进行一定程度的逻辑推理,辅助发现传统研究中难以察觉的联系。

第四是历史地理信息系统(GIS)与地图可视化。空间信息技术在历史研究中扮演着越来越重要的角色,它主要用于:

  • 历史地图数字化与校正:将古旧地图进行高精度扫描,并利用GIS技术进行地理配准和坐标校正,使其与现代地图数据对齐,解决不同年代地图测绘标准不一的问题。
  • 时空数据叠加与分析:将人口数据、灾害记录、行政区划变迁、军事路线等历史数据叠加到数字化地图上,通过空间分析揭示历史事件的空间分布规律、地理环境对历史进程的影响、人口迁徙模式等。例如,可视化某次瘟疫的传播路径或某一文化现象的地理扩散范围。

最后是多模态数据分析。除了文本和地图,AI技术也开始应用于历史图像、音视频等其他形式的史料。例如,利用计算机视觉技术对历史照片进行人物识别、场景分类,或者分析历史绘画中的符号与主题,为艺术史、文化史研究提供新的视角 4。

这些技术并非孤立存在,而是相互补充,共同构成一个强大的历史研究工具箱。OCR技术为文本挖掘提供基础数据,文本挖掘的实体识别结果则可用于构建知识图谱和GIS中的空间实体。知识图谱和GIS的可视化能力又为历史学家提供了直观的分析界面。通过这种协同作用,AI与数字人文技术正推动历史研究从“史料的数字化”迈向“历史的计算化”,使得学者能够处理前所未有的数据规模,发现宏观模式,并以全新的视角审视历史叙事 567。然而,其方法论的合理性与潜在的偏差风险,也随之成为新的关注焦点。

2. 档案OCR技术在历史文献整理中的落地应用

2.1 多类型历史档案的OCR识别优化路径

光学字符识别(OCR)技术在将纸质历史文献转化为可编辑、可检索的数字文本方面发挥着关键作用,极大地提升了史料的利用效率。然而,历史文献的复杂性和多样性对传统OCR系统提出了严峻挑战。针对手写文献、繁体古籍和残损档案等特殊史料,需要采取特定的优化路径来降低识别误差,提高数字化质量。

1. 手写文献的OCR识别优化:
手写体的多样性、不规范性和个体差异性是其识别的最大难点。为了应对这一挑战,优化路径主要集中在以下几个方面:

  • 深度学习模型与迁移学习: 现代OCR系统普遍采用基于深度学习的架构,如循环神经网络(RNN)结合长短期记忆网络(LSTM)或Transformer模型。这些模型能够捕捉手写字符的复杂模式和上下文依赖关系。针对历史手写文献,研究人员常采用迁移学习策略,即在一个大规模的通用手写数据集上预训练模型,然后利用少量标注过的历史手写文献进行微调。例如,有研究通过结合迁移学习和阿拉伯语OCR技术来数字化古代手稿,显著提高了阿拉伯语手写文本的识别准确率,为历史文献的数字化提供了有力的解决方案 89。
  • 数据增强与合成数据生成: 历史手写文献的标注数据稀缺是一个普遍问题。为了弥补训练数据的不足,可以通过数据增强技术,如对现有图像进行旋转、缩放、形变、添加噪声等操作来扩充数据集。此外,利用生成对抗网络(GANs)等技术生成合成历史手写文档也是一种有效方法,这些合成数据带有精确的真值标注,可以用于训练OCR模型,且无需大量人工标注 1011。
  • 图像预处理: 对手写文档图像进行去噪、二值化、倾斜校正、行分割和字符分割是提高识别率的关键步骤。特别是对于古老的、褪色或墨迹扩散的文献,高质量的预处理能够显著改善后续识别环节的性能。

2. 繁体古籍与多语种、多字体古文献的OCR识别:
繁体古籍的识别面临字体多样(如宋体、楷体、篆书、隶书等)、异体字、合体字、竖排版式、标点缺失以及特定阅读顺序等挑战。

  • 定制化字体模型与字符集: 针对古籍中特有的字体和字符变体,需要训练专门的OCR模型来识别这些字符。例如,有研究针对中文历史文献的复杂布局和阅读模式,提出了定制化的OCR系统,特别强调阅读顺序检测(ROD)的重要性,利用原始图像特征来推断古代文本的固有序列,显著降低了页面错误率 12。
  • 多任务学习框架: 一些先进的框架能够同时处理文本检测、识别和字符修复,尤其适用于中文古籍和碑刻图像,这些图像通常存在严重的视觉退化、字形结构缺失和各种噪声 13。
  • 阅读顺序检测(Reading Order Detection, ROD): 与现代横排文本不同,许多古籍采用竖排或复杂的版式。准确识别文本块的阅读顺序对于正确理解文本内容至关重要。结合视觉线索和机器学习模型来推断文本的阅读顺序是提升古籍OCR效果的关键 12。
  • 处理字符粘连与切割: 中文古籍中常见的字符粘连、重叠现象给字符分割带来困难。有研究提出基于过分割的方法来处理历史文档中的粘连字符,通过分析连接组件的几何信息来找到并分割粘连笔画,并利用动态规划优化分割路径 14。

3. 残损档案的OCR识别优化:
残损档案,如虫蛀、水渍、墨迹渗透、纸张破损等,导致字符模糊、缺失或不完整,这使得OCR识别变得尤为困难。

  • 图像修复与增强技术: 在进行OCR之前,可以利用图像处理技术对残损区域进行修复,例如通过深度学习模型预测并重建缺失的字符笔画。这些技术有助于提高字符的可辨识度。同时,对退化文档图像进行主观和客观质量评估,可以为优化处理流程提供指导 15。
  • 上下文信息与语言模型: 对于难以识别的模糊或缺失字符,可以借助强大的语言模型和历史语料库进行上下文推断和纠错。通过利用前后文信息,甚至特定历史时期的词汇和语法习惯,来猜测最可能的缺失字符或单词,从而提高整体识别的准确性。
  • 多引擎集成与错误校正: 结合多个OCR引擎的识别结果,并通过集成方法进行投票或加权,可以弥补单一引擎的不足。此外,训练专门的错误校正模型,利用条件随机场(CRF)等方法,对OCR输出结果进行后处理校正,能够进一步降低词错误率(WER) 1617。这些校正算法能够学习并纠正OCR常见的错误模式,即使在不同数据集上也能表现出良好的泛化能力 16。

通过上述多方面的优化路径,AI驱动的OCR技术正不断克服历史文献的固有挑战,使其能够更准确、更高效地转化为数字文本,为历史研究的深入开展提供坚实的数据基础。

2.2 OCR成果的校勘与结构化处理流程

尽管AI驱动的OCR技术在历史文献识别方面取得了显著进展,但鉴于历史档案的复杂性和多样性,OCR的输出结果往往并非完美无缺。因此,对OCR成果进行高效的校勘(Proofreading)和将非结构化文本转化为可分析的结构化数据集,是实现史料深度利用的关键步骤。这一流程通常结合人工干预和半自动化工具,以确保数据的准确性和可用性。

1. OCR成果的校勘标准与方法:
校勘的目的是修正OCR识别过程中产生的错误,包括字符错误、排版错误、漏识或误识等。针对历史文献的OCR结果,校勘并非简单的文字核对,还需要考虑历史语境、异体字、古地名、人名等特殊情况。

  • 校勘标准:

    • 准确性(Accuracy):确保文本内容与原始史料完全一致,避免因OCR错误导致的历史信息偏差。这通常通过字符错误率(Character Error Rate, CER)和词错误率(Word错误率, WER)来衡量1819。
    • 完整性(Completeness):检查是否存在遗漏的字符、词语或句子,特别是对于残损或布局复杂的文档。
    • 一致性(Consistency):在同一批次或同一系列史料中,对特定术语、人名、地名的识别和校对应保持一致,以便后续的实体链接和知识图谱构建。
    • 格式符合性(Format Compliance):校勘后的文本应符合预设的结构化格式要求,例如段落、标题、列表的标识。
  • 校勘方法:

    • 人工全量校勘:对于精度要求极高或史料数量有限的核心文献,可采用人工逐字逐句核对的方式。这种方法最准确,但成本高昂且耗时。例如,通过众包平台如Amazon Mechanical Turk,可以以较低成本实现手写文档的转录和校对,确保高精度并提供可搜索的文本20。
    • 半自动化校勘:这是目前最常用的方法,旨在平衡效率与准确性。它依赖于以下技术:
      • 词典匹配与拼写检查:利用历史词典、人名地名库对OCR结果进行初步纠错。例如,可以预加载特定历史时期的词汇表,自动标记不在词典中的词汇,供人工审核。
      • 语言模型辅助校对:通过大规模语料库训练的语言模型,可以识别出OCR结果中不符合语言习惯的错误序列,并给出修正建议。例如,大型语言模型(LLM)被用于纠正OCR错误并检测语言表层形式,有效提高了19世纪拉丁美洲报纸文本的质量21。
      • 交互式校勘工具:利用anyOCR等系统提供的交互式Web应用程序,用户可以对布局和OCR错误进行校正,从而提高数字化的准确性22。
      • 基于规则的后处理:针对OCR后处理中常见的错误模式,可以制定规则进行批量修正,这也是OCR后处理方法的一部分23。

2. 结构化处理流程:
将校勘后的非结构化文本转化为可检索、可分析的结构化数据集,是实现史料深度挖掘的前提。这一过程涉及信息抽取、实体识别和关系提取等关键技术。

  • 命名实体识别(Named Entity Recognition, NER):

    • 目标:自动识别文本中的人名、地名、组织机构、时间、事件等关键实体,并进行分类24。
    • 方法:利用基于规则、统计机器学习或深度学习(如Bi-LSTM-CRF、Transformer)的模型。由于历史文本的语言特点、命名习惯与现代文本存在差异,NER模型需要针对历史语料进行训练和微调,以提高识别精度。
    • 挑战:历史人名、地名可能存在别称、雅号、古今异义等情况,需要进行实体消歧和链接,即将不同表达指向同一个真实世界实体。
  • 关系抽取(Relation Extraction):

    • 目标:识别文本中实体之间的语义关系,例如“人物A是人物B的老师”、“事件C发生在地点D”、“机构E成立于时间F”等。
    • 方法:通常采用监督学习、远程监督或无监督学习方法。历史文献中的关系表达往往比较隐晦,需要复杂的模式匹配或语义分析才能准确提取。
    • 价值:关系抽取是构建历史知识图谱的核心环节,它将分散的实体连接起来,形成一个有意义的知识网络。
  • 事件抽取(Event Extraction):

    • 目标:识别文本中描述的事件及其参与者、时间、地点等要素。
    • 方法:通过识别事件触发词和论元角色来构建事件结构。例如,从一篇奏折中抽取“某年某月某地发生某起叛乱,主要参与者有某某某,起因是某某某”等信息。
    • 价值:事件抽取有助于历史学家从宏观层面把握历史事件的脉络、演变和影响。
  • 版面分析与数据提取:

    • 目标:针对表格、列表、栏目等复杂布局的史料,提取结构化数据。
    • 方法:结合图像处理和机器学习技术进行版面分析,识别不同区域的语义类型。例如,从历史名录中提取人名、官职、所属机构等信息,并将其转化为表格数据25。
    • 案例:针对19世纪哈布斯堡王朝的Hof- und Staatsschematismus(一部记录公务员体系的综合性文献),研究者通过机器学习驱动的版面检测优化OCR流程,显著降低了错误率,提升了从复杂布局历史文档中提取信息的能力18。
  • 元数据补充与规范化:

    • 目标:为结构化数据添加描述性元数据,如史料来源、年代、作者、主题分类等,并遵循统一的元数据标准(如Dublin Core, MODS)。
    • 价值:元数据有助于提高史料的可发现性、可管理性和互操作性。

通过上述校勘和结构化处理流程,原本杂乱无章的图像或文本数据被转化为规范化、可机器读取和分析的知识单元,为历史研究提供了前所未有的数据基础。这些结构化的数据集不仅便于高级检索,更是知识图谱构建、文本挖掘和量化分析等深层应用的核心驱动力。

3. 文本挖掘技术在历史叙事与规律发掘中的应用

3.1 大规模历史文本的内容特征挖掘

文本挖掘技术作为数字人文领域的核心工具之一,为历史研究提供了处理海量非结构化文本数据的新范式。通过计算方法,历史学家能够从近代报刊、地方志、私人文集等大规模史料中,发现传统细读方法难以捕捉的内容特征、主题演变和社会思潮,从而支撑长时段社会变迁和民众生活状态等宏大主题的研究。

1. 主题建模(Topic Modeling)的应用

主题建模是一种统计方法,旨在从大量文档中自动识别抽象的“主题”(topics)。每个主题由一组词语组成,这些词语在语义上高度相关。主题建模在历史研究中尤其适用于揭示大规模文本集合中的潜在结构和演变趋势:

  • 识别历史文献中的核心议题:研究者可以对特定历史时期的大量报纸、期刊或官方公文进行主题建模,从而识别出当时社会关注的核心议题。例如,对19世纪报纸文章进行主题建模,可以发现诸如工业化、城市化、移民、政治改革等主要讨论焦点,并追踪这些议题在不同时间段内的兴衰 26。这种方法帮助历史学家超越单一文献的限制,洞察整个社会的关注点变化。
  • 分析长时段社会思潮的演变:通过对跨越数十年甚至数百年的史料进行时间序列主题建模(如动态主题模型, DTM),可以描绘特定思潮或观念的形成、发展和衰落轨迹。例如,研究芬兰1854年至1917年间报纸和期刊的主题,可以揭示在民族主义兴起、社会变革时期,公众舆论中不同话语的动态演变和相互影响 26。这为理解思想史、文化史提供了量化依据。
  • 比较不同文本来源或群体的关注点:主题建模可以用于比较不同报纸、不同地区的地方志或不同社会阶层的私人文集中所包含的主题差异。例如,分析官方报纸与民间刊物在报道同一事件时的主题倾向,可以揭示官方叙事与民众认知的差异。通过这种比较,历史学家能够更全面地理解多声部的历史叙事。

2. 情感分析(Sentiment Analysis)的应用

情感分析旨在识别文本中所表达的情感极性(积极、消极、中性)以及情感强度。在历史研究中,情感分析有助于揭示特定事件、人物或政策在公众或特定群体中引发的情感反应,从而丰富对历史事件社会影响的理解:

  • 量化历史事件的情感语境:对大规模历史文本进行情感分析,可以评估某一历史事件(如战争、革命、经济危机)发生前后,相关报道或私人记录中的情感倾向。例如,对《纽约时报》1980年至2020年间关于中国“稳定”的报道进行情感分析,发现总体呈现负面情绪,尤其是在1990年至2020年间,与中美关系总体趋势相符 27。这表明媒体对特定议题的情感倾向与国际政治现实存在关联。
  • 追踪情感的演变和扩散:通过时间序列的情感分析,历史学家可以追踪特定情感(如恐惧、希望、不满)在历史进程中的传播和变化。例如,对英国报纸媒体中关于“去殖民化”的报道进行情感分析,发现2010年代中期之后,其含义发生了突然转变,并且涉及文化和机构层面的讨论往往带有更负面的情感 28。这揭示了社会情绪与话语变迁的复杂互动。
  • 分析不同群体的情感表达差异:情感分析可以应用于不同作者、不同地域或不同意识形态背景的文本,以比较他们在特定议题上的情感差异。例如,分析支持与反对某个历史运动的文献,可以量化双方情绪的极端程度和关注焦点,从而深入理解历史冲突的内在动力。

3. 其他内容特征挖掘方法

除了主题建模和情感分析,还有其他多种文本挖掘技术被应用于内容特征的挖掘:

  • 词频与关键词分析:通过统计特定词汇在文本中的出现频率,并与参照语料库进行比较,可以识别出反映特定历史时期或文本特点的关键词。例如,对清末报刊的词频分析,可以揭示“民主”、“科学”、“革命”等新词汇的流行程度及其在社会思想转型中的作用。
  • 共现分析(Collocation Analysis):分析词语之间共同出现的频率和模式,可以揭示词语的语义关联和特定话语模式。例如,分析“贫困”一词常与哪些词语共同出现(如“救济”、“饥荒”、“疾病”),可以重建当时人们对贫困问题的认知框架。
  • 命名实体识别(NER)与密度分析:识别文本中的人名、地名、组织机构名等实体,并对其在文本中的分布密度进行分析。高密度的实体出现可能指示文本的焦点或重要性。结合地理信息系统(GIS),可以分析特定人物或事件在不同地域史料中的提及频率,揭示区域历史的特点。

这些文本挖掘技术共同构成了历史学家分析大规模史料的强大工具箱,使得历史研究能够从“小数据”的精雕细琢走向“大数据”的宏观洞察。通过量化分析和可视化呈现,这些方法不仅辅助历史学家发现新的历史模式和规律,也促进了跨学科研究的融合,使历史研究更加实证化和多样化。

3.2 隐性历史关联的自动识别

传统的历史考据严重依赖于史学家对孤立史料的细致研读和人工关联,这在处理海量和分散的文献时效率低下,且容易遗漏深层或隐性的关联信息。文本挖掘技术通过自动化和计算化的方式,能够从大规模、异构的历史文本中抽取出传统方法难以发现的隐性历史关联,例如人物往来、事件因果、资源流动等,从而有效补充传统史料考据的信息盲区。

1. 人物关系网络的自动构建与分析:

历史文献中记载了大量的人物活动和互动,但这些信息往往分散在不同的传记、日记、书信、官方档案甚至文学作品中。文本挖掘技术可以通过以下方式自动构建人物关系网络:

  • 命名实体识别(NER)与实体消歧:首先,利用NER技术识别出文本中的所有人名。鉴于历史人物可能存在同名异人、异名同人或不同文献中记载方式不一的情况,需要结合上下文信息、时间段、地域等进行实体消歧,确保每个识别出的人名都准确地指向唯一的历史人物。
  • 共现分析与关系抽取:通过分析人名在文本中的共现频率和模式,以及利用关系抽取(Relation Extraction)技术识别人物之间的具体关系(如“师生”、“亲属”、“同事”、“敌对”等),可以构建一个初步的人物关系网络。例如,如果两个人物的名字经常在同一段落或同一事件描述中出现,则可能表明他们之间存在某种关联。更进一步,通过句法分析和语义模式识别,可以抽取出明确的关系谓语,如“A与B结识于某地”、“C师从D”、“E写信给F”等。
  • 社会网络分析(Social Network Analysis, SNA):一旦人物关系网络构建完成,SNA工具可以用于分析网络的结构特征,例如识别出核心人物(中心性分析)、派系形成(社群检测)、信息传播路径等。这有助于揭示历史群体内部的权力结构、影响力分布以及社会组织的演变。例如,对中世纪史诗《草原四十骑士》的文本挖掘,通过可视化人物间的沟通网络,揭示了金帐汗国时期中亚游牧文化的典型社会沟通模式,包括谱系内部的结构化联系和跨谱系的阶级联系,为理解语言与社会之间的关联提供了数据支持 29。

2. 事件因果链与叙事逻辑的识别:

历史事件的发生往往是多重因素相互作用的结果,其因果关系复杂且隐晦。文本挖掘可以辅助识别事件之间的潜在因果关联:

  • 事件抽取(Event Extraction):识别文本中描述的事件及其核心要素,包括事件类型、参与者、时间、地点等。这有助于将非结构化的文本描述转化为结构化的事件记录。
  • 时间关系与因果关系抽取:利用语言模型和模式匹配技术,识别事件之间的时间顺序和因果连接词(如“导致”、“因为”、“所以”、“结果是”等)。例如,从大量新闻报道或历史记载中识别“A事件发生后,紧接着B事件”,并分析是否存在重复的模式或明确的因果声明。深度学习模型如Causal BERT已被开发用于检测文本事件之间的因果关系 30。
  • 否定性事件的识别:在历史叙述中,未能发生或被否定的事件同样重要。例如,“某人未能抵达”或“计划未获批准”等。识别这些否定性的生物事件对于生物医学文献的分析至关重要 31,其方法论也可借鉴于历史文本,有助于更全面地理解事件的背景和影响。
  • 序列模式挖掘:对事件序列进行挖掘,发现反复出现的事件模式或演进路径,从而推断出潜在的因果机制或发展规律。

3. 资源流动与信息传播路径的追踪:

历史上的资源(如货物、财富、知识、技术)流动和信息传播对于理解经济史、文化交流史和政治运作至关重要。文本挖掘可以帮助勾勒这些隐性路径:

  • 命名实体识别与地理信息提取:识别文本中提及的商品、技术名称、货币类型以及相关的地理位置信息。
  • 空间共现与路径分析:通过分析特定资源或信息与不同地点在文本中的共现关系,以及对描述运输、贸易、传播行为的词汇进行抽取,可以绘制出资源和信息在特定时期的流动路径。结合历史GIS系统,可以将这些流动路径可视化,并进行空间分析。
  • 主题演化与传播:通过追踪特定主题或概念在不同地域、不同时期文献中的出现频率和语义变化,可以间接反映知识和思想的传播过程。例如,通过分析不同地区的地方志中关于某种农作物种植技术的描述,可以推断其传播路径和时间。

通过这些文本挖掘技术,历史学家能够超越单个史料的局限,从宏观层面揭示隐藏在海量文本数据中的深层结构和动态关联,从而填补传统研究的空白,为构建更全面、更细致的历史图景提供新的证据和视角。

4. 历史知识图谱的构建与应用价值

4.1 多源异构历史数据的知识图谱搭建方法

历史知识图谱(Historical Knowledge Graph, HKG)是将历史领域中的实体(如人物、事件、时间、地点)及其相互关系以结构化的方式表示出来的一种语义网络。它能够有效整合来自不同来源、格式各异的历史数据,将零散的史料转化为机器可读、可理解、可推理的知识体系。构建历史知识图谱是实现历史数据深度挖掘和智能分析的关键一步,其核心挑战在于处理历史数据的多源异构性、模糊性和不确定性。

历史知识图谱的搭建方法主要包括以下几个关键环节:

1. 实体识别与分类(Entity Recognition and Classification):
这是构建知识图谱的第一步,旨在从原始历史文本或其他数据源中识别出核心的历史实体并进行分类。核心实体通常包括:

  • 人物(Person):历史人物的姓名、别号、字、官职、生卒年月等。
  • 事件(Event):特定的历史事件,如战争、改革、起义、会议、著作发表等。
  • 时间(Time):事件发生或人物生卒的具体时间点或时间段。
  • 地点(Location):历史地名、地理区域、建筑等。
  • 组织机构(Organization):如政府部门、军队、学派、社团等。

实体识别通常采用命名实体识别(NER)技术,利用机器学习(特别是深度学习)模型进行训练。由于历史文献的语言特点和多样性,NER模型需要专门针对历史语料进行训练和优化。例如,注入时间感知知识(temporal-aware knowledge)可以显著提高历史命名实体识别的准确性 32。

2. 实体消歧与对齐(Entity Disambiguation and Alignment):
历史数据中普遍存在同名异人、异名同人、古今地名变迁等问题,这使得实体消歧和对齐成为构建高质量知识图谱的关键。

  • 同名异人消歧:例如,中国历史上可能存在多个同名同姓的人物。消歧通过比对人物的生卒年、籍贯、官职、社会关系、活动事件等上下文信息来区分。
  • 异名同人对齐:同一人物在不同史料中可能使用不同的名字(如本名、字、号、谥号、别称),或因为录入错误导致微小差异。实体对齐的目标是将这些不同的指称链接到知识图谱中唯一的实体ID上。
  • 古今地名映射:历史地名可能随着行政区划变迁、地理环境变化而改变。这需要借助历史地理信息系统(GIS)和专业的历史地名对照表进行精确映射,将历史地名与现代地理坐标或标准地名对齐。
  • 跨来源史料的关联逻辑设定:当从多个数据库或文本中提取实体时,需要确保这些实体在语义上的一致性。例如,如果一份档案将某人称为“王安石”,另一份文献称之为“临川先生”,系统需能够识别二者指代同一实体。这通常通过构建本体论(Ontology)和设定匹配规则来实现,定义实体类型和属性的规范化表示。本体论为知识图谱提供了一个共享的、形式化的概念模型,确保不同来源的数据能够在一个统一的框架下进行整合和理解 3334。

3. 关系抽取与语义建模(Relation Extraction and Semantic Modeling):
关系抽取旨在识别实体之间存在的语义关系,并将这些关系以三元组(Subject, Predicate, Object)的形式存储。

  • 关系类型定义:根据历史研究的需求,需要预先定义一系列关系类型,例如“出生于”、“卒于”、“任职于”、“参与了”、“著有”、“师从”、“配偶是”等。
  • 关系抽取方法:可以采用基于规则的方法(如模式匹配)、监督学习方法(需要大量标注数据)或半监督/无监督学习方法(如远程监督、Bootstrapping)进行关系抽取。例如,DeepKE等工具已被用于从非结构化文本中提取实体间的关系 35。
  • 语义网络构建:将抽取出的实体和关系连接起来,形成一个庞大的语义网络。这一网络不仅包含实体之间的直接关系,还能通过推理机制揭示隐性关系。例如,如果A是B的父亲,B是C的父亲,则可以推断出A是C的祖父。

4. 属性提取与知识补全(Attribute Extraction and Knowledge Completion):
除了实体和关系,知识图谱还包含实体的各种属性,如人物的生卒日期、籍贯、民族,事件的具体时间、地点、参与人数等。

  • 属性提取:通过信息抽取技术从文本中提取实体的具体属性值。
  • 知识补全:利用知识图谱嵌入(Knowledge Graph Embedding)模型,通过学习实体和关系在低维向量空间的表示,预测知识图谱中缺失的实体或关系,从而丰富图谱内容 3637。例如,当发现某人物仅有出生年而无卒年时,模型可能根据其历史同期人物的平均寿命或相关事件信息进行合理推断和补全。

5. 知识融合与图谱存储(Knowledge Fusion and Storage):

  • 多源知识融合:将从不同来源(如传记、地方志、碑刻、档案、数据库等)抽取出的知识进行整合。融合过程中需要解决实体冲突、关系冗余等问题,确保知识的唯一性和准确性。语义网技术和本体工程在知识融合中扮演重要角色,它们提供了标准化的数据模型和互操作性框架,使得来自不同来源的历史数据能够有效集成 383940。
  • 图谱存储:构建完成的知识图谱通常存储在图数据库(如Neo4j, Virtuoso)中,以便进行高效的查询和分析。图数据库能够原生支持图结构数据的存储和遍历操作,非常适合处理复杂的实体关系网络。

通过上述方法,可以搭建一个结构化、可查询、可扩展的历史知识图谱,为历史研究提供一个全新的数据基础设施。这种图谱能够将散落在海量史料中的信息碎片整合起来,形成一张互联互通的知识网络,使得历史学家能够以前所未有的方式探索和理解过去。例如,针对中国古代历史文化领域的知识图谱构建研究,通过爬取百科数据和利用BiLSTM-CNN-CRF模型进行实体识别,结合DeepKE工具进行关系抽取,形成了古籍历史文化知识图谱,旨在帮助公众更快更准确地理解相关知识 35。还有研究构建了辽代历史文化领域的智能问答系统,其核心也是基于知识图谱,通过语义匹配提升问答的准确性 41。此外,一些研究通过将手写文档转化为文本,并结合命名实体识别和历史知识图谱来构建语义搜索模型,提升了历史文档的检索效率 42。针对音乐遗产历史文本中的命名实体识别和链接问题,也通过知识图谱增强了实体链接模型的性能 43。

4.2 历史知识图谱的典型应用场景

历史知识图谱的构建为历史研究带来了革命性的变革,其强大的结构化和关联能力使得传统方法难以处理的复杂历史信息得以清晰呈现和深度分析。以下是历史知识图谱在历史研究中的几个典型应用场景:

1. 历史人物关系网络分析

知识图谱能够将散落在各种史料中的人物信息(如亲属关系、师生关系、同僚关系、社会交往、结社活动等)整合起来,构建出宏大而精细的历史人物关系网络。

  • 识别核心人物与社群结构:通过对知识图谱进行社会网络分析(Social Network Analysis, SNA),可以识别出网络中的关键人物(如拥有高中心性的人物),这些人物往往在历史进程中扮演着重要角色。同时,知识图谱能够揭示历史社会中的社群结构、派系形成与演变。例如,通过分析1809-1917年芬兰大公国超过百万封信件的知识图谱,研究人员可以重建和分析当时重要的社会和通信网络,聚焦于历史人物的传记数据和通信数据,从而从以个人为中心的网络分析转向更宏观的社会中心网络,揭示思想和信息的流动 44。
  • 人物群体行为分析:知识图谱还可以作为历史队列分析的基础。通过构建包含历史人物的知识图谱,研究人员能够系统性地研究特定人群(如某个年代出生、某个职业或某个地域出身的群体)的集体行为模式、社会流动性、职业路径等。例如,CohortVA系统利用从大规模历史数据库构建的知识图谱,生成候选队列并构建队列特征,结合可视化交互界面,帮助历史学家探索基于历史数据的队列行为,从而增强队列识别、人物认证和假设生成的能力 45。
  • 实体消歧与身份认证:在处理大量历史人物数据时,同名异人或异名同人的问题普遍存在。知识图谱通过整合多源信息和利用上下文关系,能够有效解决实体消歧问题,实现历史人物的精确身份认证。荷兰“黄金经纪人”(Golden Agents)项目就致力于创建基础设施,通过跨文化遗产机构的去中心化知识图谱来搜索模式,并开发了通过嵌入和领域知识结合的方法,对17、18世纪阿姆斯特丹档案记录进行实体解析,显著提高了实体解析的性能 4647。

2. 长时段事件演化路径推演

知识图谱能够将分散的历史事件进行时间序列的关联和空间维度的映射,从而辅助历史学家推演长时段历史事件的演化路径和复杂因果关系。

  • 事件链条与因果分析:通过将历史事件(及其参与者、时间、地点)编码为知识图谱中的节点和边,历史学家可以追踪一系列相关事件的发生顺序和相互影响。例如,基于时间知识图谱推理模型(Temporal Knowledge Graph Reasoning Model),可以联合建模相关历史事件和时间邻域事件上下文,从而预测未来事件的发生 48。此外,新的时间知识图谱推理模型如Graph Hawkes Transformer (GHT) 能够在未来时间进行实体预测和时间预测任务,特别适用于长期演化任务,有助于理解历史事件的动态演变 49。
  • 不确定性事件的视觉推理:历史数据中常存在数据缺失或不精确导致的不确定性。利用知识图谱,结合视觉分析系统,可以支持对历史人物时空事件中不确定性的视觉推理。例如,有系统通过构建知识图谱来捕捉由数据缺失和错误引起的不确定性,并利用时间轴视图、地图视图和人际关系矩阵等协调视图来描述和分析事件的异构信息,帮助识别具有缺失或不精确时空信息的不确定事件 50。
  • 主题轨迹与语义演化:知识图谱还可以结合主题建模等技术,追踪特定主题或概念在历史中的演变轨迹。例如,通过构建新闻数据中的知识图谱,可以建模特定主题在时间维度上的语义轨迹,从而支持跨学科研究,如环境变化对商业的影响分析 51。

3. 跨领域史料关联性检索

历史知识图谱通过整合来自不同史料类型(如传记、档案、报刊、地方志、碑刻等)的数据,打破了传统史料分类的壁垒,实现了跨领域、多维度的关联性检索。

  • 统一检索与发现:学者可以通过知识图谱进行复杂的查询,例如“查找在某个特定历史事件中,来自某个地域,并与某个重要人物有书信往来的所有官员”。这种多条件、跨实体的复合检索能力,远超传统关键词搜索的效率和深度。
  • 揭示隐性联系:知识图谱的核心价值在于能够通过图结构揭示传统研究中难以发现的隐性关联。例如,通过发现两个看似无关的历史人物都曾与第三方人物有过互动,或是通过追踪资源(如艺术品、技术)在不同地域和时间点的流转,揭示文化交流或经济往来的深层网络。
  • 辅助构建智能问答系统:基于历史知识图谱,可以开发智能问答系统。用户只需用自然语言提问,系统便能从知识图谱中检索并整合信息,给出精准的答案。例如,有研究基于辽代历史文化的知识图谱构建了智能问答系统,通过语义匹配融合句子和词语级别的交互特征,提高了问答的准确性 41。这不仅便利了专业研究,也为公众提供了更便捷的历史知识获取方式。
  • 支持语义搜索与推荐:知识图谱通过提供丰富的实体和关系上下文,能够显著提升搜索结果的语义相关性。例如,在推荐系统中,知识图谱已被证明能提升新闻推荐的效果,通过融合语义层面和知识层面的新闻表征,能够实现更深层次的个性化推荐 52。此外,图神经网络在推荐系统中的应用日益广泛,通过对用户和物品之间复杂关系的建模,进一步提升了推荐的准确性和多样性 53。

总而言之,历史知识图谱不仅仅是数据的结构化存储,更是一个强大的分析和发现平台。它使得历史学家能够以计算化的方式探索复杂历史现象,从宏观趋势到微观个体互动,都能获得更全面、更深入的理解,从而辅助历史学者发现传统方法难以察觉的联系和模式,推动历史研究范式向数据驱动和计算导向转型。

5. 空间信息技术支撑下的历史地图数字化与分析应用

5.1 古旧地图的数字化校正与时空匹配

历史地图是承载过去空间信息的重要载体,其记载了特定历史时期的地理面貌、行政区划、人口分布、资源利用乃至军事部署等丰富信息。然而,古旧地图在数字化和分析应用中面临诸多挑战,例如测绘标准不一、投影系统未知、精度不足以及地名变迁等。空间信息技术(GIS, Geographic Information System)的介入,为历史地图的数字化校正与时空匹配提供了强大的工具,使得构建可交互、可分析的历史时空数据库成为可能。

1. 古旧地图的数字化与预处理:
首先,需将纸质古旧地图进行高分辨率扫描,生成数字图像。随后,对图像进行预处理,包括去噪、色彩校正、几何校正(如去除扫描过程中可能产生的形变)等,以提高图像质量和后续处理的准确性。

2. 坐标校正与地理配准(Georeferencing):
这是历史地图数字化的核心环节,旨在将古旧地图上的地理要素与现代地理坐标系统建立对应关系,使其能够叠加到现代地理底图上进行分析。

  • 控制点(Ground Control Points, GCPs)的选择与采集: 地理配准的关键在于识别古旧地图上与现代地图或卫星影像上清晰可辨的共同特征点,即控制点。这些点可以是建筑物角点、道路交叉口、河流交汇处、山峰等地理要素 54。选择控制点时需注意其在地图上的均匀分布,且在古旧地图和现代参考数据上都能精确识别。对于早期地图,由于测绘精度和技术限制,控制点的选择尤为重要,往往需要识别历史建筑的角落等不易随时间变化的固定特征 54。
  • 坐标转换模型: 选择合适的数学模型将古旧地图上的像素坐标转换为现代地理坐标。常用的模型包括仿射变换(Affine Transformation)、多项式变换(Polynomial Transformation)和薄板样条函数(Thin Plate Spline, TPS)等。仿射变换适用于形变较小的地图,而多项式变换和TPS可以处理更复杂的非线性形变。然而,对于早期地图,由于其投影系统和大地基准(geodetic datum)参数未知,传统的转换方法可能导致较大误差 54。在这种情况下,迭代重投影或基于遗传算法的优化方法可能会被采用,通过不断调整参数以最小化配准误差 5455。
  • 误差评估与优化: 地理配准完成后,需要评估其位置准确性。通常使用均方根误差(Root Mean Square Error, RMSE)来衡量控制点残差,RMSE越小代表配准精度越高。对于特定历史地图的实验表明,位置精度、不确定性和特征变化检测是衡量其数字化质量的关键指标 56。研究人员会根据误差大小调整控制点或转换模型,以达到最佳配准效果。

3. 古今地名映射(Historical Toponymy Mapping):
历史地名往往随时代变迁,给跨时期空间分析带来挑战。

  • 地名数据库构建: 建立包含古今地名对照关系、别称、隶属关系和时间范围的综合性地名数据库。这需要历史学、地理学和语言学等多学科的合作。例如,唐宋时期长江流域的酒文化数据集就包含了古代和现代地名的映射表,这对于从诗歌中提取时空信息至关重要 57。
  • 命名实体识别(NER)与语义匹配: 利用自然语言处理技术,从历史文本中自动识别地名实体,并结合地名数据库进行匹配和规范化。对于无法直接匹配的地名,可能需要专家知识进行人工判断或利用模糊匹配算法。
  • 本体论与词汇表建设: 为了标准化历史地理概念,建立本体论和共享词汇表至关重要。这有助于解决历史数据中概念的时间变异性问题,确保不同历史数据源之间的互操作性和一致性 58。

4. 图层叠加与时空数据库构建:

  • 矢量化与要素提取: 对地理配准后的历史地图进行矢量化,将地图上的道路、河流、行政边界、居民点等地理要素提取出来,并赋予其属性信息(如名称、类型、年代)。
  • 时空数据库整合: 将不同年代、不同主题的历史地图矢量数据,结合其对应的年代信息,整合到一个统一的时空地理信息数据库中。这个数据库不仅包含空间位置,还包含时间维度,形成“地理本体”和“时间本体”的组合。
  • 多图层叠加与分析: 建立多图层叠加的历史GIS数据库后,可以实现不同历史时期的地理信息在一个平台上进行显示和比较。例如,可以将1905年的城市土地利用图与2003年的遥感影像叠加,分析城市100年间的土地利用变化,包括住宅、商业、工业、道路和绿地等细致类型的变迁 5960。这种叠加分析有助于揭示城市发展、人口迁徙、环境变迁等历史过程的空间规律。

通过上述技术方法,古旧地图不再是静态的图像,而是可以与现代地理信息系统融合、进行动态分析的数字资源。这为历史学家提供了前所未有的工具,以空间视角审视历史,揭示地理环境在历史进程中的作用,并可视化复杂的时空变迁过程。

5.2 历史地理信息的时空分析实践

将古旧地图数字化并进行时空匹配后,历史地理信息系统(GIS)便成为一个强大的分析平台,能够支持历史学家以空间视角探索和理解复杂的历史现象。通过对数字化历史地图进行时空分析,可以揭示人口迁徙路径、区域文化传播规律以及历史遗址分布特征等,为历史研究提供新的维度和实证依据。

1. 人口迁徙路径可视化与分析:

人口迁徙是影响社会、经济和文化变迁的关键历史过程。利用数字化历史地图和GIS技术,可以对人口迁徙的路径、规模和模式进行可视化和量化分析。

  • 数据整合与点位映射: 首先,需要整合包含人口来源地、目的地和时间信息的数据,例如移民登记记录、人口普查数据(如果可用)、家族谱系、灾民流向记录等。然后,将这些历史记录中的地名通过古今地名映射,转换为精确的地理坐标点,并在历史地图上进行标记。
  • 路径绘制与时间序列叠加: 通过连接不同时间点的人口居住地或迁徙中转点,可以绘制出个体或群体迁徙的路径。GIS能够根据时间属性,将不同时期的迁徙路径叠加到对应的历史地图图层上,从而动态展示人口流动的时空演变。例如,通过绘制精神病院病患的居住地到入院地点的空间移动路径,可以分析边缘群体(如贫困和文盲人口)的社会空间过程,揭示其在社会变迁中的被边缘化轨迹 61。
  • 空间模式分析: 结合统计学方法,可以分析迁徙流的空间模式,例如迁徙的集聚或扩散趋势、主要迁徙走廊、迁徙距离分布等。这有助于理解推动人口迁徙的地理、经济或社会因素。例如,研究伦敦铁路网络和人口密度变化的相互作用,发现人口密度与铁路网络密度之间存在正反馈效应,铁路延伸促进了郊区人口增长,而人口增长又反过来推动了更多铁路建设 62。类似的方法可以用于分析历史时期城市化进程中人口的空间流动。

2. 区域文化传播空间规律挖掘:

文化现象的传播和演变往往与地理空间密切相关。GIS为研究文化传播的地理模式提供了有效工具。

  • 文化要素的空间分布映射: 将考古遗址、宗教场所、方言区、艺术风格、习俗分布等文化要素在历史地图上进行精确标记。这些数据可能来源于考古报告、地方志、民族志、宗教文献等。例如,通过对齐家文化遗址的点数据进行空间分析,可以确定其文化核心区(如甘肃省东南部的石兆村和西山坪遗址),并识别出不同区域尺度下的扩散模式,如扩展扩散和迁移扩散 63。
  • 扩散模式分析: 利用空间统计方法(如聚类分析、核密度分析、空间自相关)识别文化要素的空间集聚或分散模式。通过观察文化边界的形成、扩展或收缩,可以推断文化传播的动力学机制,例如是沿着交通干线传播、受地理障碍影响、还是通过社会网络扩散。
  • 影响因素分析: 将文化要素的分布与自然地理因素(如河流、山脉、气候)、社会经济因素(如人口密度、贸易路线、政治中心)叠加分析,可以揭示影响文化传播的关键地理和社会环境变量。例如,研究齐家文化发现,其扩散特征与低海拔地区和沿水系扩散模式有关 63。类似地,福建省文化遗产的分布研究显示,自然因素如平原和丘陵地带、特定坡度以及河流一公里半径范围内,对文化遗产的集聚有显著影响 64。

3. 历史遗址分布特征分析:

历史遗址(如古建筑、古墓葬、历史城镇、纪念碑等)是承载历史记忆和文化价值的重要载体。GIS有助于系统分析其空间分布特征及其形成原因。

  • 遗址点位数据化与属性管理: 将历史遗址的精确地理坐标录入GIS系统,并附带详细属性信息,如遗址类型、建造年代、规模、保存状况、历史背景等。
  • 空间集聚与密度分析: 利用最近邻分析(Nearest Neighbor Analysis, NNA)、核密度估计(Kernel Density Estimation)等空间统计工具,识别历史遗址在特定区域内的集聚区和密度高点。例如,对扬州遗产建筑的研究发现,其空间分布不均,主要集中在邗江区、高邮区和宝应县 65。中国的资源型旅游景点(多为自然风光和历史遗址)的空间分布也呈现出显著的集聚模式,NNI值低至0.57,表明高密度区集中在长三角、北京、西安和洛阳等地 66。
  • 时空演变分析: 通过对不同历史时期遗址分布图层的叠加和比较,揭示遗址分布重心的迁移轨迹。例如,福建省文化遗产的重心从史前和秦汉时期的上游地区,逐渐向明清时期的东部沿海地区迁移 64。扬州遗产建筑的重心也呈现从西南向西北再向东北移动的趋势 65。
  • 环境与人文因素关联: 分析遗址分布与自然环境(如地形、水系、气候)和人文环境(如交通、政治中心、经济活动、文化习俗)之间的关联。例如,研究表明扬州遗产建筑的分布受地形、水资源、盐商文化、运河交通等自然和人文因素的共同影响,其中人文因素的影响更为深刻 65。福建省文化遗产的分布也与平原、丘陵、河流以及区域历史、文化、政治和经济环境密切相关 64。

综上所述,通过数字化历史地图与GIS技术的结合,历史学家能够超越传统二维平面地图的局限,进行多维度、动态化的时空分析。这不仅有助于发现肉眼难以察觉的空间模式和历史规律,也为历史学研究提供了直观的可视化工具,使得历史叙事更具实证性和说服力。GIS的应用将历史地理学从描述性研究推向了分析性研究,并为文化遗产保护、城市规划和旅游发展提供了重要的决策参考 6768。

6. AI与数字人文方法应用于历史研究的偏差风险

6.1 技术层面的偏差来源

尽管AI与数字人文方法为历史研究带来了前所未有的机遇,但其技术层面的固有特性也引入了诸多偏差风险,这些偏差可能对历史研究的结论产生深远影响。理解并识别这些技术偏差的来源,对于确保研究的严谨性和结论的可靠性至关重要。主要的技术层面偏差来源包括训练数据集代表性不足、算法固有偏见以及识别误差传导。

1. 训练数据集代表性不足:
AI模型,特别是深度学习模型,其性能高度依赖于训练数据的质量和代表性。在历史研究领域,训练数据集的代表性不足是普遍存在的挑战,主要体现在以下几个方面:

  • 史料的固有不均衡性:历史史料的留存本身就具有严重的偏向性。例如,官方文献往往比私人文献保存更完整;精英阶层的记录多于普通民众;男性声音多于女性声音;主流文化叙事多于边缘群体。当使用这些具有偏向性的史料来训练AI模型时,模型会内化并放大这种不均衡性。例如,如果OCR模型主要用官方印刷文本训练,那么在识别手写私人信件或地方志时,其准确率会显著下降。若命名实体识别模型主要在关于政治精英的文本上训练,可能导致其在识别普通民众、女性或少数民族的历史人物时表现不佳。
  • 标注数据的稀缺与偏差:训练AI模型需要大量的标注数据(例如,OCR需要手写文本的转录真值,NER需要实体类型标注)。历史领域标注数据的获取成本高昂且专业性强,往往只能依赖少数专家进行小规模标注。这些小规模标注数据可能无法充分覆盖历史文献的多样性和复杂性,导致模型泛化能力不足。此外,标注者自身的知识背景、关注点和主观理解也可能在标注过程中引入偏差,例如,对某些特定历史概念或人物的理解差异,会在标注数据中留下印记。
  • 迁移学习的局限性:在缺乏领域特定数据时,研究者常采用迁移学习(Transfer Learning)的方法,即使用在通用领域训练好的模型在历史数据上进行微调 69。然而,通用领域数据与历史文献之间存在的显著领域差异(Domain Shift,例如语言风格、词汇、句法结构、背景知识等),可能导致模型无法完全适应历史语境,进而产生识别或分析偏差。例如,在现代新闻语料上训练的情感分析模型,可能无法准确捕捉历史文献中特定词汇在当时语境下的情感色彩。

2. 算法固有偏见(Algorithmic Bias):
算法并非中立的工具,其设计原理和优化目标都可能隐含偏见,导致AI系统在决策或预测时产生不公平或不准确的结果 7071。在历史研究中,算法固有偏见可能表现为:

  • 历史偏见的自动化:如上所述,如果训练数据本身就反映了历史上的偏见(例如性别歧视、种族偏见、殖民主义视角),AI模型在学习这些数据后,就会自动复制并可能放大这些历史偏见。例如,一个基于历史数据训练的推荐系统可能会优先展示男性历史人物的传记,而忽略女性历史人物,即使数据中包含女性人物的信息,只是其比例较少 70。这种偏见并非开发者有意为之,而是数据驱动模型固有的风险 71。
  • 不透明的决策过程(黑箱问题):许多复杂的AI模型,特别是深度神经网络,其内部工作机制对于人类研究者而言往往是“黑箱” 7273。这意味着历史学家难以理解模型为何会得出某一特定结论或发现某一特定模式。这种不透明性使得识别和纠正算法可能存在的偏见变得异常困难,也降低了研究结论的可解释性和可信度。历史学强调解释和因果链条的构建,缺乏对AI内部机制的理解,会使得研究结论的论证基础变得脆弱。
  • 公平性与准确性的权衡:在算法设计中,有时需要权衡公平性与准确性 73。例如,为了提高在主流群体数据上的预测准确率,算法可能会牺牲在边缘群体上的表现。在历史研究中,这意味着某些历史叙事或群体可能会在自动化分析中被系统性地忽视或误读,从而进一步固化甚至加剧史料的固有偏见。
  • 度量偏差:算法可能在不同的子群体上表现出不同的准确性,这被称为度量偏差。例如,文献70中提到,算法在应用于不同种族群体时可能会表现出不同的公平性问题,导致结果的偏差。在历史文本分析中,也可能出现对特定历史时期、地域或语言风格的文本分析效果不一的情况。

3. 识别误差的传导与累积:
历史研究中的AI应用往往是多阶段串联的流程,例如,首先进行OCR识别,然后是命名实体识别,再构建知识图谱,最后进行文本挖掘。在这个链条中,前一阶段的错误会不可避免地传导并可能在后续阶段累积,从而对最终的研究结论产生显著影响。

  • OCR错误的影响:OCR技术虽然取得了巨大进步,但在面对历史文献的模糊、残损、手写体等复杂情况时,仍会产生识别错误。这些错误,即使是微小的字符差异,也可能导致命名实体识别失败,或者改变词语的语义,进而影响主题建模、情感分析和知识图谱的准确性。例如,“康熙”被误识别为“康农”,则后续所有关于康熙皇帝的分析都可能出错或被遗漏。
  • NER与关系抽取的误差放大:命名实体识别和关系抽取是构建知识图谱的基础。如果NER错误地识别了实体边界或类型,或者关系抽取错误地判断了实体间的关系,这些错误将被编码进知识图谱。在进行复杂查询或网络分析时,这些初始阶段的错误会被放大,导致错误的关联和推断。例如,将“李白”和“李白(同名商人)”混淆,则其社会关系网络和作品归属都会出现严重错误。
  • 数据质量与分析结果的关联:整体而言,AI系统的数据质量问题(包括不完整、不准确或具有偏见的数据)是导致其决策和行为存在偏差的根本原因 7475。在历史研究中,这种数据质量问题通过各个技术环节层层传递,最终导致自动化分析结果与真实历史事实产生偏差。如果研究者不对中间环节的识别误差进行审慎评估和校勘,那么基于这些误差数据得出的“发现”可能只是技术幻觉而非真实的历史规律。

综上所述,技术层面的偏差来源是AI与数字人文方法在历史研究中需要高度警惕的问题。历史学家在应用这些工具时,必须深入理解其技术原理和局限性,对数据质量和算法输出进行批判性评估,并采取多种策略(如严格的校勘流程、多模型验证、结合传统考据)来降低这些偏差带来的风险,以确保历史研究的科学性和严谨性。

6.2 数据层面的偏差来源

AI与数字人文方法在历史研究中的应用,其效用和可靠性深刻依赖于所处理的原始历史数据。然而,历史数据并非中立、全面且均衡的,其固有的不均衡性是历史研究的常态。当这些带有偏见和不完整的历史数据被算法处理时,不仅其原有的不均衡性会被放大,还可能导致边缘群体的历史叙事在量化分析中被遮蔽,从而对历史的理解产生系统性偏差。

1. 史料留存的固有不均衡性被算法放大:

历史的记录和保存本身就是一个充满偏见和选择的过程。绝大多数留存至今的史料是统治阶级、精英群体、主流文化或特定权力意志的产物。例如:

  • 官方记录的偏重:政府公文、律法典籍、官方编纂的历史等,往往占据史料的主体。这些史料在数量上远超私人信件、地方志或民间传说,在被数字化后,也更容易被OCR系统高精度识别,从而形成庞大的可分析数据集。当AI模型在这些数据上训练时,会更加“擅长”识别和分析官方叙事和精英活动,而相对忽视其他非官方或非主流的视角。
  • 社会精英的声音优势:历史记载更多关注帝王将相、文人墨客、富商巨贾等社会精英。他们的著作、传记、往来信件更容易被整理、刊印和保存。当AI算法对这些数据进行主题建模、情感分析或知识图谱构建时,自然会更多地反映精英阶层的思想、情感和活动网络,使得历史叙事呈现出以精英为中心的倾向。
  • 地域与文化偏向:某些地区(如政治经济文化中心)或强势文化区域的史料留存更为丰富和完整,而边缘地区或少数民族的史料则相对稀少或保存不善。算法在处理这些数据时,会更频繁地“发现”来自数据丰富区域的模式和联系,从而强化对这些区域的关注,进一步边缘化数据匮乏区域的历史。

当AI算法被应用于这些本身就存在偏向性的史料时,其“数据驱动”的特性使得模型会优先从数据量大、模式清晰的史料中学习和提取信息。这种倾向并非算法有意为之,而是其设计目标——即从数据中发现统计规律——的必然结果。结果就是,史料中原有的不均衡性会被算法所固化、放大,甚至在量化分析结果中表现为一种“真实”的统计模式,从而误导历史研究者对历史全貌的认知。例如,若历史文献中某个特定族裔群体或女性的声音甚少,AI在识别、分析和呈现时,也很难生成关于他们的有意义结论,进一步加剧其“隐形”的程度。

2. 边缘群体历史叙事在量化分析中被遮蔽的典型表现:

在量化分析的框架下,那些数据量小、模式不明显、或者与主流叙事不符的边缘群体历史叙事,极易被淹没或扭曲。具体表现为:

  • 数据稀缺导致模型“视而不见”:AI模型在训练过程中,会倾向于学习那些在数据中出现频率高、模式显著的特征。对于边缘群体,如女性、奴隶、劳工、少数民族、LGBTQ+群体或特定地方社区,其相关的史料往往零散、碎片化、非标准化,甚至以口述、民间传说等非书面形式存在。当这些数据不足以形成“足够大”的样本量时,AI模型很难从中发现稳定的模式,从而在分析结果中对这些群体“视而不见”。例如,在人物关系网络中,女性或非精英人物的连接度往往较低,甚至无法被识别,这并非他们没有社会关系,而是相关记载太少,不足以被算法捕捉。
  • 主流概念框架的强制套用:文本挖掘技术(如主题建模、情感分析)通常基于某种预设的语言模型或概念框架。这些框架往往是基于主流语言和观念构建的。当处理边缘群体的文本时,他们独特的词汇、表达方式和认知体系可能无法被这些主流框架准确捕捉和诠释。例如,主流情感分析模型可能无法理解特定历史时期或文化背景下边缘群体所特有的情感表达方式,从而误判其情感倾向。
  • 量化指标的局限性:量化分析往往追求可测量、可比较的指标(如词频、共现次数、网络中心性等)。然而,历史的丰富性和复杂性,特别是边缘群体的经验,往往难以被简单量化。例如,一个群体的文化影响力可能无法仅仅通过文献提及次数来衡量;一个社会运动的深远意义也无法只通过参与人数来体现。过度依赖量化指标,容易忽略那些无法被直接计数但具有重要历史意义的定性因素。
  • “数字红线”(Digital Redlining)效应:正如传统“红线”政策在地理上划分区域以实施歧视一样,数字红线是指算法和数据实践中,通过数据收集、算法设计和模型训练中的偏向,导致对某些群体或区域的系统性排斥或不公平对待 76。在历史研究中,如果某个边缘群体的史料长期缺乏数字化和标注,或其数据被认为“质量不高”,那么与该群体相关的研究就可能在数字化的历史分析中被“划红线”,难以获得算法的有效关注。这种效应可能导致对边缘群体历史的系统性低估或忽视,使得他们的历史叙事在数字时代依然难以浮现 7377。
  • 偏见累积与循环:数据层面的偏见与技术层面的偏见相互作用,形成恶性循环。例如,缺乏边缘群体史料导致算法无法有效识别其信息,这又进一步降低了研究者对边缘群体历史进行量化分析的兴趣,反过来又减少了对相关史料进行数字化和标注的动力,使得边缘群体在数字历史研究中的“不可见”状态持续存在,甚至加剧 717478798081。

因此,在应用AI与数字人文方法进行历史研究时,研究者必须对史料固有的偏见性保持高度警惕,并采取积极措施来弥补数据层面的不均衡。这包括主动寻找和数字化边缘群体的史料、开发针对特定历史语境和少数群体的定制化算法模型、以及结合定性研究方法对量化结果进行批判性审视和补充,以确保历史叙事的全面性和包容性。

7. AI驱动的历史研究的方法论争议与反思

7.1 核心争议焦点梳理

AI与数字人文方法在历史研究中的深入应用,在带来巨大潜力的同时,也引发了一系列深刻的方法论争议和反思。这些争议不仅触及了历史研究的本质,也关乎技术与人文之间关系的未来走向。当前领域内关于AI驱动的历史研究,主要聚焦于以下几个核心争议焦点:量化分析与传统定性考据的优先级、技术介入是否消解历史研究的人文属性、以及算法生成历史结论的可信度边界。

1. 量化分析与传统定性考据的优先级之争

数字人文和AI技术的核心优势在于处理和分析大规模数据,从而进行量化研究,发现宏观模式和统计规律。然而,传统历史研究长期以来高度依赖对史料的细致解读、深度考证和个案分析,即定性研究方法。这两种路径在方法论上的差异,引发了关于它们在历史研究中应如何定位和优先级的讨论。

  • 量化分析的倡导者认为,在面对海量史料时,传统定性方法力有不逮,而量化分析能够帮助历史学家从“树木”中看到“森林”,发现人类认知能力难以企及的宏观趋势、隐藏模式和统计显著性。例如,通过文本挖掘可以分析长时段社会思潮的演变,通过知识图谱可以揭示复杂的人物关系网络,这些都是传统方法难以高效完成的。他们认为,量化分析是历史学走向更科学、更实证、更具可重复性研究的重要途径。
  • 传统定性考据的捍卫者则强调,历史的独特性、复杂性和人文性无法被简单地量化。他们认为,历史事件的发生往往受到具体情境、个体能动性和偶然因素的影响,这些是抽象的量化模型难以捕捉的。过度依赖量化分析可能导致历史的“扁平化”,忽视史料的语义深度、文化语境和叙事细节。他们坚持认为,细致的文献考订、批判性解读和对历史文本多义性的体认,是历史研究不可或缺的核心,任何量化结果都必须回归到原始史料进行定性验证和深度阐释。
  • 折中观点试图弥合两者之间的鸿沟,认为量化分析和定性考据并非相互排斥,而是互补共生的。量化方法可以作为发现问题、生成假设的“探照灯”,帮助历史学家从海量数据中锁定值得深入研究的区域或现象;而定性考据则为这些量化发现提供细致的背景、深度的解释和批判性的审视。例如,先通过主题建模识别出某一新兴的历史主题,再通过细读相关文献来理解其具体内涵和演变细节。这种“远读”与“细读”相结合的策略,旨在实现宏观与微观、广度与深度的统一。

2. 技术介入是否消解历史研究的人文属性

历史学作为一门经典的人文学科,其核心在于对人类经验、意义、价值和叙事的理解与阐释。AI和数字人文技术的介入,特别是其“计算化”和“数据驱动”的特性,引发了关于历史研究人文属性是否会因此被消解的担忧。

  • 担忧者认为,当历史研究过多地依赖算法和模型,历史学家可能从“阐释者”异化为“数据分析师”,对史料的“感知力”、“同情式理解”和“直觉”等人文素养将被技术所取代。他们担心,机器无法理解人类情感、文化意涵和道德判断,算法生成的结果可能会导致历史叙事的去人文化、去道德化。此外,技术的复杂性也可能使得历史学者将更多精力投入到技术操作而非历史思考本身,从而偏离人文学科的本质。
  • 支持者则辩称,技术工具本身是中立的,它只是拓展了历史学家认知和分析的边界,而非取代人文学术的价值。他们认为,AI和数字人文能够解放历史学家从繁重的体力劳动(如史料整理、计数)中解脱出来,将更多精力投入到更高层次的思考、阐释和理论构建中。通过可视化、知识图谱等工具,历史学家反而能更清晰地呈现历史现象,更有效地进行跨学科对话。数字人文的本质,正如一些学者所强调的,是人文与计算的融合,而非简单的技术替代,它为“数字工匠精神”在人文领域的实践提供了平台 82。
  • 批判性反思则进一步指出,问题的核心不在于技术本身,而在于历史学家如何驾驭技术。技术可以作为辅助工具,但最终的解释权和意义赋予者仍然是人。关键在于培养历史学者的“数字素养”和“计算思维”,使其能够批判性地使用工具,并对算法的局限性和偏见保持警惕。同时,也需要重新审视并强调历史学作为人文学科在批判性思维、伦理判断和社会责任方面的独特价值,确保技术应用服务于人文关怀,而非被技术逻辑所主导。

3. 算法生成历史结论的可信度边界

AI模型,特别是大型语言模型(LLM),能够生成看似流畅且具有说服力的文本内容,甚至进行一定程度的推理。这使得人们开始探讨算法生成历史结论的可信度边界问题。

  • 过度信任的风险:一些研究表明,即使是先进的AI模型如ChatGPT-4,在生成科学写作(包括人文领域)时,虽然在重组现有知识方面表现出色,但在生成原创科学内容方面仍存在局限性 83。AI系统虽然能够“流畅”地表达,但在“事实性”上可能存在缺陷,即可能产生“幻觉”(hallucinations),生成看似合理但实际不准确或虚构的信息。如果历史学者对算法输出的结果缺乏批判性审查,盲目采信,可能会导致历史结论的误读或歪曲。
  • “黑箱”问题与可解释性:许多先进的AI模型(如深度学习)内部运作机制复杂,难以被人类完全理解和解释,被称为“黑箱”模型 84。这意味着当AI给出一个历史结论或发现时,历史学家很难追溯其决策过程,理解模型为何会得出这样的结果。这种缺乏可解释性使得算法生成结论的可信度面临挑战,因为它不符合历史研究中对证据链条和逻辑推演的严格要求。
  • 偏见与不透明性:如前所述,AI模型训练数据中固有的偏见以及算法设计本身可能存在的偏见,都会影响算法生成结论的客观性和公平性。如果这些偏见未被识别和纠正,算法可能会强化甚至自动化历史上的不公正叙事。例如,一个基于有偏见史料训练的AI可能生成一个对边缘群体带有歧视色彩的历史解读,而使用者可能难以察觉其中的偏见 778586。
  • 人机协作的边界:关于算法生成结论的可信度,引发了人机协作中“责任”归属的讨论。当AI生成了错误的历史信息,责任应由谁承担?是开发者、使用者,还是模型本身?业界普遍认为,AI应被视为一种辅助工具,最终的决策和解释权仍应归属于人类专家。因此,历史学家在使用AI时,需要保持高度的批判性思维,将AI的输出视为初步的发现或假设,而非最终的结论,并通过传统方法进行严格的验证和交叉检查。

这些争议焦点促使历史学界和数字人文领域的研究者不断反思,如何在充分利用AI和数字人文工具优势的同时,规避其潜在风险,确保历史研究的科学性、严谨性和人文关怀。这不仅需要跨学科的对话和合作,也需要历史学者自身在方法论上进行创新和调适。

7.2 历史研究领域技术应用的适配原则

面对AI与数字人文方法在历史研究中引发的争议与潜在风险,确立一套清晰、审慎的技术应用适配原则至关重要。这些原则旨在最大化技术工具的辅助效用,同时维护历史研究的严谨性、人文底色和批判精神,确保技术服务于历史理解,而非主导历史叙事。核心原则应包括:技术工具的辅助定位、人工研判优先的分析规则,以及历史学者与技术人员跨学科协作的规范路径。

1. 技术工具的辅助定位:

AI和数字人文工具应被明确视为历史研究的辅助性工具和赋能性手段,而非取代历史学家的核心作用。它们的目标是增强历史学家的研究能力,扩展研究的广度和深度,而非替代人类的批判性思维、解释能力和对历史意义的深刻理解。

  • 拓展认知边界而非定义历史:技术工具可以帮助历史学家处理海量数据、发现模式、生成假设、可视化复杂信息,从而拓展传统方法难以触及的认知边界。例如,大规模文本挖掘可以揭示长时段的社会思潮,但最终对这些思潮的内涵、动因和影响的解释权,仍在于历史学家。技术提供的是“什么是”,而非“为什么”或“这意味着什么”。
  • 解放劳动力而非取代思考:AI在史料数字化、信息抽取、实体识别等重复性、高强度任务中具有显著优势,能够将历史学家从繁琐的体力劳动中解放出来。然而,这应促使历史学家将更多精力投入到更深层次的理论建构、问题提出、批判性分析和创新性解释中,而非将思考的责任转嫁给机器。
  • 多元证据的补充而非单一裁决:AI生成的洞察或结论应被视为一种新的证据形式或分析视角,必须与传统的史料考证、文献批判相结合,形成多元证据链条。任何单一的技术分析结果,在未经历史学家严格验证和语境化解读之前,都不能作为最终的历史结论。

2. 人工研判优先的分析规则:

在AI驱动的历史研究流程中,人的判断和批判性审视应贯穿始终,尤其是在关键决策点和结果解释环节。任何技术输出都应被视为初步的、待验证的产物,而不是可以直接采信的“真理”。

  • 前置性批判:对数据源的审慎选择与评估:在将任何史料输入AI系统之前,历史学家必须对其来源、性质、形成背景、可能偏见和局限性进行严格的批判性评估。例如,识别史料的留存偏差、书写目的、作者立场等。对于那些代表性不足或存在严重偏见的史料,应在使用时采取额外的数据平衡策略或进行明确的局限性说明。
  • 过程性监控:对算法输出的持续校勘与验证:在整个数据处理和分析链条中,应设置多重人工校勘和验证环节。例如,OCR结果的抽样校对、命名实体识别结果的核查、知识图谱关系的逻辑检验等。对于任何异常或不符合历史常识的结果,都应追溯到原始史料进行人工查证。当模型生成与历史事实不符的“幻觉”时,历史学家必须具备识别和修正的能力。
  • 终极性解释:对分析结果的语境化解读与意义赋予:AI分析得出的统计模式、关联网络或主题趋势,本身并不具备历史意义。历史学家需要结合其深厚的历史知识、理论框架和人文素养,对这些技术结果进行深度的语境化解读,解释其背后的历史动因、社会影响和文化意涵,并将其融入更宏大的历史叙事之中。这是机器无法替代的核心环节。
  • 可解释性与透明度追求:在选择AI工具时,应优先考虑那些提供一定程度可解释性(Explainable AI, XAI)的模型,即使是“黑箱”模型,也应尽可能探索其内部决策机制,以便历史学家能够理解算法得出某一结论的原因,从而对其进行批判性评估。

3. 历史学者与技术人员跨学科协作的规范路径:

历史研究的复杂性和AI技术的专业性决定了跨学科协作是成功的关键。历史学者和技术人员需要建立高效、互信、规范的合作模式,以弥合学科壁垒,共同推动研究。

  • 建立共同语言与理解框架:历史学者需要学习基本的计算思维、数据处理概念和AI原理,理解技术的优势与局限;技术人员则需要深入了解历史研究的特性、史料的复杂性、历史学家的研究范式和关注点。通过定期交流、联合研讨、共同参与项目设计等方式,逐步建立跨学科的共同语言和理解框架。
  • 明确角色定位与职责分工:在合作项目中,应明确历史学者(作为领域专家)负责提出历史问题、定义研究目标、评估史料、解释结果;技术人员(作为方法专家)负责设计算法、开发工具、处理数据、评估技术性能。职责清晰有助于避免推诿责任或越俎代庖。
  • 从问题出发的迭代式开发:跨学科协作不应是“技术找应用”或“历史被动接受技术”,而应是从历史研究的真实问题出发,共同设计解决方案,并进行迭代式开发。历史学家提出具体的研究需求和史料特性,技术人员据此选择或开发合适的算法,并在历史学家的反馈下不断优化。例如,开发一套OCR系统,技术人员负责算法实现,历史学家负责提供高质量的标注样本并校对输出结果。
  • 共享知识与能力建设:协作的最终目标之一是促进双方的知识共享和能力提升。技术人员应向历史学者普及最新AI技术进展,历史学者则应向技术人员传授历史知识和方法论。同时,鼓励开设跨学科课程,培养既懂历史又懂计算的复合型人才,为未来历史研究的持续发展奠定人才基础。
  • 伦理与风险的共同审视:历史学者和技术人员应共同关注技术应用带来的伦理问题和偏差风险,并在项目设计、数据处理和结果呈现的每一个环节,都进行伦理审查和风险评估。例如,共同讨论如何保护史料中的个人隐私、如何避免算法放大历史偏见、如何确保研究结果的公平性和包容性。

通过遵循这些适配原则,AI与数字人文方法在历史研究中的应用将能更有效地服务于学科发展,促进对人类过去的全面、深入和批判性理解。

内容由 AI 生成,仅供参考,请仔细甄别

参考文献

1The rise of a new paradigm of literary studies: The challenge of digital humanitiesOpenAlex

Ning Wang
Digital humanities has become a heatedly discussed and even debated topic among humanities scholars. It is true that it has made a revolutionary impact on the teaching and academic research of humanities scholars, which has also raised a severe challenge to traditional humanities scholars who are not ready for the impact of this trend. But in any event, it indicates a shift of paradigm of research, which marks the rise of a new academic paradigm and reading method in humanities studies. It at least bridges the gap between science and technology and the humanities and makes humanities scholars more efficiently digitizing their research results. This is especially true in literary studies, or more specifically, comparative and world literature studies. The so-called “distant reading” is a new way of researching world literature. Comparatists could use such a distant reading method to have a general picture of the historical development and evolution of world literature. But at the same time, close reading is also important which could for literary scholars to have a deep-going research and analysis of individual literary works. Thus it is necessary for comparatists to combine the two different reading methods in comparative and world literature studies.

2The Transnational and the Text-Searchable: Digitized Sources and the Shadows They CastOpenAlex

Lara Putnam
The transnational turn has been happening simultaneously with the digital turn and the implications of this entanglement are profound, although as yet largely undiscussed. The great bulk of methodological discussions of "history in a digital age" have so far centered on approaches that harness computational tools to reveal patterns in large sets of textual or mixed sources-emerging techniques of "text--mining" and "distant reading." 2 But more pervasive shifts brought by the internet age are working a much broader impact on what historians do and how. Only a tiny fraction of us are tackling "big data" with quantitative tools. Vastly more of us use the search functions of Google, Google Books, JSTOR, digitized newspaper databases, Ancestry dot com, and the like as we track down qualitative information on particular topics, people, places, or eras. 3 1 I am very grateful to Julie Greene, Diana Paton, Christian De Vito, and Laura Edwards

3Big Data, new epistemologies and paradigm shiftsOpenAlex

Rob Kitchin
This article examines how the availability of Big Data, coupled with new data analytics, challenges established epistemologies across the sciences, social sciences and humanities, and assesses the extent to which they are engendering paradigm shifts across multiple disciplines. In particular, it critically explores new forms of empiricism that declare ‘the end of theory’, the creation of data-driven rather than knowledge-driven science, and the development of digital humanities and computational social sciences that propose radically different ways to make sense of culture, history, economy and society. It is argued that: (1) Big Data and new data analytics are disruptive innovations which are reconfiguring in many instances how research is conducted; and (2) there is an urgent need for wider critical reflection within the academy on the epistemological implications of the unfolding data revolution, a task that has barely begun to be tackled despite the rapid changes in research practices presently taking place. After critically reviewing emerging epistemological positions, it is contended that a potentially fruitful approach would be the development of a situated, reflexive and contextually nuanced epistemology.

4Detecting Treasures in Museums with Artificial IntelligenceOpenAlex

Walpola Layantha Perera, Heike Messemer, Matthias Heinz, et al.
Museums around the world possess hundreds of thousands of priceless objects, which have stories to tell about human history. While students and scholars study them, even the general public is interested in these stories. If there is a way to automate the information delivery system about these objects it will be of immense value, e.g. it will support students to study these objects and speed up research. Adaptive blended learning options are conceivable, which can perfectly merge digital analysis and onsite viewing. Thus, the preparation and post-processing of studied objects is just as conceivable as the adequate acquisition of information for on-site studies. Examples of such solutions would be mobile apps and computer software that can be used for history and archaeology education as well. However, it is important to identify these objects correctly in order to build such solutions. Computer vision technologies in artificial intelligence (AI) can be used for this. Therefore, this paper will show how AI-algorithms can be used for digital humanities in novel ways, such as for detecting museum treasures.

5Digital Humanities: Knowledge and Critique in a Digital AgeOpenAlex

David M. Berry, Anders Fagerjord
As the twenty-first century unfolds, computers continue to change the way we think about culture, society and what it is to be human: areas traditionally explored by the humanities. In a world of Big Data, Google Books, digital archives, real-time streaming systems and smart phones, our use of culture has been changing dramatically. The digital humanities give us powerful tools and methods for thinking about culture and history in the contemporary world, through the use of sophisticated computing techniques and methods. Berry and Fagerjord provide a comprehensive guide, exploring the history, intellectual work, key arguments and ideas of this emerging discipline. They also undertake a substantive critique, suggesting ways in which the humanities can be enriched through computing, but also how cultural critique can transform the digital humanities.

6Exploring big historical data: the historian's macroscopeOpenAlex

The Digital Humanities have arrived at a moment when digital Big Data is becoming more readily available, opening exciting new avenues of inquiry but also new challenges. This pioneering book describes and demonstrates the ways these data can be explored to construct cultural heritage knowledge, for research and in teaching and learning. It helps humanities scholars to grasp Big Data in order to do their work, whether that means understanding the underlying algorithms at work in search engines, or designing and using their own tools to process large amounts of information. Demonstrating what digital tools have to offer and also what 'digital' does to how we understand the past, the authors introduce the many different tools and developing approaches in Big Data for historical and humanistic scholarship, show how to use them, what to be wary of, and discuss the kinds of questions and new perspectives this new macroscopic perspective opens up. Authored 'live' online with ongoing feedback from the wider digital history community, Exploring Big Historical Data breaks new ground and sets the direction for the conversation into the future. It represents the current state-of-the-art thinking in the field and exemplifies the way that digital work can enhance public engagement in the humanities. Exploring Big Historical Data should be the go-to resource for undergraduate and graduate students confronted by a vast corpus of data, and researchers encountering these methods for the first time. It will also offer a helping hand to the interested individual seeking to make sense of genealogical data or digitized newspapers, and even the local historical society who are trying to see the value in digitizing their holdings. Readership: Researchers, graduate and undergraduate students in the field of Digital Humanities, and people looking to digitize historical archives.

7On Digital HistoryOpenAlex

Gerben Zaagsma
Digital humanities seem to be omnipresent these days and the discipline of history is no exception. This introduction is concerned with the changing practice of ‘doing’ history in the digital age, seen within a broader historical context of developments in the digital humanities and ‘digital history’. It argues that there is too much emphasis on tools and data while too little attention is being paid to how doing history in the digital age is changing as a result of the digital turn. This tendency towards technological determinism needs to be balanced by more attention to methodological and epistemological considerations. The article offers a short survey of history and computing since the 1960s with particular attention given to the situation in the Netherlands, considers various definitions of ‘digital history’ and argues for an integrative view of historical practice in the digital age that underscores hybridity as its main characteristic. It then discusses some of the major changes in historical practice before outlining the three major themes that are explored by the various articles in this thematic issue – digitisation and the archive, digital historical analysis, and historical knowledge (re)presentation and audiences. This article is part of the special issue 'Digital History'.

8Revolutionizing Historical Document Digitization: LSTM-Enhanced OCR for Arabic Handwritten ManuscriptsOpenAlex

Safiullah Faizullah, Muhammad Sohaib Ayub, Turki G. Alghamdi, et al.
Optical Character Recognition (OCR) holds immense practical value in the realm of hand-written document analysis, given its widespread use in various human transactions. This scientific process enables the conversion of diverse documents or images into analyzable, editable, and searchable data. In this paper, we present a novel approach that combines transfer learning and Arabic OCR technology to digitize ancient handwritten scripts. Our method aims to preserve and enhance accessibility to extensive collections of historically significant materials, including fragile manuscripts and rare books. Through a comprehensive examination of the challenges encountered in digitizing Arabic handwritten texts, we propose a transfer learning-based framework that leverages pre-trained models to overcome the scarcity of labeled data for training OCR systems. The experimental results demonstrate a remarkable improvement in the recognition accuracy of Arabic handwritten texts, thereby offering a highly promising solution for the digitization of historical documents. Our work enables the digitization of large collections of ancient historical materials, including manuscripts and rare books characterized by delicate physical conditions. The proposed approach signifies a significant step towards preserving our cultural heritage and facilitating advanced research in historical document analysis.

9A Survey of OCR in Arabic Language: Applications, Techniques, and ChallengesOpenAlex

Safiullah Faizullah, Muhammad Sohaib Ayub, Sajid Hussain, et al.
Optical character recognition (OCR) is the process of extracting handwritten or printed text from a scanned or printed image and converting it to a machine-readable form for further data processing, such as searching or editing. Automatic text extraction using OCR helps to digitize documents for improved productivity and accessibility and for preservation of historical documents. This paper provides a survey of the current state-of-the-art applications, techniques, and challenges in Arabic OCR. We present the existing methods for each step of the complete OCR process to identify the best-performing approach for improved results. This paper follows the keyword-search method for reviewing the articles related to Arabic OCR, including the backward and forward citations of the article. In addition to state-of-art techniques, this paper identifies research gaps and presents future directions for Arabic OCR.

10Generating Synthetic Handwritten Historical Documents with OCR Constrained GANsOpenAlex

Lars Vögtlin, Manuel Drazyk, Vinaychandran Pondenkandath, et al.

11Generating Synthetic Handwritten Historical Documents With OCR\n Constrained GANsOpenAlex

Lars Vögtlin, Manuel Drazyk, Vinaychandran Pondenkandath, et al.
We present a framework to generate synthetic historical documents with\nprecise ground truth using nothing more than a collection of unlabeled\nhistorical images. Obtaining large labeled datasets is often the limiting\nfactor to effectively use supervised deep learning methods for Document Image\nAnalysis (DIA). Prior approaches towards synthetic data generation either\nrequire expertise or result in poor accuracy in the synthetic documents. To\nachieve high precision transformations without requiring expertise, we tackle\nthe problem in two steps. First, we create template documents with\nuser-specified content and structure. Second, we transfer the style of a\ncollection of unlabeled historical images to these template documents while\npreserving their text and layout. We evaluate the use of our synthetic\nhistorical documents in a pre-training setting and find that we outperform the\nbaselines (randomly initialized and pre-trained). Additionally, with visual\nexamples, we demonstrate a high-quality synthesis that makes it possible to\ngenerate large labeled historical document datasets with precise ground truth.\n

12Reading between the Lines: Image-Based Order Detection in OCR for Chinese Historical DocumentsOpenAlex

Hsing-Yuan Ma, Hen‐Hsen Huang, Chao-Lin Liu
Chinese historical documents, with their unique layouts and reading patterns, pose significant challenges for traditional Optical Character Recognition (OCR) systems. This paper introduces a tailored OCR system designed to address these complexities, particularly emphasizing the crucial aspect of Reading Order Detection(ROD). Our system operates through a threefold process: text detection using the Differential Binarization++ model, text recognition with the SVTR Net, and a novel ROD approach harnessing raw image features. This innovative method for ROD, inspired by human perception, utilizes visual cues present in raw images to deduce the inherent sequence of ancient texts. Preliminary results show promising reductions in page error rates. By preserving both content and context, our system contributes meaningfully to the accurate and contextual digitization of Chinese historical manuscripts.

13UniText: A Unified Framework for Chinese Text Detection, Recognition, and Restoration in Ancient Document and Inscription ImagesOpenAlex

Lu Shen, Zewei Wu, Xiaoyuan Huang, et al.
Processing ancient text images presents significant challenges due to severe visual degradation, missing glyph structures, and various types of noise caused by aging. These issues are particularly prominent in Chinese historical documents and stone inscriptions, where diverse writing styles, multi-angle capturing, uneven lighting, and low contrast further hinder the performance of traditional OCR techniques. In this paper, we propose a unified neural framework, UniText, for the detection, recognition, and glyph restoration of Chinese characters in images of historical documents and inscriptions. UniText operates at the character level and processes full-page inputs, making it robust to multi-scale, multi-oriented, and noise-corrupted text. The model adopts a multi-task architecture that integrates spatial localization, semantic recognition, and visual restoration through stroke-aware supervision and multi-scale feature aggregation. Experimental results on our curated dataset of ancient Chinese texts demonstrate that UniText achieves a competitive performance in detection and recognition while producing visually faithful restorations under challenging conditions. This work provides a technically scalable and generalizable framework for image-based document analysis, with potential applications in historical document processing, digital archiving, and broader tasks in text image understanding.

14Touching character segmentation method for Chinese historical documentsOpenAlex

Xiaolu Sun, Liangrui Peng, Xiaoqing Ding
The OCR technology for Chinese historical documents is still an open problem. As these documents are hand-written or hand-carved in various styles, overlapped and touching characters bring great difficulty for character segmentation module. This paper presents an over-segmentation-based method to handle the overlapped and touching Chinese characters in historic documents. The whole segmentation process includes two parts: over-segmented and segmenting path optimization. In the former part, touching strokes will be found and segmented by analyzing the geometric information of the white and black connected components. The segmentation cost of the touching strokes is estimated with connected components' shape and location, as well as the touching stroke width. The latter part uses local optimization dynamic programming to find best segmenting path. HMM is used to express the multiple choices of segmenting paths, and Viterbi algorithm is used to search local optimal solution. Experimental results on practical Chinese documents show the proposed method is effective.

15Subjective and objective quality assessment of degraded document imagesOpenAlex

Atena Shahkolaei, Hossein Ziaei Nafchi, Somaya Al-Máadeed, et al.

16How well does multiple OCR error correction generalize?OpenAlex

William B. Lund, Eric K. Ringger, Daniel D. Walker
As the digitization of historical documents, such as newspapers, becomes more common, the need of the archive patron for accurate digital text from those documents increases. Building on our earlier work, the contributions of this paper are: 1. in demonstrating the applicability of novel methods for correcting optical character recognition (OCR) on disparate data sets, including a new synthetic training set, 2. enhancing the correction algorithm with novel features, and 3. assessing the data requirements of the correction learning method. First, we correct errors using conditional random fields (CRF) trained on synthetic training data sets in order to demonstrate the applicability of the methodology to unrelated test sets. Second, we show the strength of lexical features from the training sets on two unrelated test sets, yielding a relative reduction in word error rate on the test sets of 6.52%. New features capture the recurrence of hypothesis tokens and yield an additional relative reduction in WER of 2.30%. Further, we show that only 2.0% of the full training corpus of over 500,000 feature cases is needed to achieve correction results comparable to those using the entire training corpus, effectively reducing both the complexity of the training process and the learned correction model.

17Ensemble Methods for Historical Machine-Printed Document RecognitionOpenAlex

William B. Lund
The usefulness of digitized documents is directly related to the quality of the extracted text. Optical Character Recognition (OCR) has reached a point where well-formatted and clean machine- printed documents are easily recognizable by current commercial OCR products; however, older or degraded machine-printed documents present problems to OCR engines resulting in word error rates (WER) that severely limit either automated or manual use of the extracted text. Major archives of historical machine-printed documents are being assembled around the globe, requiring an accurate transcription of the text for the automated creation of descriptive metadata, full-text searching, and information extraction. Given document images to be transcribed, ensemble recognition methods with multiple sources of evidence from the original document image and information sources external to the document have been shown in this and related work to improve output. This research introduces new methods of evidence extraction, feature engineering, and evidence combination to correct errors from state-of-the-art OCR engines. This work also investigates the success and failure of ensemble methods in the OCR error correction task, as well as the conditions under which these ensemble recognition methods reduce the Word Error Rate (WER), improving the quality of the OCR transcription, showing that the average document word error rate can be reduced below the WER of a state-of-the-art commercial OCR system by between 7.4% and 28.6% depending on the test corpus and methods. This research on OCR error correction contributes within the larger field of ensemble methods as follows. Four unique corpora for OCR error correction are introduced: The Eisenhower Communiqués, a collection of typewritten documents from 1944 to 1945; The Nineteenth Century Mormon Articles Newspaper Index from 1831 to 1900; and two synthetic corpora based on the Enron (2001) and the Reuters (1997) datasets. The Reverse Dijkstra Heuristic is introduced as a novel admissible heuristic for the A* exact alignment algorithm. The impact of the heuristic is a dramatic reduction in the number of nodes processed during text alignment as compared to the baseline method. From the aligned text, the method developed here creates a lattice of competing hypotheses for word tokens. In contrast to much of the work in this field, the word token lattice is created from a character alignment, preserving split and merged tokens within the hypothesis columns of the lattice. This alignment method more explicitly identifies competing word hypotheses which may otherwise have been split apart by a word alignment. Lastly, this research explores, in order of increasing contribution to word error rate reduction: voting among hypotheses, decision lists based on an in-domain training set, ensemble recognition methods with novel feature sets, multiple binarizations of the same document image, and training on synthetic document images.

18Enhancing OCR in historical documents with complex layouts through machine learningOpenAlex

David Fleischhacker, Roman Kern, Wolfgang Göderle
Abstract This paper explores the challenge of processing and extracting information from large quantities of printed serial sources from the 19th century, which have been largely untapped due to the inadequacies of existing extraction techniques. We focus on the Habsburg Central Europe’s Hof- und Staatsschematismus , a comprehensive record published between 1702 and 1918 that documents the Habsburg civil service’s hierarchy and the evolution of its central administration over two centuries. Our approach sees the significant investment into machine learning-driven layout detection prior to the OCR-process. We generated synthetic data mimicking the Hof- und Staatsschematismus style for initial training of a Faster R-CNN model, followed by fine-tuning the model with a smaller dataset of manually annotated historical documents. Subsequently, we optimised Tesseract-OCR for our document style to enhance the combined structure extraction and OCR process. Our evaluation demonstrates significant improvements in OCR performance metrics (WER and CER), with the combined structure detection and fine-tuned OCR process showing a decrease in error rates of 15.68 percentage points for CER and 19.95 percentage points for WER. These findings underscore the potential of ML techniques in facilitating the extraction and analysis of historical documents.

19Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case studyOpenAlex

Mark J. Hill, Simon Hengchen
Abstract This article aims to quantify the impact optical character recognition (OCR) has on the quantitative analysis of historical documents. Using Eighteenth Century Collections Online as a case study, we first explore and explain the differences between the OCR corpus and its keyed-in counterpart, created by the Text Creation Partnership. We then conduct a series of specific analyses common to the digital humanities: topic modelling, authorship attribution, collocation analysis, and vector space modelling. The article concludes by offering some preliminary thoughts on how these conclusions can be applied to other datasets, by reflecting on the potential for predicting the quality of OCR where no ground-truth exists.

20Using Amazon Mechanical Turk to Transcribe Historical Handwritten DocumentsOpenAlex

Andrew Lang, Joshua Rio-Ross
The developing “information age” is continually unraveling new ways of discovering, presenting and sharing information. Most new academic material is digitally formatted upon its creation and is thus easy to find and query. However, there remains a good deal of material from times prior to the “information age” that has yet to be converted to digital form. Much of this material can be found in library collections—whether academic, public or private—and thus remains available only to a limited number of locals or willing-and-able sojourners. Using OCR technology, most typeset documents can be digitized and made available online; and there are several projects underway to do exactly this. However, there remains little to be done for handwritten materials. Those who own collections of handwritten documents are increasingly wanting to make the content thereof available to the general public. Unfortunately, traditional transcription models typically prove to be expensive or inefficient and pdf snapshots are not searchable. We have developed a model for digital transcription using Google Docs and Amazon's Mechanical Turk. Using this model, one can use an online workforce to efficiently transcribe handwritten texts and perform quality control at a cost much lower than professional transcription services. To illustrate the model we used Amazon’s Mechanical Turk to transcribe and then proofread the Frederick Douglass Diary which we have made available on a public searchable wiki. The total cost of transcription and proofreading for the 72 page diary was less than $25.00 with some pages being transcribed and proofread for as little as $0.04. Our results show that using Amazon’s Mechanical Turk holds great promise for providing an affordable transcription method for hand-written historical documents making them easily sharable and fully searchable.

21Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR CorrectionOpenAlex

Laura Manrique-Gómez, Tony Montes, Arturo Rodriguez Herrera, et al.
This paper presents two significant contributions: First, it introduces a novel dataset of 19thcentury Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analysis in this region.Second, it develops a flexible framework that utilizes a Large Language Model for OCR error correction and linguistic surface form detection in digitized corpora.This semi-automated framework is adaptable to various contexts and datasets and is applied to the newly created dataset.

22anyOCR: An Open-Source OCR System for Historical ArchivesOpenAlex

Syed Saqib Bukhari, Ahmad Kadi, Mohammad Ayman Jouneh, et al.
Currently an intensive amount of research is going on in the field of digitizing historical archives for converting scanned document images into searchable full text. This paper presents the "anyOCR" system which mainly emphasize the techniques requires for digitizing a historical archive with high accuracy. It is an open-source system for the research community who can easily apply the anyOCR system for digitizing historical archives. The anyOCR system supports a complete document processing pipeline, which includes layout analysis, training OCR models and text line prediction, with an addition of intelligent and interactive layout and OCR error corrections web applications. The anyOCR system can also be used for contemporary document images containing diverse, simple to complex, layouts. This paper describes the current state of the anyOCR system, its architecture, as well as its major features. This paper also provides information about the availability, documentation, and tutorials of the anyOCR system.

23Survey of Post-OCR Processing ApproachesOpenAlex

Thi Tuyet Haï Nguyen, Adam Jatowt, Mickaël Coustaty, et al.
Optical character recognition (OCR) is one of the most popular techniques used for converting printed documents into machine-readable ones. While OCR engines can do well with modern text, their performance is unfortunately significantly reduced on historical materials. Additionally, many texts have already been processed by various out-of-date digitisation techniques. As a consequence, digitised texts are noisy and need to be post-corrected. This article clarifies the importance of enhancing quality of OCR results by studying their effects on information retrieval and natural language processing applications. We then define the post-OCR processing problem, illustrate its typical pipeline, and review the state-of-the-art post-OCR processing approaches. Evaluation metrics, accessible datasets, language resources, and useful toolkits are also reported. Furthermore, the work identifies the current trend and outlines some research directions of this field.

24Named Entity Recognition and Classification in Historical Documents: A SurveyOpenAlex

Maud Ehrmann, Ahmed Hamdi, Elvys Linhares Pontes, et al.
After decades of massive digitisation, an unprecedented number of historical documents are available in digital format, along with their machine-readable texts. While this represents a major step forward with respect to preservation and accessibility, it also opens up new opportunities in terms of content mining and the next fundamental challenge is to develop appropriate technologies to efficiently search, retrieve, and explore information from this ‘big data of the past’. Among semantic indexing opportunities, the recognition and classification of named entities are in great demand among humanities scholars. Yet, named entity recognition (NER) systems are heavily challenged with diverse, historical, and noisy inputs. In this survey, we present the array of challenges posed by historical documents to NER, inventory existing resources, describe the main approaches deployed so far, and identify key priorities for future developments.

25Mining Spatio-temporal Data on Industrialization from Historical RegistriesOpenAlex

D. Berenbaum, D. Deighan, T. Marlow, et al.
Despite the growing availability of big data in many fields, historical data on socio-evironmental phenomena are often not available due to a lack of automated and scalable approaches for collecting, digitizing, and assembling them. We have developed a datamining method for extracting tabulated, geocoded data from printed directories. While scanning and optical character recognition (OCR) can digitize printed text, these methods alone do not capture the structure of the underlying data. Our pipeline integrates both page layout analysis and OCR to extract tabular, geocoded data from structured text. We demonstrate the utility of this method by applying it to scanned manufacturing registries from Rhode Island that record 41 years of industrial land use. The resulting spatio-temporal data can be used for socio-environmental analyses of industrialization at a resolution that was not previously possible. In particular, we find strong evidence for the dispersion of manufacturing from the urban core of Providence, the state’s capital, along the Interstate 95 corridor to the north and south.

26Topic Modelling Discourse Dynamics in Historical NewspapersOpenAlex

Jani Marjanen, Elaine Zosa, Simon Hengchen, et al.
This paper addresses methodological issues in diachronic data analysis for historical research. We apply two families of topic models (LDA and DTM) on a relatively large set of historical newspapers, with the aim of capturing and understanding discourse dynamics. Our case study focuses on newspapers and periodicals published in Finland between 1854 and 1917, but our method can easily be transposed to any diachronic data. Our main contributions are a) a combined sampling, training and inference procedure for applying topic models to huge and imbalanced diachronic text collections; b) a discussion on the differences between two topic models for this type of data; c) quantifying topic prominence for a period and thus a generalization of document-wise topic assignment to a discourse level; and d) a discussion of the role of humanistic interpretation with regard to analysing discourse dynamics through topic models.

27Discursive use of stability in New York Times’ coverage of China: a sentiment analysis approachOpenAlex

Guofeng Wang, Yilin Liu, Shengmeng Tu
Abstract The importance of stability has been consistently emphasized in China and the discursive use of stability is found to have legitimizing effects in Chinese newspapers, but how such political keywords are employed by the newspaper of a country that is ideologically distinct from China remains underexplored. This study addresses this gap by investigating the use of stability in The New York Times ’ coverage of China between 1980 and 2020, drawing on critical discourse analysis (particularly, the discourse-historical approach) and sentiment analysis. A diachronic quantitative analysis demonstrates an overall negative sentiment in news reports relating to China’s stability across these years, with positive sentiment evident only during the 1980s and negative sentiment prevailing from 1990 to 2020. These findings are consistent with general trends in US-China relations and US foreign policy over the four decades. Qualitative analysis reveals that negative sentiment focuses on sociopolitical and territorial issues, whereas positive sentiment focuses primarily on economic and financial aspects, indicating that the newspaper views the issue of China’s stability from a politically self-interested perspective of the US and is also concerned about the persistence of certain dominant ideologies in American society. This study contributes to a greater comprehension of the use of political keywords in national and international news discourse, especially by the media of ideologically diverse societies. Moreover, because the application of sentiment analysis to critical discourse analysis and news discourse analysis has proven to be time-efficient, verifiable, and accurate, researchers can confidently employ it to disclose hidden meanings in texts.

28Decolonisation: meaning, sentiments and implications for heritageOpenAlex

Jeroen Cant, Katelijne Nolet, Suzie Thomas
Within the printed and online newspaper media in the UK, notions of 'decolonisation' referring to various contexts, such as historical, relating to museums and institutions, cultural decolonisation and in relation to modern independence discussions, can be traced. In this article, we have applied topic modelling and natural language processing methods to carry out a classification of, and sentiment analysis on, newspaper headlines and texts from leading British newspapers covering decolonisation over the past decade. The results show an abrupt change in the meaning of decolonisation starting in the middle of the 2010s with an increased focus on cultural and institutional matters, particularly in right-leaning media. Surprisingly, the editorial slant of broadsheets seemingly only had at best moderate effects on tone, while headlines in right-wing tabloids were significantly more negative. Articles covering cultural aspects of decolonisation were substantially more negative than those applying a more traditional, territorial definition of decolonisation. Given the influence of newspaper media on public and private opinions, we discuss the heritage implications of these findings and suggest avenues for further investigation.

29Intricate Genealogies: What is Said in the Epic Poem about Nogay People?OpenAlex

Quwatbek Duysen, Botagoz Suıyerkul, Marzhan Nurmanova
This article examines communication blueprints of the Forty Knights of Steppe epic poem (original ti-tle in Kazakh "Қырымның қырық батыры"). The Forty Knights of Steppe is a cultural heritage of Turkic nations. The epic poem’s text is an assemblage of ballads that are connected with each other. It consists of thirty-five ballads and depicts relationships of around two hundred characters. The epic poem's events have a historical origin that links to the Nogay nobility from the Golden Horde period. The main event in the epic poem develops around Edige and his descendants. The historical aspect of this epic poem positions it with the paragons of heroic literature such as "Shahnameh" and "Manas". The epic poem’s text was first fully recorded in 1942 in Almaty by Muryn Zhyrau (real name 'Tilegen Sengirbekuly') by specialists from the Kazakh SSR Academy of Sciences. Unfortunately, the epic poem stayed unpublished for more than 60 years. The full text of the epic poem was published in 2005 after restoring the independence of the Republic of Kazakhstan. This research discovers the epic poem by analyzing their network. It focuses on mapping and visualizing links between characters described in the text. The authors used the methods of text mining, proposition, and semantic triangles for assembling information about contacts between characters into the databases. The databases include basic information, such as the names of characters and their contacts with other characters. The databases also contain information about the direction of every contact, relationship statuses of contactors, contactors' roles inside of genealogical lineage, kinship degrees of contactors inside their genealogies, and contactors' gender. In addition, the databases include information about the connec-tion of contactors to genealogies, and information about the contacts of contactors outside of their genealo-gies. All databases were built in spatial data analysis compatible format. The databases were launched through data visualization software by switching on the environmental settings required for this type of research. The authors used Gephi data visualization software for this research to visualize databases. The visualization shows an intricate communicative network that covers almost all characters. The visualization demonstrates two types of contacts between characters of the epic poem. The first type is a structured type of contacts. This type of contacts develops inside of genealogies. It is a common type of contacts in the epic poem, and it follows the hierarchical order. The second type is a class-based type of contacts. This type of contacts ties characters from different genealogies and does not correspond to the hierarchical order. However, this type of contacts depends on the characters' roles and statuses inside their genealogies. Thereby, the second type of contacts is connecting different genealogies into whole network. The types of contacts demonstrate two levels of social communication in the narration of the epic poem. Considering the visualization results and according to the historical origins of the characters, the authors argue that the text of the epic poem shows the patterns of social communication typical for Central Eurasian medieval nomad cultures in the Golden Horde period. On the other side, the stratification into two types of contacts may demonstrate the levels of bureaucracy in social communication in the same historical period. The authors suggest that the social networks of the epic poem describe the transitional form of the chiefdom society with communication levels typical for tribal and chiefdom societies. In general, the paper's authors suppose that the example of this research on the epic poem's communication structure may give more data to understand the correlation between language and society.

30Causal BERT: Language Models for Causality Detection Between Events Expressed in TextOpenAlex

Vivek Khetan, Roshni Ramnani, Mayuresh Anand, et al.

31Negated bio-events: analysis and identificationOpenAlex

Raheel Nawaz, Paul M. Thompson, Sophia Ananiadou
BACKGROUND: Negation occurs frequently in scientific literature, especially in biomedical literature. It has previously been reported that around 13% of sentences found in biomedical research articles contain negation. Historically, the main motivation for identifying negated events has been to ensure their exclusion from lists of extracted interactions. However, recently, there has been a growing interest in negative results, which has resulted in negation detection being identified as a key challenge in biomedical relation extraction. In this article, we focus on the problem of identifying negated bio-events, given gold standard event annotations. RESULTS: We have conducted a detailed analysis of three open access bio-event corpora containing negation information (i.e., GENIA Event, BioInfer and BioNLP'09 ST), and have identified the main types of negated bio-events. We have analysed the key aspects of a machine learning solution to the problem of detecting negated events, including selection of negation cues, feature engineering and the choice of learning algorithm. Combining the best solutions for each aspect of the problem, we propose a novel framework for the identification of negated bio-events. We have evaluated our system on each of the three open access corpora mentioned above. The performance of the system significantly surpasses the best results previously reported on the BioNLP'09 ST corpus, and achieves even better results on the GENIA Event and BioInfer corpora, both of which contain more varied and complex events. CONCLUSIONS: Recently, in the field of biomedical text mining, the development and enhancement of event-based systems has received significant interest. The ability to identify negated events is a key performance element for these systems. We have conducted the first detailed study on the analysis and identification of negated bio-events. Our proposed framework can be integrated with state-of-the-art event extraction systems. The resulting systems will be able to extract bio-events with attached polarities from textual documents, which can serve as the foundation for more elaborate systems that are able to detect mutually contradicting bio-events.

32Injecting Temporal-Aware Knowledge in Historical Named Entity RecognitionOpenAlex

Carlos-Emiliano González-Gallardo, Emanuela Boroş, Edward Giamphy, et al.

33Factoid-based prosopography and computer ontologies: towards an integrated approachOpenAlex

Michele Pasin, John Bradley
Structured Prosopography provides a formal model for representing prosopography: a branch of historical research that traditionally has focused on the identification of people that appear in historical sources. Since the 1990s, KCL’s Department of Digital Humanities has been involved in the development of structured prosopographical databases using a general ‘factoid-oriented’ model of structure that links people to the information about them via spots in primary sources that assert that information. Recent developments, particularly the World Wide Web, and its related technologies around the Semantic Web, have promoted the possibility to both interconnecting dispersed data, and allowing it to be queried semantically. To the purpose of making available our prosopographical databases on the Semantic Web, in this article we review the principles behind our established factoid-based model and reformulate it using a more interoperable approach, based on knowledge representation principles and formal ontologies. In particular, we are going to focus primarily on a high-level semantic analysis of the factoid notion, on its relation to other cultural heritage standards such as CIDOC-CRM, and on the modularity and extensibility of the proposed solutions.

34Cultural Heritage Information Retrieval: Past, Present, and Future TrendsOpenAlex

Babak Ranjgar, Abolghasem Sadeghi‐Niaraki, Maryam Shakeri, et al.
Knowledge organization and development of better information retrieval techniques were of great importance from a very early time period in human history. The need has grown high for such systems with the advent of digitization and the web era. Computer systems and web have offered easier retrieval of information in almost no time. However, as the amount of data increased, these systems were not able to work well in terms of accuracy and precision of retrieval. Semantic Web concept was introduced to overcome the issue by converting the web of documents to a web of data. Semantic Web technologies makes data machine-understandable so that information retrieval can be more precise and accurate. The Cultural Heritage (CH) community, with the goal of preserving and dissemination of the historical information to people and society, is one of the first domains to adopt Semantic Web recommendations and technologies, which can provide interoperability between various organizations by creating a shared understanding in the community. The data in the CH domain differs widely with types and formats. Also, a lot of organizations and experts from various fields interact through different processes within this community. Due to the mentioned needs, the CH community employed Semantic Web technologies step by step along its evolution process for better knowledge management and a uniform understanding among the community. In this study, we present a comprehensive conceptual framework that spans cultural heritage, information modeling, and information retrieval. Our model addresses early solutions in knowledge organization systems, highlighting the evolution from classification systems and controlled vocabularies to the significance of metadata schemas. We delve into the limitations of traditional knowledge organization systems and the necessity of formal ontologies, particularly in the cultural heritage domain. The comparative analysis of CRM vs. EDM, ontology-based metadata interoperability, and ontology technologies elucidate our contributions to the field. This paper outlines the process from the initial steps of adopting Semantic Web technologies in the CH domain to the latest developments in CH information retrieval. In this paper, we also reviewed intelligent applications and services developed in the CH domain after establishing semantic data models and Knowledge Organization Systems. Finally, challenges and possible future research directions are discussed. The findings revealed that GLAMs (Galleries, Libraries, Archives, and Museums) are excellent and comprehensive sources of CH information. The CH community has put in a lot of time and effort to develop data models and knowledge organization tools; now it’s time to use this valuable resource to construct smart applications that are still in their early phases. This could benefit the CH industry even more.

35Preliminary Study on the Knowledge Graph Construction of Chinese Ancient History and CultureOpenAlex

Shuang Liu, Hui Yang, Jiayi Li, et al.
The domestic population has paid increasing attention to ancient Chinese history and culture with the continuous improvement of people’s living standards, the rapid economic growth, and the rapid advancement of information science and technology. The use of information technology has been proven to promote the spread and development of historical culture, and it is becoming a necessary means to promote our traditional culture. This paper will build a knowledge graph of ancient Chinese history and culture in order to facilitate the public to more quickly and accurately understand the relevant knowledge of ancient Chinese history and culture. The construction process is as follows: firstly, use crawler technology to obtain text and table data related to ancient history and culture on Baidu Encyclopedia (similar to Wikipedia) and ancient Chinese history and culture related pages. Among them, the crawler technology crawls the semi-structured data in the information box (InfoBox) in the Baidu Encyclopedia to directly construct the triples required for the knowledge graph, crawls the introductory text information of the entries in Baidu Encyclopedia, and specialized historical and cultural websites (history Chunqiu.com, On History.com) to extract unstructured entities and relationships. Secondly, entity recognition and relationship extraction are performed on an unstructured text. The entity recognition part uses the Bidirectional Long Short-Term Memory-Convolutional Neural Networks-Conditions Random Field (BiLSTM-CNN-CRF) model for entity extraction. The relationship extraction between entities is performed by using the open source tool DeepKE (information extraction tool with language recognition ability developed by Zhejiang University) to extract the relationships between entities. After obtaining the entity and the relationship between the entities, supplement it with the triple data that were constructed from the semi-structured data in the existing knowledge base and Baidu Encyclopedia information box. Subsequently, the ontology construction and the quality evaluation of the entire constructed knowledge graph are performed to form the final knowledge graph of ancient Chinese history and culture.

36A Survey on Knowledge Graphs: Representation, Acquisition, and ApplicationsOpenAlex

Shaoxiong Ji, Shirui Pan, Erik Cambria, et al.
Human knowledge provides a formal understanding of the world. Knowledge graphs that represent structural relations between entities have become an increasingly popular research direction toward cognition and human-level intelligence. In this survey, we provide a comprehensive review of the knowledge graph covering overall research topics about: 1) knowledge graph representation learning; 2) knowledge acquisition and completion; 3) temporal knowledge graph; and 4) knowledge-aware applications and summarize recent breakthroughs and perspective directions to facilitate future research. We propose a full-view categorization and new taxonomies on these topics. Knowledge graph embedding is organized from four aspects of representation space, scoring function, encoding models, and auxiliary information. For knowledge acquisition, especially knowledge graph completion, embedding methods, path inference, and logical rule reasoning are reviewed. We further explore several emerging topics, including metarelational learning, commonsense reasoning, and temporal knowledge graphs. To facilitate future research on knowledge graphs, we also provide a curated collection of data sets and open-source libraries on different tasks. In the end, we have a thorough outlook on several promising research directions.

37Knowledge Graphs: Opportunities and ChallengesOpenAlex

Ciyuan Peng, Feng Xia, Mehdi Naseriparsa, et al.
With the explosive growth of artificial intelligence (AI) and big data, it has become vitally important to organize and represent the enormous volume of knowledge appropriately. As graph data, knowledge graphs accumulate and convey knowledge of the real world. It has been well-recognized that knowledge graphs effectively represent complex information; hence, they rapidly gain the attention of academia and industry in recent years. Thus to develop a deeper understanding of knowledge graphs, this paper presents a systematic overview of this field. Specifically, we focus on the opportunities and challenges of knowledge graphs. We first review the opportunities of knowledge graphs in terms of two aspects: (1) AI systems built upon knowledge graphs; (2) potential application fields of knowledge graphs. Then, we thoroughly discuss severe technical challenges in this field, such as knowledge graph embeddings, knowledge acquisition, knowledge graph completion, knowledge fusion, and knowledge reasoning. We expect that this survey will shed new light on future research and the development of knowledge graphs.

38Semantic technologies for historical research: A surveyOpenAlex

Albert Meroño-Peñuela, A. Ashkpour, Marieke van Erp, et al.
During the nineties of the last century, historians and computer scientists created together a research agenda around the life cycle of historical information. It comprised the tasks of creation, design, enrichment, editing, retrieval, analysis and presentation of historical information with help o f information technology. They also identified a number of problems and challenges in this field, some of them closely related to semantics and meaning. In this survey paper we study the joint work of historians and computer scientists in the use of Semantic Web methods and technologies in historical research. We analyse to what extent these contributions help in solving the open problems in the agenda of historians, and we describe open challenges and possible lines of research pushing further a still young, but promising, historical Semantic Web.

39Integrating building information modelling and semantic web technologies for the management of built heritage informationOpenAlex

Pieter Pauwels, Rens Bod, Danilo Di Mascio, et al.
The historical built environment is acknowledged as a valuable but complex material and cultural resource that needs to be preserved. Digital technologies give the opportunity to improve and expand the comprehension of the complex artefacts present in this built environment. Building information modelling (BIM) and semantic web technologies are two technologies that are often used for the documentation of the built environment and of cultural heritage resources. With our research, we investigate to what extent those technologies can be integrated and which advantages this combination can produce for the analysis and interpretation of our built environment. In this paper, we present the application of BIM software and semantic web technologies to a case study: the Book Tower in Ghent, Belgium. The Book Tower is one of the most important early 20th century buildings in the city of Ghent. Through the paper we will show how BIM and semantic web technologies were integrated, which advantages this combination can produce and which future developments could be considered. The recorded information can be essential to plan and manage a recovery plan and/or a maintenance program taking into consideration also aspects linked to cultural diversity and environmental sustainability.

40The Aggregate Dutch Historical CensusesOpenAlex

A. Ashkpour, Albert Meroño-Peñuela, Kees Mandemakers
Historical censuses have an enormous potential for research. In order to fully use this potential, harmonization of these censuses is essential. During the last decades, enormous efforts have been undertaken in digitizing the published aggregated outcomes of the Dutch historical censuses . Although the accessibility has been improved enormously, researchers must cope with hundreds of heterogeneous and disconnected Excel tables. As a result, the census is still for the most part an untapped source of information. The authors describe the main harmonization challenges of the census and how they work toward one harmonized dataset. They propose a specific approach and model in creating an interlinked census dataset in the Semantic Web using the Resource Description Framework technology.

41An Intelligent Question Answering System of the Liao Dynasty Based on Knowledge GraphOpenAlex

Shuang Liu, Nannan Tan, Hui Yang, et al.
Abstract The Liao Dynasty was a minority regime established by the Khitan on the grasslands of northern China. To promote and spread the cultural knowledge of the Liao Dynasty, an intelligent question-and-answer system is constructed based on the knowledge graph in the historical and cultural field of the Liao Dynasty. In the traditional question answering system, the quality of answers was not high due to incomplete data and distinctive vocabulary. To solve this problem, a combination method of Liao Dynasty question-and-answer database and KB is proposed to realize knowledge graph question answering, and a joint model of Siamese LSTM and fusion MatchPyramid is proposed for semantic matching between questions in the question-and-answer database. With the joint model, it is easy to perform semantic matching by fusing sentence-level and word-level interactive features through LSTM and MatchPyramid. Furthermore, the question sentence with the same semantics as the user input question sentence is retrieved in the question-and-answer database, and the answer corresponding to the question sentence is returned as the result. The experimental results show that our proposed method has achieved relatively good performance in the historical domain of the Liao Dynasty and the open-domain knowledge graph, and improved the accuracy of question and answer.

42A Semantic Search Engine for Historical Handwritten Document ImagesOpenAlex

Vuong M. Ngo, Gary Munnelly, Fabrizio Orlandi, et al.
Abstract A very large number of historical manuscript collections are available in image formats and require extensive manual processing in order to search through them. So, we propose and build a search engine for automatically storing, indexing and efficiently searching the manuscript images. Firstly, a handwritten text recognition technique is used to convert the images into textual representations. In the next steps, we apply the named entity recognition and historical knowledge graph to build a semantic search model, which can understand the user’s intent in the query and the contextual meaning of concepts in documents, to return correctly the transcriptions and their corresponding images for users.

43Musical heritage historical entity linkingOpenAlex

Arianna Graciotti, Nicolas Lazzari, Valentina Presutti, et al.
Abstract Linking named entities occurring in text to their corresponding entity in a Knowledge Base (KB) is challenging, especially when dealing with historical texts. In this work, we introduce Musical Heritage named Entities Recognition, Classification and Linking ( mhercl ), a novel benchmark consisting of manually annotated sentences extrapolated from historical periodicals of the music domain. mhercl contains named entities under-represented or absent in the most famous KBs. We experiment with several State-of-the-Art models on the Entity Linking (EL) task and show that mhercl is a challenging dataset for all of them. We propose a novel unsupervised EL model and a method to extend supervised entity linkers by using Knowledge Graphs (KGs) to tackle the main difficulties posed by historical documents. Our experiments reveal that relying on unsupervised techniques and improving models with logical constraints based on KGs and heuristics to predict entities (entities not represented in the KB of reference) results in better EL performance on historical documents.

44Exploring Cultural Heritage Knowledge Graphs – Case of Correspondence Networks in Grand Duchy of Finland 1809–1917OpenAlex

Henna Poikkimäki, Petri Leskinen, Eero Hyvönen
This paper argues for using methods and tools of Network Analysis (NA) to study contents of knowledge graphs (KG) in Digital Humanities (DH) research. As a case study, social and correspondence networks in the Grand Duchy of Finland 1809–1917 are considered with a focus on prosopographical data about historical people and, in particular, their correspondences (epistolary data). Letters have been an important form of communication, and networks based on letter metadata, letter’s content, and related biographical information can be used for rebuilding and analyzing historical social networks and for studying the flow of ideas and information. In correspondence network analysis, ego-networks focusing on only one person and his correspondents are common due to the nature of letter collections. Combining letter collections and biographical data helps move from ego-centric network approach towards sociocentric networks, as the larger network starts to emerge when letter collections from many individuals are brought together, although analyses still suffer from missing data. In this paper, we present results of the Constellations of Correspondence (CoCo) project that so far has created a KG of over million letters exchanged during 1809–1917 in Finland, reusing data of prosopographical KGs of the same period of time.

45CohortVA: A Visual Analytic System for Interactive Exploration of Cohorts Based on Historical DataOpenAlex

Wei Zhang, Wong Kam-Kwai, Xumeng Wang, et al.
In history research, cohort analysis seeks to identify social structures and figure mobilities by studying the group-based behavior of historical figures. Prior works mainly employ automatic data mining approaches, lacking effective visual explanation. In this paper, we present CohortVA, an interactive visual analytic approach that enables historians to incorporate expertise and insight into the iterative exploration process. The kernel of CohortVA is a novel identification model that generates candidate cohorts and constructs cohort features by means of pre-built knowledge graphs constructed from large-scale history databases. We propose a set of coordinated views to illustrate identified cohorts and features coupled with historical events and figure profiles. Two case studies and interviews with historians demonstrate that CohortVA can greatly enhance the capabilities of cohort identifications, figure authentications, and hypothesis generation.

46Adding Domain Knowledge to Improve Entity Resolution in 17th and 18th Century Amsterdam Archival RecordsOpenAlex

Jurian Baas, Leon van Wissen, J. Reinders, et al.
The problem of entity resolution is central in the field of Digital Humanities. It is also one of the major issues in the Golden Agents project, which aims at creating an infrastructure that enables researchers to search for patterns that span across decentralised knowledge graphs from cultural heritage institutes. To this end, we created a method to perform entity resolution on complex historical knowledge graphs. In previous work, we encoded and embedded the relevant (duplicate) entities in a vector space to derive similarities between them based on sharing a similar context in RDF graphs. In some cases, however, available domain knowledge or rational axioms can be applied to improve entity resolution performance. We show how domain knowledge and rational axioms relevant to the task at hand can be expressed as (probabilistic) rules, and how the information derived from rule application can be combined with quantitative information from the embedding. In this work, we perform our entity resolution method on two data sets. First, we apply it to a data set for which we have a detailed ground truth for validation. This experiment shows that the combination of embedding and the application of domain knowledge and rational axioms leads to improved resolution performance. Second, we perform a case study by applying our method to a larger data set for which there is no ground truth and where the outcome is subsequently validated by a domain expert. Results of this demonstrate that our method achieves a very high precision.

47Adding Domain Knowledge to Improve Entity Resolution in 17th and 18th Century Amsterdam Archival RecordsOpenAlex

Baas, J., van Wissen, Leon, Reinders, Jirsi, et al.
The problem of entity resolution is central in the field of Digital Humanities. It is also one of the major issues in the Golden Agents project, which aims at creating an infrastructure that enables researchers to search for patterns that span across decentralised knowledge graphs from cultural heritage institutes. To this end, we created a method to perform entity resolution on complex historical knowledge graphs. In previous work, we encoded and embedded the relevant (duplicate) entities in a vector space to derive similarities between them based on sharing a similar context in RDF graphs. In some cases, however, available domain knowledge or rational axioms can be applied to improve entity resolution performance. We show how domain knowledge and rational axioms relevant to the task at hand can be expressed as (probabilistic) rules, and how the information derived from rule application can be combined with quantitative information from the embedding. In this work, we perform our entity resolution method on two data sets. First, we apply it to a data set for which we have a detailed ground truth for validation. This experiment shows that the combination of embedding and the application of domain knowledge and rational axioms leads to improved resolution performance. Second, we perform a case study by applying our method to a larger data set for which there is no ground truth and where the outcome is subsequently validated by a domain expert. Results of this demonstrate that our method achieves a very high precision.

48TKGR-RHETNE: A New Temporal Knowledge Graph Reasoning Model via Jointly Modeling Relevant Historical Event and Temporal Neighborhood Event ContextOpenAlex

Jinze Sun, Yongpan Sheng, Zhan Ling, et al.

49Graph Hawkes Transformer for Extrapolated Reasoning on Temporal Knowledge GraphsOpenAlex

Haohai Sun, Shangyi Geng, Jialun Zhong, et al.
Temporal Knowledge Graph (TKG) reasoning has attracted increasing attention due to its enormous potential value, and the critical issue is how to model the complex temporal structure information effectively. Recent studies use the method of encoding graph snapshots into hidden vector space and then performing heuristic deductions, which perform well on the task of entity prediction. However, these approaches cannot predict when an event will occur and have the following limitations: 1) there are many facts not related to the query that can confuse the model; 2) there exists information forgetting caused by long-term evolutionary processes. To this end, we propose a Graph Hawkes Transformer (GHT) for both TKG entity prediction and time prediction tasks in the future time. In GHT, there are two variants of Transformer, which capture the instantaneous structural information and temporal evolution information, respectively, and a new relational continuous-time encoding function to facilitate feature evolution with the Hawkes process. Extensive experiments on four public datasets demonstrate its superior performance, especially on long-term evolutionary tasks.

50Visual Reasoning for Uncertainty in Spatio-Temporal Events of Historical FiguresOpenAlex

Wei Zhang, Siwei Tan, Siming Chen, et al.
The development of digitized humanity information provides a new perspective on data-oriented studies of history. Many previous studies have ignored uncertainty in the exploration of historical figures and events, which has limited the capability of researchers to capture complex processes associated with historical phenomena. We propose a visual reasoning system to support visual reasoning of uncertainty associated with spatio-temporal events of historical figures based on data from the China Biographical Database Project. We build a knowledge graph of entities extracted from a historical database to capture uncertainty generated by missing data and error. The proposed system uses an overview of chronology, a map view, and an interpersonal relation matrix to describe and analyse heterogeneous information of events. The system also includes uncertainty visualization to identify uncertain events with missing or imprecise spatio-temporal information. Results from case studies and expert evaluations suggest that the visual reasoning system is able to quantify and reduce uncertainty generated by the data.

51Modeling semantic business trajectories of territories for multidisciplinary studies through controlled vocabulariesOpenAlex

Muhammad Arslan, Christophe Cruz
Environmental changes influence society and this impact extends to businesses across several sectors. The risks linked with environmental changes have both advantages and disadvantages for companies. Therefore, business analysts must understand how their companies should deal with the opportunities and challenges that environmental changes present. On the other hand, environmental analysts working on different aspects (e.g. atmospheric sciences, ecology, geosciences, social sciences, etc.) of an environment for human health and well-being are also interested in studying the impact of businesses (e.g. the launch of a factory) on society. Though, business and environmental analysts are from different educational backgrounds. They work on different vocabularies and data formats for analyzing data. These differences create semantic heterogeneities to understand key concepts from data. This makes understanding cross-domain data a complex and time-consuming process for them, eventually creating a knowledge barrier. To address this knowledge barrier, this work presents a proof-of-concept modeling of business trajectories using online news data. The proposed method consists of five stages: a) news data preprocessing, and topic model training for identifying the relevant thematic concepts related to business trajectories using the historical dataset of news articles, b) semantic enrichment of thematic concepts using the controlled vocabularies, c) processing latest news articles using the trained topic model to acquire the most similar thematic concepts, and obtaining the location and temporal entities using the Named-Entity Recognition model (NER) process, d) constructing a knowledge graph using the collected thematic, spatial and temporal entities, along with enrichment of the relevant environmental data, and e) visualizing the semantic trajectories from knowledge graphs for understanding insights in data. This article presents the potential for retrieving the same trajectory data using different business and environmental viewpoints for multidisciplinary data analysis.

52DKNOpenAlex

Hongwei Wang, Fuzheng Zhang, Xing Xie, et al.
Online news recommender systems aim to address the information explosion of news and make personalized recommendation for users. In general, news language is highly condensed, full of knowledge entities and common sense. However, existing methods are unaware of such external knowledge and cannot fully discover latent knowledge-level connections among news. The recommended results for a user are consequently limited to simple patterns and cannot be extended reasonably. To solve the above problem, in this paper, we propose a deep knowledge-aware network (DKN) that incorporates knowledge graph representation into news recommendation. DKN is a content-based deep recommendation framework for click-through rate prediction. The key component of DKN is a multi-channel and word-entity-aligned knowledge-aware convolutional neural network (KCNN) that fuses semantic-level and knowledge-level representations of news. KCNN treats words and entities as multiple channels, and explicitly keeps their alignment relationship during convolution. In addition, to address users» diverse interests, we also design an attention module in DKN to dynamically aggregate a user»s history with respect to current candidate news. Through extensive experiments on a real online news platform, we demonstrate that DKN achieves substantial gains over state-of-the-art deep recommendation models. We also validate the efficacy of the usage of knowledge in DKN.

53A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and DirectionsOpenAlex

Chen Gao, Yu Zheng, Nian Li, et al.
Recommender system is one of the most important information services on today’s Internet. Recently, graph neural networks have become the new state-of-the-art approach to recommender systems. In this survey, we conduct a comprehensive review of the literature on graph neural network-based recommender systems. We first introduce the background and the history of the development of both recommender systems and graph neural networks. For recommender systems, in general, there are four aspects for categorizing existing works: stage, scenario, objective, and application. For graph neural networks, the existing methods consist of two categories: spectral models and spatial ones. We then discuss the motivation of applying graph neural networks into recommender systems, mainly consisting of the high-order connectivity, the structural property of data and the enhanced supervision signal. We then systematically analyze the challenges in graph construction, embedding propagation/aggregation, model optimization, and computation efficiency. Afterward and primarily, we provide a comprehensive overview of a multitude of existing works of graph neural network-based recommender systems, following the taxonomy above. Finally, we raise discussions on the open problems and promising future directions in this area. We summarize the representative papers along with their code repositories in https://github.com/tsinghua-fib-lab/GNN-Recommender-Systems .

54Accuracy of 1908 high to medium scale cartography of Rome and its surroundings and related georeferencing problemsOpenAlex

Valerio Baiocchi, Keti Lelo
Preliminary attempts to georeference maps of early twentieth century made by the Military Geographic Institute (IGM, the Italian geodetic agency) for the city of Rome and its surroundings, reported residual errors larger than errors observed on similar maps. Previous studies carried out on one or two century older maps of the same area, showed similar or even smaller errors (Baiocchi and Lelo 2005).Six sheets of the “City of Rome and its surroundings” map in scale 1:5 000 dated 1908 have been studied. The identified errors can be referred to the different system of geodetic projection and geodetic datum or to the derivation of some details from maps at smaller scale, but in this case historic documents seem to suggest a different explanation.Parameters useful to perform the transformation of the geodetic systems used in historical maps to modern systems are not known; for this reason until now the various attempts of georeferencing maps of this type were based on collimation of points recognizable on modern cartographies such as corners of historical buildings. This method has often given unsatisfactory results; therefore it was decided to proceed by determining the parameters for the transformation of geodetic datum.The history of geodetic systems used in Italy at the beginning of the 20th century is complex and, in the past, this has led some researcher to misinterpretations. For this reason a full explanation of geodetic systems used in Italy in this period is reported below. Since the parameters of the projection used for the maps in our case study are not known for sure, the reprojection was considered the only way for a correct georeferencing.

55Border Reconstruction of Bosnia and Herzegovina's Access to the Adriatic Sea at Sutorina by Consulting Old MapsOpenAlex

Nedim Tuno, Admir Mulahusić, Mithad Kozličić, et al.
The paper presents the results of scientific research into the former southernmost Bosnian border by analyzing historical maps. In cartographic representations of the area (created between the mid-17th and mid-20th centuries), state and administrative boundary lines are clearly demarcated. They provide indisputable proof that the Sutorina area belonged to Bosnia and Herzegovina through many centuries, providing access to the Adriatic Sea. The maps presented follow the course of the historical changes in the area which shaped its borders. The extent of the narrow Sutorina corridor was observed by combining data on boundary lines taken from historical maps with the current situation in the area. The technique of georeferencing old maps based on a genetic algorithm was developed for this purpose.

56Positional accuracy, positional uncertainty, and feature change detection in historical maps: Results of an experimentOpenAlex

Michele Tucci, Alberto Giordano

57A dataset the geographic wine culture in Tang-Song poetry in the Yangtze River BasinOpenAlex

Xueying Zhang, Xinchen SUN, Minxuan FENG, et al.
<p indent="0mm">In China, wine culture has a history of thousands of years and is deeply embedded in the fabric of Chinese civilization. Starting with the <italic>Book of Songs</italic> (<italic>Shijing</italic>), generations of scholars have expressed the spiritual connotations of wine through poetry, forming a unique “fusion of poetry and wine culture”. As a major cultural cradle of China, the Yangtze River Basin has long been an important source of inspiration for literary creation. The close integration of wine and poetry in this region has become one of its distinctive cultural features. To explore this phenomenon, this study collected and organized place names, wine culture vocabulary, and poetry related to wine from Tang and Song dynasties, drawing upon classical texts and online repositories of ancient poetry, resulting in a dataset the geographic distribution of wine culture in Poetry of Tang and Song dynasties in the Yangtze River Basin. This dataset consists of a wine culture classification and code table, a glossary of ancient Chinese wine culture terms, a table mapping ancient and modern place name in the Yangtze River Basin, a metadata table of Tang-Song poetry in the Yangtze River Basin, a table extracting place names and wine culture vocabulary from Tang-Song poetry, and a table mapping place names in the Tang-Song poetry with their corresponding wine culture classifications. Data entries have undergone multiple rounds of manual verification to ensure its completeness and accuracy. Using this dataset, researchers can extract temporal and spatial information from poetry and uncover the deep connection between wine culture and poetic creation by mapping wine culture terms to their corresponding classifications. The dataset provides new data support for the dissemination and study of wine culture in the Yangtze River Basin.

58Common Language for Accessibility, Interoperability, and Reusability in Historical DemographyOpenAlex

Rick J. Mourits, Tim Riswick, Rombert Stapel
Abstract One of the biggest challenges in the transition to open science is making data interoperable. Ideally, existing schemas and vocabularies are (re-)used to describe data, but these are generally problematic for historical data, as they exclude historical concepts and are insensitive to temporal variations in meaning. Therefore, the subdiscipline of historical demography has designed its own schemas and vocabularies to standardize historical data, as researchers require them to make and study large-scale reconstructions of populations and life courses. We introduce a web environment called CLAIR-HD that helps researchers to find vocabularies to standardize historical demographic data, and determine lacunae in the standardization of data within the field of historical demography.

59Methodology for classifying and detecting intra-urban land use change: A case study of Changchun city during the last 100 yearsOpenAlex

Quanqin Shao
Urban internal function and structure which is represented by intra-urban land use information is important spa- tial-explicitly information in urban geography research and urban planning or management; however, its extraction of fine-scale spatial information is still a difficult question in geographic information science (GIS). To overcome the difficulty in detecting intra-urban land use types using remotely sensed data alone and the positioning errors of mismatch in overlaying multi-temporal images, we have developed a digital reconstructing method to classify and detect intra-urban land use change through combining the hierarchical classification and object-oriented segmentation methods. This paper traces back to the historical developing process of Changchun city based on the above methods to classify and detect the detail intra-urban land use types including residential land, commercial land, industrial land, roads, water body and urban green space. The study has resolved the key problems on accurate spatial positioning from multi-scale and different data sources, classifying intra-urban land from function characteristics and detecting historical developing process of metropolis. Furthermore, we combined the SPOT5 imagery, 1:10000 topographic maps, historical maps, urban planning map and other auxiliary data of Changchun City to classify and detect its intra-urban land use change from 1905 to 2003. The results indicate that the methods are significant and effective in classifying and detecting urban land use spatial information through combining the human-computer interactive interpretation with the expert-based knowledge from different source data. The methods can not only enhance the accuracy of urban land use classification, but also improve spatial information extraction efficiency and positioning accuracy of multi-temporal spatial overlaying analysis.

60Methodology for classifying and detecting intra-urban land use change: A case study of Changchun city during the last 100 yearsOpenAlex

匡文慧, 中国科学院地理科学与资源研究所, 张树文, et al.
Urban internal function and structure which is represented by intra-urban land use information is important spatial-explicitly information in urban geography research and urban planning or management; however, its extraction of fine-scale spatial information is still a difficult question in geographic information science (GIS).To overcome the difficulty in detecting intra-urban land use types using remotely sensed data alone and the positioning errors of mismatch in overlaying multi-temporal images, we have developed a digital reconstructing method to classify and detect intra-urban land use change through combining the "hierarchical classification" and "object-oriented segmentation" methods.This paper traces back to the historical developing process of Changchun city based on the above methods to classify and detect the detail intra-urban land use types including residential land, commercial land, industrial land, roads, water body and urban green space.The study has resolved the key problems on accurate spatial positioning from multi-scale and different data sources, classifying intra-urban land from function characteristics and detecting historical developing process of metropolis.Furthermore, we combined the SPOT5 imagery, 1:10000 topographic maps, historical maps, urban planning map and other auxiliary data of Changchun City to classify and detect its intra-urban land use change from 1905 to 2003.The results indicate that the methods are significant and effective in classifying and detecting urban land use spatial information through combining the human-computer interactive interpretation with the expert-based knowledge from different source data.The methods can not only enhance the accuracy of urban land use classification, but also improve spatial information extraction efficiency and positioning accuracy of multi-temporal spatial overlaying analysis.

61Mapping Madness: HGIS and the Analysis of Irish Patient RecordsOpenAlex

Oonagh Walsh, Stuart Clancy
Abstract The Connaught District Lunatic Asylum (CDLA) opened at Ballinasloe, Co. Galway in 1833 as one of the first of a nationwide network of Irish District Asylums. Intended to serve the curable pauper lunatics of the counties of Mayo, Sligo, Leitrim, Galway, and Roscommon, the institution found itself at the heart of significant social, economic, and political change in the West of Ireland. From its opening, the asylum maintained a full and complex series of records that provide an exceptional level of detail on a cohort – the very poor and illiterate population of Connaught – who otherwise often lived and died unrecorded on the margins of Irish society. The CDLA admission records include information on age, sex, occupation, education, religion, marital status, places of origin and residence, migration, and family structures as well as the medical information, both mental and physical, required for treatment in the asylum. This paper will examine the potential benefits of implementing spatial epidemiological methods into historical studies of mental illness. Using a database of patient records this paper will conduct a demographic analysis of a sample of the population of the CDLA. The paper will outline the process of transforming the data extracted from these records into visual maps using Historical GIS (HGIS). Using the geographic co-ordinates of these two locations, the unique patterns of movement of those that entered the asylum can be mapped using GIS. These maps enable the examination of the socio-spatial processes which affected the marginalised population of the asylum. (The sources used in this paper come from the Connaught District Lunatic Asylum records, held at the National Archives of Ireland in Dublin. As the material is drawn from committal warrants of patients admitted in 1889, it falls outside of the ‘100 year rule’ for accessing sensitive historic patient information in Ireland.)

62Density and dispersion: the co-development of land use and rail in LondonOpenAlex

David Levinson
This article examines the changes that occurred in the rail network and density of population in London during the 19th and 20th centuries. It aims to disentangle the 'chicken and egg' problem of which came first, network or land development, through a set of statistical analyses clearly distinguishing events by order. Using panel data representing the 33 boroughs of London over each decade from 1871 to 2001, the research finds that there is a positive feedback effect between population density and network density. Additional rail stations (either Underground or surface) are positive factors leading to subsequent increases in population in the suburbs of London, while additional population density is a factor in subsequently deploying more rail. These effects differ in central London, where the additional accessibility produced by rail led to commercial development and concomitant depopulation. There are also differences in the effects associated with surface rail stations and Underground stations, as the Underground was able to get into central London in a way that surface rail could not. However, the two networks were weak (and statistically insignificant) substitutes for each other in the suburbs, while the density of surface rail stations was a complement to the Underground in the center, though not vice versa.

63Field Model-Based Cultural Diffusion Patterns and GIS Spatial Analysis Study on the Spatial Diffusion Patterns of Qijia Culture in ChinaOpenAlex

Yuanyuan Wang, Nai’ang Wang, Xuepeng Zhao, et al.
Cultural diffusion is one of the core issues among researchers in the field of cultural geography. This study aimed to examine the spatial diffusion patterns of the Qijia culture (QJC) to clarify the origin and formation process of Chinese field model-based cultural diffusion patterns (FM-CDP) and geographic information system (GIS) spatial analysis methods. It used the point data of Qijia cultural sites without time information and combined them with the relevant records of Qijia cultural and historical documents, as well as archaeological excavation materials. Starting with the spatial location information of cultural distribution, it comprehensively analysed the cultural hearth, regions, diffusion patterns, and diffusion paths. The results indicated the following. (1) The QJC’s heart is in the southeast of Gansu Province, where the Shizhaocun and Xishanping sites are distributed. (2) Five different levels of cultural regions were formed, which demonstrated different diffusion patterns at different regional scales. On a large regional scale, many cultural regions belong to relocation diffusion patterns. Meanwhile, at the small regional scale (in the Gansu–Qinghai region), there are two patterns of diffusion: expansion diffusion and relocation diffusion; however, the expansion diffusion pattern is the main one. (3) Based on the relationship between the QJC, altitude, and the water system, the culture also has the characteristics of diffusion to low altitude areas and a pattern of diffusion along water systems. (4) There is a circular structure of the core, periphery, and fringe regions of the QJC. Finally, (5) the dry and cold climate around 4000a B.P., the cultural exchange between Europe and the Asian continent (the introduction of barley, wheat, livestock and sheep, and copper smelting technology), and the war in the late Neolithic period were important factors affecting the diffusion of the QJC.

64Study on the spatiotemporal distribution patterns and influencing factors of cultural heritage: a case study of Fujian ProvinceOpenAlex

Junjie Fu, Huasong Mao
Abstract The spatiotemporal distribution characteristics of cultural heritage reveal the trajectory of human activity changes, and a deep analysis of its natural and cultural factors holds significant reference value for the overall conservation and management of cultural heritages. This study focuses on the cultural heritage at the provincial level and above in Fujian, utilizing GIS spatial analysis to explore the spatiotemporal evolution of cultural heritages and their natural and human influencing factors. The research findings are as follows: (1) The distribution of cultural heritage in Fujian exhibits a clustering pattern, with dense areas transitioning from the upstream regions of the prehistoric and pre-Qin periods to the eastern coastal areas gradually. (2) The Ming and Qing dynasties have the highest number of cultural heritages, with the type of heritage transitioning from ancient sites in the early periods to ancient architecture, and in modern times, mainly important historical sites and representative architectural heritages. (3) The overall centroid coordinates of cultural heritage reveal a shift from the northern part of Fujian to the eastern and southern parts. (4) Natural factors significantly influence the distribution of cultural heritage, with a higher concentration in plain and hilly areas, on slight slopes with gradients between 0.5° and 2.0°, and on the southern and southeastern slopes, especially within a 1-kilometer radius of rivers. (5) The creation of cultural heritage during historical periods is closely linked to the regional history, culture, political, and economic environments. The positive development of these socio-cultural factors has a promotional effect on the quantity of cultural heritage. This study demonstrates the utility and applicability of GIS spatial analysis techniques in cultural heritage research, providing a methodological framework that can be adapted and applied internationally. The findings offer insightful data that can inform targeted conservation and development strategies for cultural heritage, ensuring their effective preservation and sustainable management across different regions.

65Spatial and Temporal Distribution Characteristics of Heritage Buildings in Yangzhou and Influencing Factors and Tourism Development StrategiesOpenAlex

Kexin Wei, Xuemei Jiang, Rong Zhu, et al.
Heritage buildings are significant humanistic tourism resources for a city. Yangzhou’s heritage buildings have conservation and utilization value and are a key vehicle for promoting urban tourism development. However, there is a lack of research on their spatiotemporal distribution characteristics and subdivision types. This study aims to explore the spatial and temporal clustering and distribution characteristics of Yangzhou’s heritage buildings, as well as the factors contributing to the formation of these distribution patterns, as a means of promoting the tourism development of Yangzhou. Using mathematical statistics and GIS spatial analysis methods, this study analyzes the geographical distribution patterns of 528 heritage buildings and their influencing factors by using average nearest neighbor analysis, an imbalance index, and density mapping. This study reveals the following findings: (1) The temporal distribution shows an “Λ” shape, in which ancient buildings, modern historical sites, and important modern historical sites and representative buildings account for a significant proportion. (2) The temporal center shows a trend of shifting over time, moving from the southwest to the northwest and then to the northeast. (3) The spatial distribution is uneven; most of these are clustered in Hanjiang District, Gaoyou District, and Baoying County, while few are distributed in other regions. (4) The distribution is influenced by both natural and human factors, including topography, water resources, salt merchant culture, revolutionary culture, war culture, and canal transportation culture, with humans and human factors having a more profound impact than natural factors. Based on these findings, strategies such as regional integration and route planning, the prioritization of sustainable tourism development and preservation, and culture fusion and innovative promotion are proposed in this study as references for the all-for-one tourism development and cultural dissemination of Yangzhou.

66Tourism Spatial Structure of Resources-based Attractions in ChinaOpenAlex

Bihu Wu
With the rapid development in economy,the tourism industry of China is also booming.Though there are more and more man-made attractions such as Themed Amusement Park,traditional attractions with natural beauty and/or historical sites are still within the most popular choices of tourists.In this article such kind of attractions are named resources-based attraction.Based on former study,the paper makes definition of resources-based attraction.Then it selects 509 resources-based attractions as research samples from 671 National AAAA Tourist Attractions,which were authorized by China National Tourism Administration in 2005.By means of GIS spatial analysis tools and some quantitative analysis methods such as NNI(Nearest Neighbor Index),GCI(Geographic Concentration Index),Gini Coefficient and Lorenz Curve,the paper analyses the spatial structure of 509 resources-based attractions and observes their distribution in 8 geographical regions and 31 provinces in China.The result shows that the value of NNI is as low as 0.57,which means the distribution of 509 resources-based attractions is a type of agglomeration.And the distribution of 509 resources-based attractions in 8 geographical and 31 provinces is asymmetric.According to the Lorenz Curve,more than half of 509 resources-based attractions concentrate in 9 provinces,such as Jiangsu and Zhejiang.Changjiang River Delta,Beijing,Xi'an and Luoyang are the places with a high density of resources-based attractions.

67Research on Application of GIS in Historical Cultural Resource ProtectionOpenAlex

Chao Tu
Protection of historical cultural resource based on GIS is using GIS to query,search and display detail of historical cultural resource and its establishment.Supervise and plan historical cultural resource using analytic functions of GIS,and offer decision-making function for protection of historical cultural resource based on forecast model.The method can be applied in protection of cultural resource,urban-development planning,travel resource development,sight planning and etc.

68Application and Perspective of GIS in Research on Historical Geography and Cultural GeographyOpenAlex

Fan Li
It is becoming current trend to use GIS for research on historical geography and cultural geography.By integrated analysis of theses and literatures,this paper summarizes progress of using GIS in historical geography,cultural geography,archaeology and cultural resource management,spatial social science,and introduces some main practices of developing historical cultural GIS in Chinese Mainland,Taiwan and Hongkong.The author thinks that though some problems exist in the course of using GIS,it holds enormous potential to apply GIS for research on historical geography and cultural geography.Therefore,scholars of historical geography and cultural geography should recognize function of GIS,develop GIS software for these disciplines together with spatial information technology specialists,build cooperative research platform from different disciplines and areas,and make using GIS for study of history and culture becoming a part of plan of digital city.

69A survey of transfer learningOpenAlex

Karl R. Weiss, Taghi M. Khoshgoftaar, Dingding Wang
Machine learning and data mining techniques have been used in numerous real-world applications. An assumption of traditional machine learning methodologies is the training data and testing data are taken from the same domain, such that the input feature space and data distribution characteristics are the same. However, in some real-world machine learning scenarios, this assumption does not hold. There are cases where training data is expensive or difficult to collect. Therefore, there is a need to create high-performance learners trained with more easily obtained data from different domains. This methodology is referred to as transfer learning. This survey paper formally defines transfer learning, presents information on current solutions, and reviews applications applied to transfer learning. Lastly, there is information listed on software downloads for various transfer learning solutions and a discussion of possible future research work. The transfer learning solutions surveyed are independent of data size and can be applied to big data environments.

70Detecting racial bias in algorithms and machine learningOpenAlex

Nicol Turner Lee
Purpose The online economy has not resolved the issue of racial bias in its applications. While algorithms are procedures that facilitate automated decision-making, or a sequence of unambiguous instructions, bias is a byproduct of these computations, bringing harm to historically disadvantaged populations. This paper argues that algorithmic biases explicitly and implicitly harm racial groups and lead to forms of discrimination. Relying upon sociological and technical research, the paper offers commentary on the need for more workplace diversity within high-tech industries and public policies that can detect or reduce the likelihood of racial bias in algorithmic design and execution. Design/methodology/approach The paper shares examples in the US where algorithmic biases have been reported and the strategies for explaining and addressing them. Findings The findings of the paper suggest that explicit racial bias in algorithms can be mitigated by existing laws, including those governing housing, employment, and the extension of credit. Implicit, or unconscious, biases are harder to redress without more diverse workplaces and public policies that have an approach to bias detection and mitigation. Research limitations/implications The major implication of this research is that further research needs to be done. Increasing the scholarly research in this area will be a major contribution in understanding how emerging technologies are creating disparate and unfair treatment for certain populations. Practical implications The practical implications of the work point to areas within industries and the government that can tackle the question of algorithmic bias, fairness and accountability, especially African-Americans. Social implications The social implications are that emerging technologies are not devoid of societal influences that constantly define positions of power, values, and norms. Originality/value The paper joins a scarcity of existing research, especially in the area that intersects race and algorithmic development.

71Algorithmic BiasOpenAlex

Sara Hajian, Francesco Bonchi, Carlos Castillo
Algorithms and decision making based on Big Data have become pervasive in all aspects of our daily lives lives (offline and online), as they have become essential tools in personal finance, health care, hiring, housing, education, and policies. It is therefore of societal and ethical importance to ask whether these algorithms can be discriminative on grounds such as gender, ethnicity, or health status. It turns out that the answer is positive: for instance, recent studies in the context of online advertising show that ads for high-income jobs are presented to men much more often than to women [Datta et al., 2015]; and ads for arrest records are significantly more likely to show up on searches for distinctively black names [Sweeney, 2013]. This algorithmic bias exists even when there is no discrimination intention in the developer of the algorithm. Sometimes it may be inherent to the data sources used (software making decisions based on data can reflect, or even amplify, the results of historical discrimination), but even when the sensitive attributes have been suppressed from the input, a well trained machine learning algorithm may still discriminate on the basis of such sensitive attributes because of correlations existing in the data. These considerations call for the development of data mining systems which are discrimination-conscious by-design. This is a novel and challenging research area for the data mining community.

72Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AIOpenAlex

Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, et al.

73Fairer machine learning in the real world: Mitigating discrimination without collecting sensitive dataOpenAlex

Michael Veale, Reuben Binns
Decisions based on algorithmic, machine learning models can be unfair, reproducing biases in historical data used to train them. While computational techniques are emerging to address aspects of these concerns through communities such as discrimination-aware data mining (DADM) and fairness, accountability and transparency machine learning (FATML), their practical implementation faces real-world challenges. For legal, institutional or commercial reasons, organisations might not hold the data on sensitive attributes such as gender, ethnicity, sexuality or disability needed to diagnose and mitigate emergent indirect discrimination-by-proxy, such as redlining. Such organisations might also lack the knowledge and capacity to identify and manage fairness issues that are emergent properties of complex sociotechnical systems. This paper presents and discusses three potential approaches to deal with such knowledge and information deficits in the context of fairer machine learning. Trusted third parties could selectively store data necessary for performing discrimination discovery and incorporating fairness constraints into model-building in a privacy-preserving manner. Collaborative online platforms would allow diverse organisations to record, share and access contextual and experiential knowledge to promote fairness in machine learning systems. Finally, unsupervised learning and pedagogically interpretable algorithms might allow fairness hypotheses to be built for further selective testing and exploration. Real-world fairness challenges in machine learning are not abstract, constrained optimisation problems, but are institutionally and contextually grounded. Computational fairness tools are useful, but must be researched and developed in and with the messy contexts that will shape their deployment, rather than just for imagined situations. Not doing so risks real, near-term algorithmic harm.

74Bias in data‐driven artificial intelligence systems—An introductory surveyOpenAlex

Eirini Ntoutsi, Pavlos Fafalios, Ujwal Gadiraju, et al.
Abstract Artificial Intelligence (AI)‐based systems are widely employed nowadays to make decisions that have far‐reaching impact on individuals and society. Their decisions might affect everyone, everywhere, and anytime, entailing concerns about potential human rights issues. Therefore, it is necessary to move beyond traditional AI algorithms optimized for predictive performance and embed ethical and legal principles in their design, training, and deployment to ensure social good while still benefiting from the huge potential of the AI technology. The goal of this survey is to provide a broad multidisciplinary overview of the area of bias in AI systems, focusing on technical challenges and solutions as well as to suggest new research directions towards approaches well‐grounded in a legal frame. In this survey, we focus on data‐driven AI, as a large part of AI is powered nowadays by (big) data and powerful machine learning algorithms. If otherwise not specified, we use the general term bias to describe problems related to the gathering or processing of data that might result in prejudiced decisions on the bases of demographic features such as race, sex, and so forth. This article is categorized under: Commercial, Legal, and Ethical Issues &gt; Fairness in Data Mining Commercial, Legal, and Ethical Issues &gt; Ethical Considerations Commercial, Legal, and Ethical Issues &gt; Legal Issues

75Consumers and Artificial Intelligence: An Experiential PerspectiveOpenAlex

Stefano Puntoni, Rebecca Walker Reczek, Markus Giesler, et al.
Artificial intelligence (AI) helps companies offer important benefits to consumers, such as health monitoring with wearable devices, advice with recommender systems, peace of mind with smart household products, and convenience with voice-activated virtual assistants. However, although AI can be seen as a neutral tool to be evaluated on efficiency and accuracy, this approach does not consider the social and individual challenges that can occur when AI is deployed. This research aims to bridge these two perspectives: on one side, the authors acknowledge the value that embedding AI technology into products and services can provide to consumers. On the other side, the authors build on and integrate sociological and psychological scholarship to examine some of the costs consumers experience in their interactions with AI. In doing so, the authors identify four types of consumer experiences with AI: (1) data capture, (2) classification, (3) delegation, and (4) social. This approach allows the authors to discuss policy and managerial avenues to address the ways in which consumers may fail to experience value in organizations’ investments into AI and to lay out an agenda for future research.

76Redlining Culture: A Data History of Racial Inequality and Postwar FictionOpenAlex

Richard Jean So
The canon of postwar American fiction has changed over the past few decades to include far more writers of color. It would appear that we are making progress-recovering marginalized voices and including those who were for far too long ignored. However, is this celebratory narrative borne out in the data?Richard Jean So draws on big data, literary history, and close readings to offer an unprecedented analysis of racial inequality in American publishing that reveals the persistence of an extreme bias toward white authors. In fact, a defining feature of the publishing industry is its vast whiteness, which has denied nonwhite authors, especially black writers, the coveted resources of publishing, reviews, prizes, and sales, with profound effects on the language, form, and content of the postwar novel. Rather than seeing the postwar period as the era of multiculturalism, So argues that we should understand it as the invention of a new form of racial inequality-one that continues to shape the arts and literature today.Interweaving data analysis of large-scale patterns with a consideration of Toni Morrison's career as an editor at Random House and readings of individual works by Octavia Butler, Henry Dumas, Amy Tan, and others, So develops a form of criticism that brings together qualitative and quantitative approaches to the study of literature. A vital and provocative work for American literary studies, critical race studies, and the digital humanities, Redlining Culture shows the importance of data and computational methods for understanding and challenging racial inequality

77Data FeminismOpenAlex

Catherine D’Ignazio, Lauren Klein
A new way of thinking about data science and data ethics that is informed by the ideas of intersectional feminism. Today, data science is a form of power. It has been used to expose injustice, improve health outcomes, and topple governments. But it has also been used to discriminate, police, and surveil. This potential for good, on the one hand, and harm, on the other, makes it essential to ask: Data science by whom? Data science for whom? Data science with whose interests in mind? The narratives around big data and data science are overwhelmingly white, male, and techno-heroic. In Data Feminism, Catherine D'Ignazio and Lauren Klein present a new way of thinking about data science and data ethics—one that is informed by intersectional feminist thought. Illustrating data feminism in action, D'Ignazio and Klein show how challenges to the male/female binary can help challenge other hierarchical (and empirically wrong) classification systems. They explain how, for example, an understanding of emotion can expand our ideas about effective data visualization, and how the concept of invisible labor can expose the significant human efforts required by our automated systems. And they show why the data never, ever “speak for themselves.” Data Feminism offers strategies for data scientists seeking to learn how feminism can help them work toward justice, and for feminists who want to focus their efforts on the growing field of data science. But Data Feminism is about much more than gender. It is about power, about who has it and who doesn't, and about how those differentials of power can be challenged and changed. The open access edition of this book was made possible by generous funding from the MIT Libraries.

78Fairness and Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, and Mitigation StrategiesOpenAlex

Emilio Ferrara
The significant advancements in applying artificial intelligence (AI) to healthcare decision-making, medical diagnosis, and other domains have simultaneously raised concerns about the fairness and bias of AI systems. This is particularly critical in areas like healthcare, employment, criminal justice, credit scoring, and increasingly, in generative AI models (GenAI) that produce synthetic media. Such systems can lead to unfair outcomes and perpetuate existing inequalities, including generative biases that affect the representation of individuals in synthetic data. This survey study offers a succinct, comprehensive overview of fairness and bias in AI, addressing their sources, impacts, and mitigation strategies. We review sources of bias, such as data, algorithm, and human decision biases—highlighting the emergent issue of generative AI bias, where models may reproduce and amplify societal stereotypes. We assess the societal impact of biased AI systems, focusing on perpetuating inequalities and reinforcing harmful stereotypes, especially as generative AI becomes more prevalent in creating content that influences public perception. We explore various proposed mitigation strategies, discuss the ethical considerations of their implementation, and emphasize the need for interdisciplinary collaboration to ensure effectiveness. Through a systematic literature review spanning multiple academic disciplines, we present definitions of AI bias and its different types, including a detailed look at generative AI bias. We discuss the negative impacts of AI bias on individuals and society and provide an overview of current approaches to mitigate AI bias, including data pre-processing, model selection, and post-processing. We emphasize the unique challenges presented by generative AI models and the importance of strategies specifically tailored to address these. Addressing bias in AI requires a holistic approach involving diverse and representative datasets, enhanced transparency and accountability in AI systems, and the exploration of alternative AI paradigms that prioritize fairness and ethical considerations. This survey contributes to the ongoing discussion on developing fair and unbiased AI systems by providing an overview of the sources, impacts, and mitigation strategies related to AI bias, with a particular focus on the emerging field of generative AI.

79Bias in artificial intelligence algorithms and recommendations for mitigationOpenAlex

Lama Nazer, Razan Zatarah, Shai Waldrip, et al.
The adoption of artificial intelligence (AI) algorithms is rapidly increasing in healthcare. Such algorithms may be shaped by various factors such as social determinants of health that can influence health outcomes. While AI algorithms have been proposed as a tool to expand the reach of quality healthcare to underserved communities and improve health equity, recent literature has raised concerns about the propagation of biases and healthcare disparities through implementation of these algorithms. Thus, it is critical to understand the sources of bias inherent in AI-based algorithms. This review aims to highlight the potential sources of bias within each step of developing AI algorithms in healthcare, starting from framing the problem, data collection, preprocessing, development, and validation, as well as their full implementation. For each of these steps, we also discuss strategies to mitigate the bias and disparities. A checklist was developed with recommendations for reducing bias during the development and implementation stages. It is important for developers and users of AI-based algorithms to keep these important considerations in mind to advance health equity for all populations.

80Guiding Principles to Address the Impact of Algorithm Bias on Racial and Ethnic Disparities in Health and Health CareOpenAlex

Marshall H. Chin, Nasim Afsarmanesh, Arlene S. Bierman, et al.
Importance: Health care algorithms are used for diagnosis, treatment, prognosis, risk stratification, and allocation of resources. Bias in the development and use of algorithms can lead to worse outcomes for racial and ethnic minoritized groups and other historically marginalized populations such as individuals with lower income. Objective: To provide a conceptual framework and guiding principles for mitigating and preventing bias in health care algorithms to promote health and health care equity. Evidence Review: The Agency for Healthcare Research and Quality and the National Institute for Minority Health and Health Disparities convened a diverse panel of experts to review evidence, hear from stakeholders, and receive community feedback. Findings: The panel developed a conceptual framework to apply guiding principles across an algorithm's life cycle, centering health and health care equity for patients and communities as the goal, within the wider context of structural racism and discrimination. Multiple stakeholders can mitigate and prevent bias at each phase of the algorithm life cycle, including problem formulation (phase 1); data selection, assessment, and management (phase 2); algorithm development, training, and validation (phase 3); deployment and integration of algorithms in intended settings (phase 4); and algorithm monitoring, maintenance, updating, or deimplementation (phase 5). Five principles should guide these efforts: (1) promote health and health care equity during all phases of the health care algorithm life cycle; (2) ensure health care algorithms and their use are transparent and explainable; (3) authentically engage patients and communities during all phases of the health care algorithm life cycle and earn trustworthiness; (4) explicitly identify health care algorithmic fairness issues and trade-offs; and (5) establish accountability for equity and fairness in outcomes from health care algorithms. Conclusions and Relevance: Multiple stakeholders must partner to create systems, processes, regulations, incentives, standards, and policies to mitigate and prevent algorithmic bias. Reforms should implement guiding principles that support promotion of health and health care equity in all phases of the algorithm life cycle as well as transparency and explainability, authentic community engagement and ethical partnerships, explicit identification of fairness issues and trade-offs, and accountability for equity and fairness.

81Measuring discrimination in algorithmic decision makingOpenAlex

Indrė Žliobaitė

82Digital craftsmanship in South African sculptural practice and the impact of new technologies on fine art and craftOpenAlex

Jonathan van der Walt
In a modern and fast-evolving technological world, precarity has become more notable. Digital transformation&nbsp; has ushered in an era of ‘datafication’, profoundly impacting societies and individuals in such a way that there are emerging complexities and potential vulnerabilities in our interactions with technology. Thus, it is crucial that the Humanities subjects focus on human beings, their culture and values. This book focuses on the challenges and opportunities experienced in the Digital Humanities. The main thesis of this book is on Digital Humanities in precarious times, while also reporting on topics and research methods in a variety of Humanities subject fields. Digital Humanities is a dynamic multidisciplinary, interdisciplinary and transdisciplinary field that encompasses a wide array of disciplines, methodologies and approaches. It represents a fusion of computational methods with humanistic inquiry, leveraging technology to explore and analyse various facets of human culture, society and history. At its core, this field’s nature allows scholars from diverse backgrounds – including literature, history, linguistics, cultural studies and more – to collaborate and engage in innovative research projects that transcend traditional disciplinary boundaries. All the chapters in this book represent a scholarly discourse and provide original research, they are based on different methodologies ranging from an interdisciplinary approach, a philosophical desk study, case studies, qualitative studies and a semi-structured survey.

83Fluent but Not Factual: A Comparative Analysis of ChatGPT and Other AI Chatbots’ Proficiency and Originality in Scientific Writing for HumanitiesOpenAlex

Edisa Lozić, Benjamin Štular
Historically, mastery of writing was deemed essential to human progress. However, recent advances in generative AI have marked an inflection point in this narrative, including for scientific writing. This article provides a comprehensive analysis of the capabilities and limitations of six AI chatbots in scholarly writing in the humanities and archaeology. The methodology was based on tagging AI-generated content for quantitative accuracy and qualitative precision by human experts. Quantitative accuracy assessed the factual correctness in a manner similar to grading students, while qualitative precision gauged the scientific contribution similar to reviewing a scientific article. In the quantitative test, ChatGPT-4 scored near the passing grade (−5) whereas ChatGPT-3.5 (−18), Bing (−21) and Bard (−31) were not far behind. Claude 2 (−75) and Aria (−80) scored much lower. In the qualitative test, all AI chatbots, but especially ChatGPT-4, demonstrated proficiency in recombining existing knowledge, but all failed to generate original scientific content. As a side note, our results suggest that with ChatGPT-4, the size of large language models has reached a plateau. Furthermore, this paper underscores the intricate and recursive nature of human research. This process of transforming raw data into refined knowledge is computationally irreducible, highlighting the challenges AI chatbots face in emulating human originality in scientific writing. Our results apply to the state of affairs in the third quarter of 2023. In conclusion, while large language models have revolutionised content generation, their ability to produce original scientific contributions in the humanities remains limited. We expect this to change in the near future as current large language model-based AI chatbots evolve into large language model-powered software.

84Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial IntelligenceOpenAlex

Sajid Ali, Tamer Abuhmed, Shaker El–Sappagh, et al.
Artificial intelligence (AI) is currently being utilized in a wide range of sophisticated applications, but the outcomes of many AI models are challenging to comprehend and trust due to their black-box nature. Usually, it is essential to understand the reasoning behind an AI model’s decision-making. Thus, the need for eXplainable AI (XAI) methods for improving trust in AI models has arisen. XAI has become a popular research subject within the AI field in recent years. Existing survey papers have tackled the concepts of XAI, its general terms, and post-hoc explainability methods but there have not been any reviews that have looked at the assessment methods, available tools, XAI datasets, and other related aspects. Therefore, in this comprehensive study, we provide readers with an overview of the current research and trends in this rapidly emerging area with a case study example. The study starts by explaining the background of XAI, common definitions, and summarizing recently proposed techniques in XAI for supervised machine learning. The review divides XAI techniques into four axes using a hierarchical categorization system: (i) data explainability, (ii) model explainability, (iii) post-hoc explainability, and (iv) assessment of explanations. We also introduce available evaluation metrics as well as open-source packages and datasets with future research directions. Then, the significance of explainability in terms of legal demands, user viewpoints, and application orientation is outlined, termed as XAI concerns. This paper advocates for tailoring explanation content to specific user types. An examination of XAI techniques and evaluation was conducted by looking at 410 critical articles, published between January 2016 and October 2022, in reputed journals and using a wide range of research databases as a source of information. The article is aimed at XAI researchers who are interested in making their AI models more trustworthy, as well as towards researchers from other disciplines who are looking for effective XAI methods to complete tasks with confidence while communicating meaning from data.

85AI, Concepts of Intelligence, and Chatbots: The “Figure of Man,” the Rise of Emotion, and Future Visions of EducationOpenAlex

Bernadette Baker, Kathy A. Mills, Peter McDonald, et al.
Background: Artificial intelligence (AI) applications have been implemented across all levels of education, with the rapid developments of chatbots and AI language models, like ChatGPT, demonstrating the urgent need to conceptualize the key debates and their implications for a new era of learning and assessment. This adoption occurs in a context where AI is dramatically remapping “the human,” the purposes of schooling, and pedagogy. Focus of Study: The paper examines how different formulations of “human” became interwoven with the sliding signifier of “intelligence” through a series of violent exclusions, and how the shifting contour of “intelligence” produces uneven and unjust ontological scales undergirding both education and AI fields. Its purpose is to engage the education research community in dialogue about biases, the nature of ethics, and decision-making concerning AI in education. Research Design: This paper adapts a historical-philosophical method. It traces the effects of colonialism and racialization within humanism’s emergence through Sylvia Wynter’s historiography of “figure of Man,” especially via the invention of “intelligence,” which has linked education and computer science. It also investigates themes central to modern education such as justice, equity, and in/exclusion through a philosophical examination of the ontological scales of “human.” Conclusions: After outlining how “intelligence” has shifted from reason-as-morality to concepts of natural intelligence, we argue that current examples of AI in Education (AIEd), like classroom chatbots and social agents, constitute an intermediary point in the arc toward a new computational superintelligence—the emergence of man3—illustrating the opportunities, risks, and ethical issues in pedagogical applications based on emotion. We outline three differing visions of AIEd’s future, concluding with a series of provocations (onto-epistemological, practice-based, and purposes of schooling) that exceed such models and that, given rapid innovations in machine learning, require urgent consideration from multiple stakeholders.

86Artificial Intelligence, Machine Learning, and Cardiovascular DiseaseOpenAlex

Pankaj Mathur, Shweta Srivastava, Xiaowei Xu, et al.
Artificial intelligence (AI)-based applications have found widespread applications in many fields of science, technology, and medicine. The use of enhanced computing power of machines in clinical medicine and diagnostics has been under exploration since the 1960s. More recently, with the advent of advances in computing, algorithms enabling machine learning, especially deep learning networks that mimic the human brain in function, there has been renewed interest to use them in clinical medicine. In cardiovascular medicine, AI-based systems have found new applications in cardiovascular imaging, cardiovascular risk prediction, and newer drug targets. This article aims to describe different AI applications including machine learning and deep learning and their applications in cardiovascular medicine. AI-based applications have enhanced our understanding of different phenotypes of heart failure and congenital heart disease. These applications have led to newer treatment strategies for different types of cardiovascular diseases, newer approach to cardiovascular drug therapy and postmarketing survey of prescription drugs. However, there are several challenges in the clinical use of AI-based applications and interpretation of the results including data privacy, poorly selected/outdated data, selection bias, and unintentional continuance of historical biases/stereotypes in the data which can lead to erroneous conclusions. Still, AI is a transformative technology and has immense potential in health care.