- 人工智能
- NLP
- 信息抽取
【免费下载链接】gliner2.5-multi-v1
本指南以
fastino/gliner2.5-multi-v1模型卡为核心,系统讲解如何在一次模型中同时完成命名实体识别(NER)、文本分类、关系抽取、Span 属性打分与结构化记录解析。读完本文,你将掌握AutoExtractor的正确加载方式、边界(Boundary)架构的工作机制,以及从独立任务解码到联合信息抽取(JointIE)的完整调用链,能够直接用于多语言场景下的本地推理。
GLiNER2.5 Multi 是 GLiNER2 系列中的多语言边界(boundary)检查点,基于 mDeBERTa-v3-base 构建,是全系列中默认的“一个模型搞定实体、分类、记录与关系”的多语言方案。它通过稀疏的 start/end 配对代替固定宽度网格来定位任意长度的 span,并借助约束解码(Classifier与JointIE)实现跨任务标签约束与带类型端点的实体—关系图搜索。与传统的序列标注式 NER 相比,它的核心创新在于:所有任务共享同一套边界打分器,任务差异只体现在 Schema 的定义上。
GLiNER2.5 模型家族与定位
GLiNER2.5 系列包含三个检查点,三者共享同一套公开 API:
| 模型 | 参数量 | 编码器 | 语言 | 适用场景 |
|---|---|---|---|---|
fastino/gliner2.5-small-v1 | 74M | DeBERTa-v3-xsmall | 英语 | 快速 CPU 抽取 / 分类 |
fastino/gliner2.5-base-v1 | 194M | DeBERTa-v3-base | 英语 | 默认英语多任务检查点 |
fastino/gliner2.5-multi-v1 | 287M | mDeBERTa-v3-base | 多语言 | 默认多语言多任务检查点 |
本仓库承载的正是fastino/gliner2.5-multi-v1的模型文件,包括 config.json(边界提取器完整配置)、model.safetensors(约 594MB 权重,多为 FP16)、tokenizer.json 与 tokenizer_config.json(DebertaV2 分词器,含[SEP_STRUCT]、[SEP_TEXT]、[P]、[C]、[E]、[R]、[L]等任务专用特殊 token)。三个检查点使用同一套公开 API,因此本文所有代码对-small-v1、-base-v1同样适用。
为什么选择 GLiNER2.5
- 一个模型,多种任务:实体、分类、结构化记录、关系与 Span 属性可以写在同一个 Schema里,一次
extract调用全部返回; - 边界架构:用稀疏的 start/end 配对代替固定 span 宽度网格,只要 span 落在编码窗口内,任意长度都能表示;
- 约束解码:
Classifier用于跨任务标签约束,JointIE用于带类型端点的实体—关系图搜索; - 本地推理:通过
gliner2[local]在 CPU、CUDA 或 MPS 上运行,无需外部 API。
边界架构原理:稀疏 start/end 配对
GLiNER2.5 的架构字段在仓库根目录 config.json 中明确标注为"architecture": "boundary"、"architectures": ["BoundaryExtractor"]。与旧版 GLiNER 的span架构(在稠密[L, W]宽度网格上打分)不同,边界架构由两个协同工作的模块组成:
- 候选搜索:先用边界打分器独立选出 top-k 的 start 与 end token(配置中
start_top_k: 24、end_top_k: 24、starts_per_end: 12、ends_per_start: 12),再做稀疏配对,而非枚举所有宽度; - span 打分与内容编码:对配对的候选 span 计算兼容性分数(
pair_dim: 128、multihead_pair_compat_heads: 8),并编码 span 内部内容(enable_span_content: true、content_dim: 64、use_inside_evidence: true),供下游分类、关系、记录任务复用。
边界头(boundary_head配置段)还包含bidirectional_proposals: true(双向候选提议)、enable_abstention: true与abstention_threshold: 0.5(弃权机制,允许模型对不存在的实体输出“无”)以及overlap_policy: "flat"(默认用加权区间调度处理重叠 span)。整个提取器支持max_len: 4096的编码窗口,因此单窗口内可表示任意长度的 span;但注意,它并不会拼接一个首尾从未共现过的 mention。
安装与加载模型
安装
pip install "gliner2[local]"需要 Python 3.10 或更高版本。[local]extra 会一并安装 PyTorch,从而可以直接加载 Hub 上的检查点。
通过 AutoExtractor 加载
加载 GLiNER2.5 必须使用AutoExtractor。GLiNER2.from_pretrained(...)是旧版的span加载器,不会分发这个边界检查点:
from gliner2 import AutoExtractor model = AutoExtractor.from_pretrained("fastino/gliner2.5-multi-v1") print(type(model).__name__) print(model.config.architecture) # BoundaryExtractor # boundary可选设备、fp16 与编译参数:
model = AutoExtractor.from_pretrained( "fastino/gliner2.5-multi-v1", map_location="cuda", # or "cpu" / "mps" quantize=True, # fp16 weights on GPU compile=True, # torch.compile after the first tracing call ) print(type(model).__name__, next(model.parameters()).device) # BoundaryExtractor cuda:0从源码结构看,AutoExtractor会读取检查点config.json中的architecture字段并自动分派到BoundaryExtractor;同样地,Classifier、JointIE等高级组件也都基于同一个边界检查点构建,内部共享实体与关系候选打分结果。因此千万不要用GLiNER2/SpanExtractor加载本检查点——它们期望的是旧版 span 架构。
实体抽取
基础用法
text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = model.extract_entities( text, ["company", "person", "product", "location"], include_confidence=True, include_spans=True, ) print(result) # { # "entities": { # "company": [{"text": "Apple", "start": 0, "end": 5, "confidence": 0.98}], # "person": [{"text": "Tim Cook", "start": 10, "end": 18, "confidence": 0.97}], # "product": [{"text": "iPhone 15", "start": 29, "end": 38, "confidence": 0.96}], # "location": [{"text": "Cupertino", "start": 42, "end": 51, "confidence": 0.95}], # } # }返回的偏移量是半开字符区间,直接作用于原始字符串:text[start:end] == entity["text"],可放心用于下游的高亮、脱敏或文本改写。
为领域标签补充描述
当标签具有领域特异性时,用字典传入标签描述,模型会把这些描述编码为查询(query),显著提升抽取质量:
result = model.extract_entities( "Patient received 400mg ibuprofen for severe headache at 2 PM.", { "medication": "Names of drugs or pharmaceutical substances", "dosage": "Amounts such as 400mg, 2 tablets, or 5ml", "symptom": "Reported symptoms or conditions", "time": "Clock times or relative times", }, include_spans=True, ) print(result) # { # "entities": { # "medication": [{"text": "ibuprofen", "start": 23, "end": 32}], # "dosage": [{"text": "400mg", "start": 17, "end": 22}], # "symptom": [{"text": "severe headache", "start": 37, "end": 52}], # "time": [{"text": "2 PM", "start": 56, "end": 60}], # } # }文本分类
classify_text按任务独立解码。单标签示例:
result = model.classify_text( "This laptop has amazing performance but terrible battery life!", {"sentiment": ["positive", "negative", "neutral"]}, ) print(result) # {"sentiment": "negative"}多标签示例(aspect 分析),通过multi_label与cls_threshold控制输出:
result = model.classify_text( "Great camera quality, decent performance, but poor battery life.", { "aspects": { "labels": ["camera", "performance", "battery", "display", "price"], "multi_label": True, "cls_threshold": 0.4, } }, ) print(result) # {"aspects": ["camera", "performance", "battery"]}约束分类:跨任务标签规则
当一个任务的标签在逻辑上约束另一个任务时(例如“intent=delete 必然带来 effects 含 delete”),classify_text的独立解码不会强制这些规则,需要改用gliner2.classification.Classifier:
from gliner2.classification import ( Classifier, ClassificationSchema, ClassificationConfig, ) from gliner2.classification import constraints as C clf = Classifier.from_pretrained("fastino/gliner2.5-multi-v1") schema = ( ClassificationSchema() .single("intent", ["read", "write", "delete"]) .multi("effects", ["read_only", "create", "modify", "delete"], min_labels=1) .constrain( C.implies(("intent", "delete"), ("effects", "delete")), C.excludes(("intent", "read"), ("effects", "delete")), ) ) result = clf.classify("Delete the temporary file from /tmp", schema) print(result.value("intent")) print(result.value("effects")) print(result.feasible) print(result.to_dict()) # delete # ['delete'] # True # { # "intent": { # "value": "delete", # "confidence": 0.93, # "probabilities": {"read": 0.02, "write": 0.05, "delete": 0.93}, # }, # "effects": { # "value": ["delete"], # "confidence": 0.88, # "probabilities": { # "read_only": 0.04, "create": 0.03, "modify": 0.05, "delete": 0.88 # }, # }, # "_meta": {"feasible": True, "decoder": "exact"}, # }这里的result.to_dict()会同时返回各标签的完整概率分布与_meta元信息;result.feasible表示硬约束是否被满足。
预测相关的参数(解码器、beam 大小等)属于调用时的ClassificationConfig,而不是from_pretrained:
result = clf.classify( "Preview the report", schema, config=ClassificationConfig(decoder="beam", beam_size=16), ) print(result.value("intent"), result.value("effects"), result.feasible) # read ['read_only'] True关系抽取
本检查点以enable_relations=True训练(见 config.json 中boundary_head段),支持独立解码:
text = "Alice works for Acme in Paris." result = model.extract_relations( text, ["works_for", "located_in"], include_spans=True, include_confidence=True, ) print(result) # { # "relation_extraction": { # "works_for": [{ # "head": {"text": "Alice", "start": 0, "end": 5, "confidence": 0.91}, # "tail": {"text": "Acme", "start": 16, "end": 20, "confidence": 0.91}, # }], # "located_in": [{ # "head": {"text": "Acme", "start": 16, "end": 20, "confidence": 0.87}, # "tail": {"text": "Paris", "start": 24, "end": 29, "confidence": 0.87}, # }], # } # }也可以通过 Schema 指定每个关系的阈值:
schema = model.create_schema().relations( {"works_for": {"threshold": 0.6}, "located_in": {"threshold": 0.6}} ) result = model.extract(text, schema, include_spans=True) print(result) # { # "relation_extraction": { # "works_for": [{ # "head": {"text": "Alice", "start": 0, "end": 5}, # "tail": {"text": "Acme", "start": 16, "end": 20}, # }], # "located_in": [{ # "head": {"text": "Acme", "start": 16, "end": 20}, # "tail": {"text": "Paris", "start": 24, "end": 29}, # }], # } # }需要特别提醒:独立抽取不保证类型约束——works_for的 head 不一定是人、tail 不一定是组织。若需要这类硬性保证,请使用下一节的JointIE。
联合信息抽取(JointIE)
JointIE对 mention 与关系候选统一打分后,在带类型端点和唯一性约束的全局图上搜索一致解:
from gliner2.joint_ie import JointIE, JointIEConfig joint = JointIE.from_pretrained("fastino/gliner2.5-multi-v1") schema = ( joint.create_schema() .entities(["person", "organization", "location"]) .relation("works_for", "person", "organization", unique_head=True) .relation("located_in", "organization", "location") .no_self_loops() ) result = joint.extract( "Alice works for Acme in Paris. Bob joined Acme last year.", schema, config=JointIEConfig(optimizer="beam", beam_size=32), ) print(result.feasible) print(result.to_dict()) # True # { # "entities": [ # {"id": "e1", "type": "person", "text": "Alice", "start": 0, "end": 5, "confidence": 0.94}, # {"id": "e2", "type": "organization", "text": "Acme", "start": 16, "end": 20, "confidence": 0.92}, # {"id": "e3", "type": "location", "text": "Paris", "start": 24, "end": 29, "confidence": 0.90}, # {"id": "e4", "type": "person", "text": "Bob", "start": 31, "end": 34, "confidence": 0.91}, # ], # "relations": [ # {"type": "works_for", "head": "e1", "tail": "e2", "confidence": 0.88}, # {"type": "works_for", "head": "e4", "tail": "e2", "confidence": 0.81}, # {"type": "located_in", "head": "e2", "tail": "e3", "confidence": 0.86}, # ], # }注意unique_head=True使每个 head(人)最多拥有一个works_for指向,而Acme作为 tail 可以被多人同时指向;no_self_loops()禁止实体指向自身的自环关系。
务必检查result.feasible:False表示硬约束无法被满足,这与“文本中不包含事实”是两回事:
for rel in result.relations: head = result.entity(rel.head) tail = result.entity(rel.tail) print(f"{head.text} -{rel.type}-> {tail.text}") # Alice -works_for-> Acme # Bob -works_for-> Acme # Acme -located_in-> ParisSpan 属性:给实体打上情感标签
Span 属性是span 条件化的:模型先找到实体,再在这些精确 span 上为属性打分。它不是额外的实体类型,也不是文档级分类。以人物情感为例:
from gliner2 import AutoExtractor, AttributeGroup model = AutoExtractor.from_pretrained("fastino/gliner2.5-multi-v1") text = ( "Alice was delighted with the promotion, " "but Bob sounded frustrated about the delay." ) schema = ( model.create_schema() .entities(["person"]) .entity_attributes({ "sentiment": AttributeGroup( ["positive", "negative", "neutral"], applies_to=["person"], qualify_labels=True, ) }) ) result = model.extract( text, schema, include_spans=True, include_confidence=True, ) print(result) # { # "entities": { # "person": [ # { # "text": "Alice", # "start": 0, # "end": 5, # "confidence": 0.96, # "sentiment": {"label": "positive", "confidence": 0.89}, # }, # { # "text": "Bob", # "start": 44, # "end": 47, # "confidence": 0.95, # "sentiment": {"label": "negative", "confidence": 0.84}, # }, # ] # } # }参数说明:applies_to=["person"]让情感属性只挂在 person 上,不影响其他实体类型;qualify_labels=True让模型侧查询编码为sentiment: positive,而返回给用户的仍是短标签positive。
把情感限制在人物上、同时照常抽取组织:
schema = ( model.create_schema() .entities(["person", "organization"]) .entity_attributes({ "sentiment": AttributeGroup( ["positive", "negative", "neutral"], applies_to=["person"], qualify_labels=True, ) }) ) result = model.extract( "Alice praised Microsoft, but Bob criticized OpenAI.", schema, include_spans=True, include_confidence=True, ) print(result) # { # "entities": { # "person": [ # { # "text": "Alice", # "start": 0, # "end": 5, # "confidence": 0.96, # "sentiment": {"label": "positive", "confidence": 0.88}, # }, # { # "text": "Bob", # "start": 29, # "end": 32, # "confidence": 0.95, # "sentiment": {"label": "negative", "confidence": 0.86}, # }, # ], # "organization": [ # {"text": "Microsoft", "start": 14, "end": 23, "confidence": 0.97}, # {"text": "OpenAI", "start": 44, "end": 50, "confidence": 0.96}, # ], # } # }注意输出差异:组织 span 没有sentiment字段,person span 才有——这正是applies_to约束生效的体现。
结构化记录:保留实例身份
普通字段抽取会把所有字段拍平进互不关联的列表,丢失“谁买了什么”的实例对应关系。记录模式通过anchor 字段绑定实例身份,本检查点以enable_records=True训练。使用natural模式并指定 anchor:
schema = ( model.create_schema() .structure("purchase", mode="natural", anchor="buyer") .field("buyer", dtype="str", cardinality="required_one") .field("item", dtype="str", cardinality="required_one") ) result = model.extract( "Alice bought apples and Bob bought oranges.", schema, ) print(result) # { # "purchase": [ # {"buyer": "Alice", "item": "apples"}, # {"buyer": "Bob", "item": "oranges"}, # ] # }cardinality支持required_one(每个实例必须恰好一个)等取值;anchor 字段用于把同名不同实例的字段值正确分组到对应记录里。
任务组合:一次 extract 完成全部任务
实体、Span 属性、分类、关系与结构化记录可以写进同一个 Schema,在一次extract调用中全部返回:
from gliner2 import AttributeGroup schema = ( model.create_schema() .entities({ "person": "Named people", "organization": "Companies or teams", "product": "Named products or services", }) .entity_attributes({ "sentiment": AttributeGroup( ["positive", "negative", "neutral"], applies_to=["person"], qualify_labels=True, ) }) .classification("topic", ["technology", "business", "sports", "politics"]) .relations(["works_for", "announced"]) .structure("announcement", mode="natural", anchor="product") .field("company", dtype="str") .field("product", dtype="str", cardinality="required_one") ) text = "Apple CEO Tim Cook unveiled the iPhone 15 Pro for $999." result = model.extract(text, schema, include_spans=True, include_confidence=True) print(result) # { # "entities": { # "person": [{ # "text": "Tim Cook", # "start": 10, # "end": 18, # "confidence": 0.97, # "sentiment": {"label": "positive", "confidence": 0.82}, # }], # "organization": [{"text": "Apple", "start": 0, "end": 5, "confidence": 0.98}], # "product": [{"text": "iPhone 15 Pro", "start": 32, "end": 45, "confidence": 0.96}], # }, # "topic": {"label": "technology", "confidence": 0.94}, # "relation_extraction": { # "works_for": [{ # "head": {"text": "Tim Cook", "start": 10, "end": 18, "confidence": 0.86}, # "tail": {"text": "Apple", "start": 0, "end": 5, "confidence": 0.86}, # }], # "announced": [{ # "head": {"text": "Tim Cook", "start": 10, "end": 18, "confidence": 0.84}, # "tail": {"text": "iPhone 15 Pro", "start": 32, "end": 45, "confidence": 0.84}, # }], # }, # "announcement": [{ # "company": "Apple", # "product": "iPhone 15 Pro", # }], # }注意语义边界:文档级topic与逐人sentiment是相互独立的;sentiment只挂在本条中Tim Cook的 person 实体上。
批量推理
texts = [ "Google hired Jane Doe in London.", "Tesla launched the Model 3 in California.", ] results = model.batch_extract_entities( texts, ["company", "person", "product", "location"], batch_size=8, include_spans=True, ) print(results) # [ # { # "entities": { # "company": [{"text": "Google", "start": 0, "end": 6}], # "person": [{"text": "Jane Doe", "start": 13, "end": 21}], # "product": [], # "location": [{"text": "London", "start": 25, "end": 31}], # } # }, # { # "entities": { # "company": [{"text": "Tesla", "start": 0, "end": 5}], # "person": [], # "product": [{"text": "Model 3", "start": 19, "end": 26}], # "location": [{"text": "California", "start": 30, "end": 40}], # } # }, # ]batch_extract既可以接收单个 Schema(作用于所有文档),也可以接收一个 Schema 列表(每篇文档对应一个 Schema)。
长文档处理
extract(...)传入max_len会截断文本。长上下文助手则通过扫描重叠的词块并在块内映射回文档偏移来解决长文问题:
long_text = ("Quarterly overview. " * 40) + "Satya Nadella spoke in Redmond about Microsoft." result = model.extract_entities_long( long_text, ["person", "organization", "location"], chunk_size=384, chunk_overlap=64, include_spans=True, ) print(result) # { # "entities": { # "person": [{"text": "Satya Nadella", "start": 800, "end": 813}], # "organization": [{"text": "Microsoft", "start": 837, "end": 846}], # "location": [{"text": "Redmond", "start": 823, "end": 830}], # } # } result = model.extract_long(long_text, schema, chunk_size=384, chunk_overlap=64) print(result["topic"]) # technology同样的思路也适用于Classifier.classify_long与JointIE.extract_long。
长文档处理的限制:
- 只有 start 和 end落在同一块中的 span 才会被保留;
- 只有当关系的两个端点在同一块中都被抽取时,关系才会被保留;
- 边界模型可以在一个编码窗口内表示任意长的 span,但它不会拼接首尾从未在同一块中共同出现的 mention。
模型细节
以下信息来自本仓库的 config.json 与 encoder_config/config.json(编码器 mDeBERTa-v3-base 的原始配置),可直接核对:
- 架构:GLiNER2boundary提取器(
BoundaryExtractor) - 候选搜索:稀疏 start/end 配对(非稠密
[L, W]宽度网格) - Span 长度:任意不超过编码窗口的长度(
max_len=4096) - 编码器:
microsoft/mdeberta-v3-base(12 层、hidden size 768、DebertaV2 结构) - 参数量:287M
- 权重体积:约 594 MB(大部分为 FP16)
- 语言:多语言
- 启用的任务头:分类、记录(
enable_records=True)、关系(enable_relations=True) - 重叠默认策略:
flat(加权区间调度);可在每次调用时用overlap_policy覆盖 - 输入 / 输出:文本 → 实体、标签、Span 属性、记录与关系边
边界头的关键超参数(来自boundary_head段)还包括:candidate_budget: 192与training_candidate_budget: 192(候选预算)、pool_size: 192与pool_boundary_top_k: 32(候选池)、ends_per_start: 12与starts_per_end: 12(start/end 双向配对比例)、enable_abstention: true与abstention_threshold: 0.5(弃权机制)、record_anchor_threshold: 0.5与record_field_threshold: 0.5(记录 anchor 与字段阈值)、relation_argument_proposal_threshold: 0.2(关系参数提议阈值)等。这些参数直接决定了推理时每个 query 的候选规模与阈值行为,是理解模型运行开销的关键入口。
再次强调:不要用GLiNER2/SpanExtractor加载本检查点,这两个类面向旧版 span 架构,无法正确分发与解码。
引用与许可
如果在研究或产品中使用了本模型,请引用:
@misc{zaratiana2025gliner2efficientmultitaskinformation, title={GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface}, author={Urchade Zaratiana and Gil Pasternak and Oliver Boyd and George Hurn-Maloney and Ash Lewis}, year={2025}, eprint={2507.18546}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.18546}, }本模型以 Apache License 2.0 发布。
- 人工智能
- NLP
- 信息抽取
【免费下载链接】gliner2.5-multi-v1
相关推荐
GLiNER2.5-Multi多语言信息抽取与微调部署:跨语言场景落地完整指南
GLiNER2.5 Multi多语言信息抽取与微调部署:跨语言场景落地完整指南 GLiNER2.5 Multi( fastino/gliner2.5 multi
人工智能NLP信息抽取如何从自然语言中抽取结构化JSON记录:GLiNER2.5-Multi结构记录抽取指南
如何从自然语言中抽取结构化JSON记录:GLiNER2.5 Multi结构记录抽取指南 GLiNER2.5 Multi 是多语言信息抽取模型,一条自然语言即可提
人工智能NLP信息抽取GLiNER2.5-Multi JointIE联合信息抽取全解:如何构建类型化实体-关系图
GLiNER2.5 Multi JointIE联合信息抽取全解:如何构建类型化实体 关系图 GLiNER2.5 Multi 是一个多语言统一信息抽取模型,其 J
人工智能NLP信息抽取
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考