计算机应用 ›› 2005, Vol. 25 ›› Issue (09): 2025-2027.DOI: 10.3724/SP.J.1087.2005.02025

• 人工智能 • 上一篇    下一篇

一种基于提取上下文信息的分词算法

曾华琳,李堂秋,史晓东   

  1. 厦门大学计算机科学系
  • 发布日期:2011-04-11 出版日期:2005-09-01
  • 基金资助:

    国家863计划资助项目(2002AA117010)

Segmentation algorithm for Chinese based on extraction of context information

 ZENG Hua-lin,LI Tang-qiu,SHI Xiao-dong   

  1. Department of Computer Science,Xiamen University,Xiamen Fujian 361005,China
  • Online:2011-04-11 Published:2005-09-01

摘要: 汉语分词在汉语文本处理过程中是一个特殊而重要的组成部分。传统的基于词典的分词算法存在很大的缺陷,无法对未登录词进行很好的处理。基于概率的算法只考虑了训练集语料的概率模型,对于不同领域的文本的处理不尽如人意。文章提出一种基于上下文信息提取的概率分词算法,能够将切分文本的上下文信息加入到分词概率模型中,以指导文本的切分。这种切分算法结合经典n元模型以及EM算法,在封闭和开放测试环境中分别取得了比较好的效果。

关键词: 中文分词, n元模型, 上下文信息

Abstract: Chinese segmentation is a special and important issue in Chinese texts processing.The traditional segmentation methods based on an existing dictionary have an obvious defect when they are used to segment texts which may contain words unknown to the dictionary.And the probabilistic methods those consider the probabilistic model of the training set only also do a bad job on the texts of a specific domain.In this paper,a probabilistic segmentation method based on extracting context information was proposed,which adds the context information of the segmenting text into the segmentation probabilistic model so as to guide the processing.The method combining n-gram model and EM algorithm achieves a good effect in the close and opening test.

Key words: Chinese segmentation, n-gram model, context information

中图分类号: