电子与信息学报
電子與信息學報
전자여신식학보
JOURNAL OF ELECTRONICS & INFORMATION TECHNOLOGY
2010年
3期
700-704
,共5页
中文分词%词性标注%一体化系统%无向图模型
中文分詞%詞性標註%一體化繫統%無嚮圖模型
중문분사%사성표주%일체화계통%무향도모형
Chinese words segmentation%Part-Of-Speech (POS) tagging%Joint system%Undirected graphical model
在中文词法分析中,分词是词性标注必须经历的阶段.为了能在分词阶段就充分利用词性标注的信息和减少两阶段错误的累计,最好的方法是将两个阶段,整合到一个架构中.该文以无向图模型为基础,将分词和词性标注有机地统一在一个序列标注模型中.由于可以采用更深层次的依赖关系作为特征,一体化系统在1998年人民日报语料上取得了97.19%的分词精确率和95.34%的词性标注精确率,是目前同类系统,在这一语料上取得的最好结果.
在中文詞法分析中,分詞是詞性標註必鬚經歷的階段.為瞭能在分詞階段就充分利用詞性標註的信息和減少兩階段錯誤的纍計,最好的方法是將兩箇階段,整閤到一箇架構中.該文以無嚮圖模型為基礎,將分詞和詞性標註有機地統一在一箇序列標註模型中.由于可以採用更深層次的依賴關繫作為特徵,一體化繫統在1998年人民日報語料上取得瞭97.19%的分詞精確率和95.34%的詞性標註精確率,是目前同類繫統,在這一語料上取得的最好結果.
재중문사법분석중,분사시사성표주필수경력적계단.위료능재분사계단취충분이용사성표주적신식화감소량계단착오적루계,최호적방법시장량개계단,정합도일개가구중.해문이무향도모형위기출,장분사화사성표주유궤지통일재일개서렬표주모형중.유우가이채용경심층차적의뢰관계작위특정,일체화계통재1998년인민일보어료상취득료97.19%적분사정학솔화95.34%적사성표주정학솔,시목전동류계통,재저일어료상취득적최호결과.
For Chinese Part-Of-Speech(POS) tagging, word segmentation is a preliminary step. To reduce accumulated errors between two steps and improve the segmentation performance by utilizing POS information, segmentation and POS tagging can be performed simultaneously. In this paper, a joint segmentation and POS tagging system is proposed based on undirected graphical models which can make full use of the dependencies between the two stages. In the joint system, segmenting and tagging are viewed as the sequence labeling; moreover any connected sub-graph can be viewed as a certain dependency which can be used to find the final opinion labeling. The joint model achieves high performances with 97.19% in segmentation precision and 95.34% in POS tagging precision, which are the state-of-art performances for Chinese word segmentation and tagging on 1998-year People's Daily corpus.