计算机工程
計算機工程
계산궤공정
COMPUTER ENGINEERING
2010年
5期
173-175
,共3页
分词词典%跳跃表%分词算法%概率算法
分詞詞典%跳躍錶%分詞算法%概率算法
분사사전%도약표%분사산법%개솔산법
segmentation dictionary%leaping form%segmentation algorithm%probabilistic algorithm
结合顺序表和跳跃表的快速查询特性,提出一种改进的整词分词词典结构,主要采用哈希法和二分法进行分词匹配,并针对机械分词算法的特点,引入随机数,探讨一种基于最大匹配的分词概率算法.实验表明,该算法具有较高的分词效率和准确率,对消去歧义词也有较好的性能.
結閤順序錶和跳躍錶的快速查詢特性,提齣一種改進的整詞分詞詞典結構,主要採用哈希法和二分法進行分詞匹配,併針對機械分詞算法的特點,引入隨機數,探討一種基于最大匹配的分詞概率算法.實驗錶明,該算法具有較高的分詞效率和準確率,對消去歧義詞也有較好的性能.
결합순서표화도약표적쾌속사순특성,제출일충개진적정사분사사전결구,주요채용합희법화이분법진행분사필배,병침대궤계분사산법적특점,인입수궤수,탐토일충기우최대필배적분사개솔산법.실험표명,해산법구유교고적분사효솔화준학솔,대소거기의사야유교호적성능.
Combined with the sequence table and leaping form fast inquery characteristic,this paper presents an improvement structure of segmentation dictionary.Hashing and binary search is used to segmentation match for enquiring,and in view ofthe characteristics of the mechanical Chinese word segmentation,by introducing the random number,a Chinese word automatic segmentation probabilistic algorithm is discussed.Experiment indicates that the arithmetic can improve the speed of Chinese segmentation and precision,also,strengthen the processing of dispelling ambiguity.