Python找出文件中使用率最高的漢字實例詳解

2020-02-23 01:30:25

字體：大中小

來源：轉載

供稿：網友

本文實例講述了Python找出文件中使用率最高的漢字的方法。分享給大家供大家參考。具體分析如下：

這是我初學Python時寫的，為了簡便，我并沒在排序完后再去掉非中文字符，稍微會影響性能（大約增加了25％的時間）。

# -*- coding: gbk -*- import codecs from time import time from operator import itemgetter def top_words(filename, size=10, encoding='gbk'):   count = {}   for line in codecs.open(filename, 'r', encoding):     for word in line:       if u'/u4E00' <= word <= u'/u9FA5' or u'/uF900' <= word <= u'/uFA2D':         count[word] = 1 + count.get(word, 0)   top_words = sorted(count.iteritems(), key=itemgetter(1), reverse=True)[:size]   print '/n'.join([u'%s : %s次' % (word, times) for word, times in top_words]) begin = time() top_words('空之境界.txt') print '一共耗時 : %s秒' % (time()-begin)

如果想用上新方法，以及讓join的可讀性更高的話，這樣也是可以的：

# -*- coding: gbk -*- import codecs from time import time from operator import itemgetter from heapq import nlargest def top_words(filename, size=10, encoding='gbk'):   count = {}   for line in codecs.open(filename, 'r', encoding):     for word in line:       if u'/u4E00' <= word <= u'/u9FA5' or u'/uF900' <= word <= u'/uFA2D':         count[word] = 1 + count.get(word, 0)   top_words = nlargest(size, count.iteritems(), key=itemgetter(1))   for word, times in top_words:     print u'%s : %s次' % (word, times) begin = time() top_words('空之境界.txt') print '一共耗時 : %s秒' % (time()-begin)

或者讓行數更少（好囧的列表綜合）：

# -*- coding: gbk -*- import codecs from time import time from operator import itemgetter def top_words(filename, size=10, encoding='gbk'):   count = {}   for word in [word for word in codecs.open(filename, 'r', encoding).read() if u'/u4E00' <= word <= u'/u9FA5' or u'/uF900' <= word <= u'/uFA2D']:     count[word] = 1 + count.get(word, 0)   top_words = sorted(count.iteritems(), key=itemgetter(1), reverse=True)[:size]   print '/n'.join([u'%s : %s次' % (word, times) for word, times in top_words]) begin = time() top_words('空之境界.txt') print '一共耗時 : %s秒' % (time()-begin)

此外還可以引入with語句，這樣只需一行就能獲得異常安全性。
3者性能幾乎一樣，結果如下：

的 : 17533次是 : 8581次不 : 6375次我 : 6168次了 : 5586次一 : 5197次這 : 4394次在 : 4264次有 : 4188次人 : 4025次一共耗時 : 0.5秒

引入psyco模塊的成績：

的 : 17533次是 : 8581次不 : 6375次我 : 6168次了 : 5586次一 : 5197次這 : 4394次在 : 4264次有 : 4188次人 : 4025次一共耗時 : 0.280999898911秒

注：測試文件為778KB的GBK編碼，40余萬字。

上一篇：對于Python裝飾器使用的一些建議

下一篇：python使用xmlrpclib模塊實現對百度google的ping功能

學習交流

筆記本開機提示error loading os錯誤的問

筆記本開機提示error loading os錯誤的問題怎么解決...

熱門圖片

猜你喜歡的新聞

猜你喜歡的關注

国产探花免费观看_亚洲丰满少妇自慰呻吟_97日韩有码在线_资源在线日韩欧美_一区二区精品毛片,辰东完美世界有声小说,欢乐颂第一季,yy玄幻小说排行榜完本

Python找出文件中使用率最高的漢字實例詳解