Matthew Honnibal
|
711ed0f636
|
* Whitespace
|
2014-11-02 14:22:32 +11:00 |
Matthew Honnibal
|
3352e89e21
|
* Use LIKE_URL and LIKE_NUMBER flag features. Seems to improve accuracy on onto web
|
2014-11-02 13:19:54 +11:00 |
Matthew Honnibal
|
f67cb9a5a3
|
* Add count_tags functionto pos.pyx, which should probably live in another file. Feature set achieves 97.9 on wsj19-21, 95.85 on onto web.
|
2014-10-31 17:42:04 +11:00 |
Matthew Honnibal
|
e6b87766fe
|
* Remove lexemes vector from Lexicon, and the id and hash attributes from Lexeme
|
2014-10-30 15:21:38 +11:00 |
Matthew Honnibal
|
889b7b48b4
|
* Fix POS tagger, so that it loads correctly. Lexemes are being read in.
|
2014-10-30 13:38:55 +11:00 |
Matthew Honnibal
|
13909a2e24
|
* Rewriting Lexeme serialization.
|
2014-10-29 23:19:38 +11:00 |
Matthew Honnibal
|
08ce602243
|
* Large refactor, particularly to Python API
|
2014-10-24 00:59:17 +11:00 |
Matthew Honnibal
|
96b835a3d4
|
* Upd for refactored Tokens class. Now gets 95.74, 185ms training on swbd_wsj_ewtb, eval on onto_web, Google POS tags.
|
2014-10-23 03:20:02 +11:00 |
Matthew Honnibal
|
ea1d4a81eb
|
* Refactoring get_atoms, improving tokens API
|
2014-10-22 13:10:56 +11:00 |
Matthew Honnibal
|
ad49e2482e
|
* Tagger now gets 97pc on wsj, parsing 19-21 in 500ms. Gets 92.7 on web text.
|
2014-10-22 12:57:06 +11:00 |
Matthew Honnibal
|
5ebe14f353
|
* Add greedy pos tagger
|
2014-10-22 10:17:26 +11:00 |