Ines Montani
|
a624ae0675
|
Remove POS, TAG and LEMMA from tokenizer exceptions
|
2020-07-22 23:09:01 +02:00 |
Ines Montani
|
b507f61629
|
Tidy up and move noun_chunks, token_match, url_match
|
2020-07-22 22:18:46 +02:00 |
Ines Montani
|
db55577c45
|
Drop Python 2.7 and 3.5 (#4828)
* Remove unicode declarations
* Remove Python 3.5 and 2.7 from CI
* Don't require pathlib
* Replace compat helpers
* Remove OrderedDict
* Use f-strings
* Set Cython compiler language level
* Fix typo
* Re-add OrderedDict for Table
* Update setup.cfg
* Revert CONTRIBUTING.md
* Revert lookups.md
* Revert top-level.md
* Small adjustments and docs [ci skip]
|
2019-12-22 01:53:56 +01:00 |
Ines Montani
|
6279d74c65
|
Tidy up and auto-format
|
2019-09-11 11:38:22 +02:00 |
Pavle Vidanović
|
d03401f532
|
Lemmatizer lookup dictionary for Serbian and basic tag set adde… (#4251)
* Serbian stopwords added. (cyrillic alphabet)
* spaCy Contribution agreement included.
* Test initialize updated
* Serbian language code update. --bugfix
* Tokenizer exceptions added. Init file updated.
* Norm exceptions and lexical attributes added.
* Examples added.
* Tests added.
* sr_lang examples update.
* Tokenizer exceptions updated. (Serbian)
* Lemmatizer created. Licence included.
* Test updated.
* Tag map basic added.
* tag_map.py file removed since it uses default spacy tags.
|
2019-09-08 14:19:15 +02:00 |
Ines Montani
|
a8752a569d
|
Auto-format [ci skip]
|
2019-08-22 11:44:39 +02:00 |
Pavle Vidanović
|
60e10a9f93
|
Serbian language improvement (#4169)
* Serbian stopwords added. (cyrillic alphabet)
* spaCy Contribution agreement included.
* Test initialize updated
* Serbian language code update. --bugfix
* Tokenizer exceptions added. Init file updated.
* Norm exceptions and lexical attributes added.
* Examples added.
* Tests added.
* sr_lang examples update.
* Tokenizer exceptions updated. (Serbian)
|
2019-08-22 11:43:07 +02:00 |