Kevin Humphreys
|
597df5bf83
|
add test
|
2018-01-03 13:00:05 -08:00 |
Ines Montani
|
ff9fc945ab
|
Merge pull request #1749 from sorenlind/da_ud_tokenization
Tune Danish tokenizer to more closely match Universal Dependencies
|
2017-12-22 16:00:49 +00:00 |
ines
|
26f313dabc
|
Fix missing import
|
2017-12-22 16:21:44 +01:00 |
ines
|
8dc1c27841
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-12-22 16:01:00 +01:00 |
ines
|
b10ba848b8
|
xfail test that causes MemoryError on Python 2 on Windows
Need to investigate this further!
|
2017-12-22 16:00:58 +01:00 |
Ines Montani
|
a3dd167d7f
|
Merge branch 'master' into da_ud_tokenization
|
2017-12-20 21:05:34 +00:00 |
Ines Montani
|
d682a8803e
|
Merge pull request #1672 from cbilgili/master
Adds Turkish Lemmatization
|
2017-12-20 21:01:00 +00:00 |
Søren Lind Kristiansen
|
15d13efafd
|
Tune Danish tokenizer to more closely match tokenization in Universal Dependencies.
|
2017-12-20 17:36:52 +01:00 |
Ines Montani
|
9c1ee65268
|
Add regression test for #1698
|
2017-12-12 10:36:11 +01:00 |
Isaac Sijaranamual
|
38021fbb00
|
Switch from python 3 only TemporaryDirectory to pytest's tmpdir
|
2017-12-11 00:16:04 +01:00 |
Isaac Sijaranamual
|
568130ce7c
|
Adds regression test_issue1622
|
2017-12-10 23:00:48 +01:00 |
Matthew Honnibal
|
36b47e3fa6
|
Fix (and test) vector pickling
|
2017-12-07 09:53:30 +01:00 |
Canbey Bilgili
|
abe098b255
|
Adds Turkish Lemmatization
|
2017-12-01 17:04:32 +03:00 |
Vadim Mazaev
|
4ba7ddf651
|
Bugfixies
|
2017-11-30 12:29:38 +03:00 |
Matthew Honnibal
|
6bc0f4d29f
|
Merge pull request #1611 from fsonntag/master
Solving #1494
|
2017-11-29 23:11:23 +01:00 |
Matthew Honnibal
|
f9ed9ea529
|
Merge pull request #1624 from GreenRiverRUS/russian
Add support for Russian
|
2017-11-29 23:10:01 +01:00 |
ines
|
a31506e060
|
Fix off-by-one error in nlp.add_pipe(after=name) (fixes #1654)
|
2017-11-28 20:37:55 +01:00 |
ines
|
b62739fbfe
|
Add regression test for #1654
|
2017-11-28 20:27:54 +01:00 |
ines
|
2e50dbb9d7
|
Simplify test
|
2017-11-28 20:27:27 +01:00 |
Felix Sonntag
|
724ae7dc55
|
Fixed issue of infix capturing prefixes
|
2017-11-28 17:17:12 +01:00 |
Søren Lind Kristiansen
|
0ffd27b0f6
|
Add several Danish alternative spellings
|
2017-11-27 13:35:41 +01:00 |
Vadim Mazaev
|
53e7c38637
|
Fixed tests depends on pymorphy2
|
2017-11-26 21:04:44 +03:00 |
Vadim Mazaev
|
cacd859dcd
|
Added tag map, fixed tests fails, added more exceptions
|
2017-11-26 20:54:48 +03:00 |
Ines Montani
|
a7bb8f1b42
|
Merge pull request #1637 from sorenlind/da_tokenization
Improve Danish tokenization
|
2017-11-26 15:41:38 +00:00 |
ines
|
c699aec089
|
Add offsets_from_biluo_tags helper and tests (see #1626)
|
2017-11-26 16:38:01 +01:00 |
Søren Lind Kristiansen
|
6aa241bcec
|
Add day of month tokenizer exceptions for Danish.
|
2017-11-24 15:03:24 +01:00 |
Søren Lind Kristiansen
|
0c276ed020
|
Add weekday abbreviations and remove abiguous month abbreviations for Danish.
|
2017-11-24 14:43:29 +01:00 |
Søren Lind Kristiansen
|
056547e989
|
Add multiple tokenizer exceptions for Danish.
|
2017-11-24 11:51:26 +01:00 |
Søren Lind Kristiansen
|
8dc265ac0c
|
Add test for tokenization of 'i.' for Danish.
|
2017-11-24 11:29:37 +01:00 |
Matthew Honnibal
|
30ba81f881
|
Merge pull request #1576 from ligser/master
Actually reset caches in pipe [wip]
|
2017-11-23 12:54:48 +01:00 |
ines
|
c90fe92e15
|
Fix displaCy test
|
2017-11-22 05:04:39 +01:00 |
ines
|
a6f33ac27d
|
Fix displaCy test
|
2017-11-22 04:19:28 +01:00 |
Vadim Mazaev
|
81314f8659
|
Fixed tokenizer: added char classes; added first lemmatizer and
tokenizer tests
|
2017-11-21 22:23:59 +03:00 |
Burton DeWilde
|
635792997c
|
Add regression test for #1612
|
2017-11-20 12:05:35 -06:00 |
ines
|
d70a64d78b
|
Fix syntax error and formatting in test (see #1617)
|
2017-11-20 14:01:25 +01:00 |
ines
|
17849dee4b
|
Fix French test (see #1617)
|
2017-11-20 13:59:59 +01:00 |
Felix Sonntag
|
8be3392302
|
Added regression text for 1494
|
2017-11-19 16:30:35 +01:00 |
Motoki Wu
|
b818afaa0e
|
Added failing test for Issue #1207.
The noun chunk iterator should work for `Doc` but not for `Span`.
|
2017-11-17 17:04:27 -08:00 |
ines
|
a3d4dd1a5d
|
Test adding of lots of pipeline components (see #1585)
Just to make sure that there's no error now or in the future with adding a large number of pipeline components.
|
2017-11-15 17:28:06 +01:00 |
Roman Domrachev
|
505c6a2f2f
|
Completely cleanup tokenizer cache
Tokenizer cache can have be different keys than string
That modification can slow down tokenizer and need to be measured
|
2017-11-15 17:55:48 +03:00 |
Roman Domrachev
|
3e21680814
|
Use safer method to get string without hit
|
2017-11-14 22:58:46 +03:00 |
Roman Domrachev
|
4e378dc4a4
|
Remove all obsolete code and test only initial problem
|
2017-11-14 20:45:04 +03:00 |
Roman
|
47ce2347b0
|
Create test that fails when actual cleanup caused
|
2017-11-14 20:28:13 +03:00 |
Roman Domrachev
|
3d247d2bb8
|
Get back previous testcase
|
2017-11-14 18:01:37 +03:00 |
Roman Domrachev
|
a2745b0e84
|
StringStore now actually cleaned
Do not lose docs in ref tracking
|
2017-11-14 17:45:50 +03:00 |
Roman Domrachev
|
ee60a52ee7
|
Fix test imports and last batch cleanup
|
2017-11-11 11:32:16 +03:00 |
Roman Domrachev
|
3c600adf23
|
Try to fix StringStore clean up (see #1506)
|
2017-11-11 03:11:27 +03:00 |
ines
|
ee97fd3cb4
|
Add regression test for #1547
|
2017-11-11 00:14:03 +01:00 |
ines
|
2df27db671
|
Add unicode declaration
|
2017-11-11 00:13:56 +01:00 |
ines
|
1c218397f6
|
Ensure path in Doc.to_disk/from_disk (resolves ##1521)
Also add Doc serialization tests with both Path and string path options
|
2017-11-09 02:29:03 +01:00 |