spaCy

History

adrianeboyd 82159b5c19 Updates/bugfixes for NER/IOB converters (#4186 ) * Updates/bugfixes for NER/IOB converters * Converter formats `ner` and `iob` use autodetect to choose a converter if possible * `iob2json` is reverted to handle sentence-per-line data like `word1\|pos1\|ent1 word2\|pos2\|ent2` * Fix bug in `merge_sentences()` so the second sentence in each batch isn't skipped * `conll_ner2json` is made more general so it can handle more formats with whitespace-separated columns * Supports all formats where the first column is the token and the final column is the IOB tag; if present, the second column is the POS tag * As in CoNLL 2003 NER, blank lines separate sentences, `-DOCSTART- -X- O O` separates documents * Add option for segmenting sentences (new flag `-s`) * Parser-based sentence segmentation with a provided model, otherwise with sentencizer (new option `-b` to specify model) * Can group sentences into documents with `n_sents` as long as sentence segmentation is available * Only applies automatic segmentation when there are no existing delimiters in the data * Provide info about settings applied during conversion with warnings and suggestions if settings conflict or might not be not optimal. * Add tests for common formats * Add '(default)' back to docs for -c auto * Add document count back to output * Revert changes to converter output message * Use explicit tabs in convert CLI test data * Adjust/add messages for n_sents=1 default * Add sample NER data to training examples * Update README * Add links in docs to example NER data * Define msg within converters		2019-08-29 12:04:01 +02:00
..
converters	Updates/bugfixes for NER/IOB converters (#4186 )	2019-08-29 12:04:01 +02:00
__init__.py	Move UD scripts to bin	2019-03-20 01:19:34 +01:00
_schemas.py	Store JSON schemas in Python and tidy up (#3235 )	2019-02-07 19:44:31 +11:00
convert.py	Updates/bugfixes for NER/IOB converters (#4186 )	2019-08-29 12:04:01 +02:00
debug_data.py	Tidy up and auto-format	2019-08-18 15:09:16 +02:00
download.py	Require downloaded model in pkg_resources (#4090 )	2019-08-07 13:18:11 +02:00
evaluate.py	Open file as utf-8 (closes #4138 )	2019-08-18 13:55:34 +02:00
info.py	Small CLI improvements (#3030 )	2018-12-08 11:49:43 +01:00
init_model.py	Tidy up and auto-format	2019-08-18 15:09:16 +02:00
link.py	Small CLI improvements (#3030 )	2018-12-08 11:49:43 +01:00
package.py	Also support "requirements" in model.json	2019-07-27 13:34:57 +02:00
pretrain.py	Fix absolute imports and avoid importing from cli	2019-08-20 15:08:59 +02:00
profile.py	Fix cytoolz import cytoolz	2018-12-06 16:04:12 +01:00
train.py	Fix #3830 : 'subtok' label being added even if learn_tokens=False (#4188 )	2019-08-23 17:54:00 +02:00
validate.py	Strip out .dev versions in spacy validate [ci skip]	2019-03-17 12:16:53 +01:00