spaCy/spacy/cli/converters/iob2json.py

# coding: utf8
from __future__ import unicode_literals

import re

from ...gold import iob_to_biluo
from ...util import minibatch


def iob2json(input_data, n_sents=10, *args, **kwargs):
    """
    Convert IOB files into JSON format for use with train cli.
    """
    sentences = read_iob(input_data.split("\n"))
    docs = merge_sentences(sentences, n_sents)
    return docs


def read_iob(raw_sents):
    sentences = []
    for line in raw_sents:
        if not line.strip():
            continue
        tokens = [re.split("[^\w\-]", line.strip())]
        if len(tokens[0]) == 3:
            words, pos, iob = zip(*tokens)
        elif len(tokens[0]) == 2:
            words, iob = zip(*tokens)
            pos = ["-"] * len(words)
        else:
            raise ValueError(
                "The iob/iob2 file is not formatted correctly. Try checking whitespace and delimiters."
            )
        biluo = iob_to_biluo(iob)
        sentences.append(
            [
                {"orth": w, "tag": p, "ner": ent}
                for (w, p, ent) in zip(words, pos, biluo)
            ]
        )
    sentences = [{"tokens": sent} for sent in sentences]
    paragraphs = [{"sentences": [sent]} for sent in sentences]
    docs = [{"id": 0, "paragraphs": [para]} for para in paragraphs]
    return docs


def merge_sentences(docs, n_sents):
    merged = []
    for group in minibatch(docs, size=n_sents):
        group = list(group)
        first = group.pop(0)
        to_extend = first["paragraphs"][0]["sentences"]
        for sent in group[1:]:
            to_extend.extend(sent["paragraphs"][0]["sentences"])
        merged.append(first)
    return merged
Add incomplete iob converter 2017-05-19 18:27:51 +00:00			`# coding: utf8`
			`from __future__ import unicode_literals`

Tidy up merge conflict leftovers 2018-12-18 12:58:30 +00:00			`import re`

Fix converters 2017-05-26 16:32:34 +00:00			`from ...gold import iob_to_biluo`
Replace cytoolz.partition_all with util.minibatch 2019-05-11 19:12:09 +00:00			`from ...util import minibatch`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00

💫 New JSON helpers, training data internals & CLI rewrite (#2932) * Support nowrap setting in util.prints * Tidy up and fix whitespace * Simplify script and use read_jsonl helper * Add JSON schemas (see #2928) * Deprecate Doc.print_tree Will be replaced with Doc.to_json, which will produce a unified format * Add Doc.to_json() method (see #2928) Converts Doc objects to JSON using the same unified format as the training data. Method also supports serializing selected custom attributes in the doc._. space. * Remove outdated test * Add write_json and write_jsonl helpers * WIP: Update spacy train * Tidy up spacy train * WIP: Use wasabi for formatting * Add GoldParse helpers for JSON format * WIP: add debug-data command * Fix typo * Add missing import * Update wasabi pin * Add missing import * 💫 Refactor CLI (#2943) To be merged into #2932. ## Description - [x] refactor CLI To use [`wasabi`](https://github.com/ines/wasabi) - [x] use [`black`](https://github.com/ambv/black) for auto-formatting - [x] add `flake8` config - [x] move all messy UD-related scripts to `cli.ud` - [x] make converters function that take the opened file and return the converted data (instead of having them handle the IO) ### Types of change enhancement ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Update wasabi pin * Delete old test * Update errors * Fix typo * Tidy up and format remaining code * Fix formatting * Improve formatting of messages * Auto-format remaining code * Add tok2vec stuff to spacy.train * Fix typo * Update wasabi pin * Fix path checks for when train() is called as function * Reformat and tidy up pretrain script * Update argument annotations * Raise error if model language doesn't match lang * Document new train command 2018-11-30 19:16:14 +00:00			`def iob2json(input_data, n_sents=10, args, *kwargs):`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00			`"""`
			`Convert IOB files into JSON format for use with train cli.`
			`"""`
Fix .iob converter (closes #3620) 2019-05-11 17:15:26 +00:00			`sentences = read_iob(input_data.split("\n"))`
			`docs = merge_sentences(sentences, n_sents)`
💫 New JSON helpers, training data internals & CLI rewrite (#2932) * Support nowrap setting in util.prints * Tidy up and fix whitespace * Simplify script and use read_jsonl helper * Add JSON schemas (see #2928) * Deprecate Doc.print_tree Will be replaced with Doc.to_json, which will produce a unified format * Add Doc.to_json() method (see #2928) Converts Doc objects to JSON using the same unified format as the training data. Method also supports serializing selected custom attributes in the doc._. space. * Remove outdated test * Add write_json and write_jsonl helpers * WIP: Update spacy train * Tidy up spacy train * WIP: Use wasabi for formatting * Add GoldParse helpers for JSON format * WIP: add debug-data command * Fix typo * Add missing import * Update wasabi pin * Add missing import * 💫 Refactor CLI (#2943) To be merged into #2932. ## Description - [x] refactor CLI To use [`wasabi`](https://github.com/ines/wasabi) - [x] use [`black`](https://github.com/ambv/black) for auto-formatting - [x] add `flake8` config - [x] move all messy UD-related scripts to `cli.ud` - [x] make converters function that take the opened file and return the converted data (instead of having them handle the IO) ### Types of change enhancement ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Update wasabi pin * Delete old test * Update errors * Fix typo * Tidy up and format remaining code * Fix formatting * Improve formatting of messages * Auto-format remaining code * Add tok2vec stuff to spacy.train * Fix typo * Update wasabi pin * Fix path checks for when train() is called as function * Reformat and tidy up pretrain script * Update argument annotations * Raise error if model language doesn't match lang * Document new train command 2018-11-30 19:16:14 +00:00			`return docs`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00

Fix concatenation in iob2json converter 2017-10-02 14:50:26 +00:00			`def read_iob(raw_sents):`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00			`sentences = []`
Fix concatenation in iob2json converter 2017-10-02 14:50:26 +00:00			`for line in raw_sents:`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00			`if not line.strip():`
			`continue`
Merge branch 'master' into develop 2018-12-18 12:48:10 +00:00			`tokens = [re.split("[^\w\-]", line.strip())]`
Handle iob with no tag in converter 2017-05-28 13:11:39 +00:00			`if len(tokens[0]) == 3:`
			`words, pos, iob = zip(*tokens)`
iob converter: add 'exception' for error 'too many values' (#3159) * added contributor agreement * issue #3128 throw exception on bad IOB/2 formatting * Update spacy/cli/converters/iob2json.py with ValueError Co-Authored-By: gavrieltal <gtloria@protonmail.com> 2019-01-16 12:44:16 +00:00			`elif len(tokens[0]) == 2:`
Handle iob with no tag in converter 2017-05-28 13:11:39 +00:00			`words, iob = zip(*tokens)`
💫 New JSON helpers, training data internals & CLI rewrite (#2932) * Support nowrap setting in util.prints * Tidy up and fix whitespace * Simplify script and use read_jsonl helper * Add JSON schemas (see #2928) * Deprecate Doc.print_tree Will be replaced with Doc.to_json, which will produce a unified format * Add Doc.to_json() method (see #2928) Converts Doc objects to JSON using the same unified format as the training data. Method also supports serializing selected custom attributes in the doc._. space. * Remove outdated test * Add write_json and write_jsonl helpers * WIP: Update spacy train * Tidy up spacy train * WIP: Use wasabi for formatting * Add GoldParse helpers for JSON format * WIP: add debug-data command * Fix typo * Add missing import * Update wasabi pin * Add missing import * 💫 Refactor CLI (#2943) To be merged into #2932. ## Description - [x] refactor CLI To use [`wasabi`](https://github.com/ines/wasabi) - [x] use [`black`](https://github.com/ambv/black) for auto-formatting - [x] add `flake8` config - [x] move all messy UD-related scripts to `cli.ud` - [x] make converters function that take the opened file and return the converted data (instead of having them handle the IO) ### Types of change enhancement ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Update wasabi pin * Delete old test * Update errors * Fix typo * Tidy up and format remaining code * Fix formatting * Improve formatting of messages * Auto-format remaining code * Add tok2vec stuff to spacy.train * Fix typo * Update wasabi pin * Fix path checks for when train() is called as function * Reformat and tidy up pretrain script * Update argument annotations * Raise error if model language doesn't match lang * Document new train command 2018-11-30 19:16:14 +00:00			`pos = ["-"] * len(words)`
iob converter: add 'exception' for error 'too many values' (#3159) * added contributor agreement * issue #3128 throw exception on bad IOB/2 formatting * Update spacy/cli/converters/iob2json.py with ValueError Co-Authored-By: gavrieltal <gtloria@protonmail.com> 2019-01-16 12:44:16 +00:00			`else:`
Merge branch 'master' into develop 2019-02-07 19:54:07 +00:00			`raise ValueError(`
			`"The iob/iob2 file is not formatted correctly. Try checking whitespace and delimiters."`
			`)`
Fix converters 2017-05-26 16:32:34 +00:00			`biluo = iob_to_biluo(iob)`
💫 New JSON helpers, training data internals & CLI rewrite (#2932) * Support nowrap setting in util.prints * Tidy up and fix whitespace * Simplify script and use read_jsonl helper * Add JSON schemas (see #2928) * Deprecate Doc.print_tree Will be replaced with Doc.to_json, which will produce a unified format * Add Doc.to_json() method (see #2928) Converts Doc objects to JSON using the same unified format as the training data. Method also supports serializing selected custom attributes in the doc._. space. * Remove outdated test * Add write_json and write_jsonl helpers * WIP: Update spacy train * Tidy up spacy train * WIP: Use wasabi for formatting * Add GoldParse helpers for JSON format * WIP: add debug-data command * Fix typo * Add missing import * Update wasabi pin * Add missing import * 💫 Refactor CLI (#2943) To be merged into #2932. ## Description - [x] refactor CLI To use [`wasabi`](https://github.com/ines/wasabi) - [x] use [`black`](https://github.com/ambv/black) for auto-formatting - [x] add `flake8` config - [x] move all messy UD-related scripts to `cli.ud` - [x] make converters function that take the opened file and return the converted data (instead of having them handle the IO) ### Types of change enhancement ## Checklist <!--- Before you submit the PR, go over this checklist and make sure you can tick off all the boxes. [] -> [x] --> - [x] I have submitted the spaCy Contributor Agreement. - [x] I ran the tests, and all new and existing tests passed. - [x] My changes don't require a change to the documentation, or if they do, I've added all required information. * Update wasabi pin * Delete old test * Update errors * Fix typo * Tidy up and format remaining code * Fix formatting * Improve formatting of messages * Auto-format remaining code * Add tok2vec stuff to spacy.train * Fix typo * Update wasabi pin * Fix path checks for when train() is called as function * Reformat and tidy up pretrain script * Update argument annotations * Raise error if model language doesn't match lang * Document new train command 2018-11-30 19:16:14 +00:00			`sentences.append(`
			`[`
			`{"orth": w, "tag": p, "ner": ent}`
			`for (w, p, ent) in zip(words, pos, biluo)`
			`]`
			`)`
			`sentences = [{"tokens": sent} for sent in sentences]`
			`paragraphs = [{"sentences": [sent]} for sent in sentences]`
			`docs = [{"id": 0, "paragraphs": [para]} for para in paragraphs]`
Add incomplete iob converter 2017-05-19 18:27:51 +00:00			`return docs`
Fix .iob converter (closes #3620) 2019-05-11 17:15:26 +00:00

			`def merge_sentences(docs, n_sents):`
			`merged = []`
Replace cytoolz.partition_all with util.minibatch 2019-05-11 19:12:09 +00:00			`for group in minibatch(docs, size=n_sents):`
Fix .iob converter (closes #3620) 2019-05-11 17:15:26 +00:00			`group = list(group)`
			`first = group.pop(0)`
			`to_extend = first["paragraphs"][0]["sentences"]`
			`for sent in group[1:]:`
			`to_extend.extend(sent["paragraphs"][0]["sentences"])`
			`merged.append(first)`
			`return merged`