Part-of-speech tagging and lemmatization of English text.
The tagger assigns Penn-Treebank-style POS tags and adds base-form lemmas for nouns (excluding proper nouns), verbs, and adjectives. It runs a unigram pass for candidate tags, rewrites tags with ordered context rules that read neighboring tokens, then lemmatizes. Tagging needs no I/O.
[dependencies]
english-pos-tagger = "0.2"Tag a sentence:
use english_pos_tagger::Tagger;
let tagger = Tagger::new();
let tokens = tagger.tag_sentence("A bear just crossed the road.");
assert_eq!(tokens[1].value, "bear");
assert_eq!(tokens[1].pos.as_deref(), Some("NN"));
assert_eq!(tokens[1].lemma.as_deref(), Some("bear"));Tag pre-tokenized input with tag, raw strings with tag_raw_tokens, and
extend the lexicon with update_lexicon. Each key normalizes before storage,
and each value lists candidate tags with the most frequent first.
use english_pos_tagger::Tagger;
let mut tagger = Tagger::new();
tagger.update_lexicon([("Obama", vec!["NNP"])]);Each token carries value and tag from tokenization. Tagging fills normal
and pos, and adds lemma for nouns, verbs, adjectives, and contraction forms.
entity_type carries through from named-entity input.
Licensed under the MIT license.