Skip to content

About

Penn Treebank POS tagger and lemmatizer for English text.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

english-pos-tagger

Part-of-speech tagging and lemmatization of English text.

The tagger assigns Penn-Treebank-style POS tags and adds base-form lemmas for nouns (excluding proper nouns), verbs, and adjectives. It runs a unigram pass for candidate tags, rewrites tags with ordered context rules that read neighboring tokens, then lemmatizes. Tagging needs no I/O.

Installation

[dependencies]
english-pos-tagger = "0.2"

Usage

Tag a sentence:

use english_pos_tagger::Tagger;

let tagger = Tagger::new();
let tokens = tagger.tag_sentence("A bear just crossed the road.");
assert_eq!(tokens[1].value, "bear");
assert_eq!(tokens[1].pos.as_deref(), Some("NN"));
assert_eq!(tokens[1].lemma.as_deref(), Some("bear"));

Tag pre-tokenized input with tag, raw strings with tag_raw_tokens, and extend the lexicon with update_lexicon. Each key normalizes before storage, and each value lists candidate tags with the most frequent first.

use english_pos_tagger::Tagger;

let mut tagger = Tagger::new();
tagger.update_lexicon([("Obama", vec!["NNP"])]);

Token fields

Each token carries value and tag from tokenization. Tagging fills normal and pos, and adds lemma for nouns, verbs, adjectives, and contraction forms. entity_type carries through from named-entity input.

License

Licensed under the MIT license.

About

Penn Treebank POS tagger and lemmatizer for English text.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages