Jump to content

Archive:Brill tagger: Difference between revisions

From IdeaWazaWiki
No edit summary
wikademia>Daveliney
m sp
Line 14: Line 14:
** Select the best (higher score) rule.  
** Select the best (higher score) rule.  
** Add it to the rule set and apply it to the text.
** Add it to the rule set and apply it to the text.
** Repeat until no rule has a score above a given threshold (that is, untill applying new rules leaves the text in the same state, which is then supposed to be the final state of the tagging).
** Repeat until no rule has a score above a given threshold (that is, until applying new rules leaves the text in the same state, which is then supposed to be the final state of the tagging).


== Rules ==
== Rules ==

Revision as of 15:33, 21 July 2005

The Brill Tagger was exposed by Eric Brill in his 1993 PhD thesis [1]. It can be summarised as an "error-driven transformation-based tagger". It is

  • error-driven in the sense that is recourses to supervised learning
  • transformation-based in the sense that it uses rules

Algorithm

The algorithm goes as follow:

  • Initialisation:
    • Known words (in vocabulary): assigning the most frequent tag associated to a form of the word
    • Unknown words (out of vocabulary) :
      • Proper noun if capitalised and simple noun else (1992)
      • Learning or guessing rules on the same basis as contextual rules (1994)
  • Learning Phase
    • Iteratively compute the error score of each candidate rule (difference between the number of errors before and after applying the rule)
    • Select the best (higher score) rule.
    • Add it to the rule set and apply it to the text.
    • Repeat until no rule has a score above a given threshold (that is, until applying new rules leaves the text in the same state, which is then supposed to be the final state of the tagging).

Rules

Lexical rules are used for the initialisation, and contextual rules are used to correct the tags.

  • Lexical rules: word → tag IF Condition (example: identification of suffixes like "-tion")
  • Contextual rules: tag1 → tag2 IF Condition (example: "preceding/following tag is X", "preceding/following word is w")

Template:Compu-stub