English grammar checker and corrector: the determiners

Auersperger, Michal

Korektor anglické gramatiky: určité a neurčité členy

dc.contributor.advisor	Pecina, Pavel
dc.creator	Auersperger, Michal
dc.date.accessioned	2017-06-28T10:02:03Z
dc.date.available	2017-06-28T10:02:03Z
dc.date.issued	2017
dc.identifier.uri	http://hdl.handle.net/20.500.11956/85647
dc.description.abstract	Předkládaná práce přistupuje ke kontrole členů v anglickém textu jako ke klasi- fikační úloze řešené metodami strojového učení s učitelem. Každé jmenné frázi v textu je přiřazena jedna ze tří tříd reprezentující určitý, neurčitý nebo nulový člen. V rámci úvodní rešerše byl definován článek dosahující na takto pojaté úloze ne- jlepších výsledků. Daný experiment byl pak zreplikován a překonán. Pomocí jiných signálů a volbou rozdílného učícího algoritmu došlo k poklesu chyby klasifikace o cca. 34%. Výsledný model byl pak porovnán s výkonem expertů na dané úloze. Přes problémy srovnání způsobené rozdílností dat se zdá, že je-li model použit na typu dat, na kterém byl trénován, je jeho úspěšnost srovnatelná s lidskou silou. Použití modelu na jiných datech se ale neosvědčilo. Stejně tak se neosvědčila ani náhrada klasifikátoru za jazykový model, který by předpovídal potenciální člen pro každou pozici ve větě. 1	cs_CZ
dc.description.abstract	Correction of the articles in English texts is approached as an article generation task, i.e. each noun phrase is assigned with a class corresponding to the definite, indefinite or zero article. Supervised machine learning methods are used to first replicate and then improve upon the best reported result in the literature known to the author. By feature engineering and a different choice of the learning method, about 34% drop in error is achieved. The resulting model is further compared to the performance of expert annotators. Although the comparison is not straightforward due to the differences in the data, the results indicate the performance of the trained model is comparable to the human-level performance when measured on the in-domain data. On the other hand, the model does not generalize well to different types of data. Using a large-scale language model to predict an article (or no article) for each word of the text has not proved successful. 1	en_US
dc.language	English	cs_CZ
dc.language.iso	en_US
dc.publisher	Univerzita Karlova, Matematicko-fyzikální fakulta	cs_CZ
dc.subject	Angličtina	cs_CZ
dc.subject	členy	cs_CZ
dc.subject	kontrola pravopisu	cs_CZ
dc.subject	English	en_US
dc.subject	determiners	en_US
dc.subject	grammar checker	en_US
dc.title	English grammar checker and corrector: the determiners	en_US
dc.type	diplomová práce	cs_CZ
dcterms.created	2017
dcterms.dateAccepted	2017-06-07
dc.description.department	Institute of Formal and Applied Linguistics	en_US
dc.description.department	Ústav formální a aplikované lingvistiky	cs_CZ
dc.description.faculty	Matematicko-fyzikální fakulta	cs_CZ
dc.description.faculty	Faculty of Mathematics and Physics	en_US
dc.identifier.repId	127237
dc.title.translated	Korektor anglické gramatiky: určité a neurčité členy	cs_CZ
dc.contributor.referee	Straňák, Pavel
thesis.degree.name	Mgr.
thesis.degree.level	navazující magisterské	cs_CZ
thesis.degree.discipline	Matematická lingvistika	cs_CZ
thesis.degree.discipline	Computational Linguistics	en_US
thesis.degree.program	Computer Science	en_US
thesis.degree.program	Informatika	cs_CZ
uk.thesis.type	diplomová práce	cs_CZ
uk.taxonomy.organization-cs	Matematicko-fyzikální fakulta::Ústav formální a aplikované lingvistiky	cs_CZ
uk.taxonomy.organization-en	Faculty of Mathematics and Physics::Institute of Formal and Applied Linguistics	en_US
uk.faculty-name.cs	Matematicko-fyzikální fakulta	cs_CZ
uk.faculty-name.en	Faculty of Mathematics and Physics	en_US
uk.faculty-abbr.cs	MFF	cs_CZ
uk.degree-discipline.cs	Matematická lingvistika	cs_CZ
uk.degree-discipline.en	Computational Linguistics	en_US
uk.degree-program.cs	Informatika	cs_CZ
uk.degree-program.en	Computer Science	en_US
thesis.grade.cs	Výborně	cs_CZ
thesis.grade.en	Excellent	en_US
uk.abstract.cs	Předkládaná práce přistupuje ke kontrole členů v anglickém textu jako ke klasi- fikační úloze řešené metodami strojového učení s učitelem. Každé jmenné frázi v textu je přiřazena jedna ze tří tříd reprezentující určitý, neurčitý nebo nulový člen. V rámci úvodní rešerše byl definován článek dosahující na takto pojaté úloze ne- jlepších výsledků. Daný experiment byl pak zreplikován a překonán. Pomocí jiných signálů a volbou rozdílného učícího algoritmu došlo k poklesu chyby klasifikace o cca. 34%. Výsledný model byl pak porovnán s výkonem expertů na dané úloze. Přes problémy srovnání způsobené rozdílností dat se zdá, že je-li model použit na typu dat, na kterém byl trénován, je jeho úspěšnost srovnatelná s lidskou silou. Použití modelu na jiných datech se ale neosvědčilo. Stejně tak se neosvědčila ani náhrada klasifikátoru za jazykový model, který by předpovídal potenciální člen pro každou pozici ve větě. 1	cs_CZ
uk.abstract.en	Correction of the articles in English texts is approached as an article generation task, i.e. each noun phrase is assigned with a class corresponding to the definite, indefinite or zero article. Supervised machine learning methods are used to first replicate and then improve upon the best reported result in the literature known to the author. By feature engineering and a different choice of the learning method, about 34% drop in error is achieved. The resulting model is further compared to the performance of expert annotators. Although the comparison is not straightforward due to the differences in the data, the results indicate the performance of the trained model is comparable to the human-level performance when measured on the in-domain data. On the other hand, the model does not generalize well to different types of data. Using a large-scale language model to predict an article (or no article) for each word of the text has not proved successful. 1	en_US
uk.file-availability	V
uk.publication.place	Praha	cs_CZ
uk.grantor	Univerzita Karlova, Matematicko-fyzikální fakulta, Ústav formální a aplikované lingvistiky	cs_CZ