German Compounds in Transformer Models

Neumannová, Kristýna

Německé složeniny v modelech typu Transformer

dc.contributor.advisor	Bojar, Ondřej
dc.creator	Neumannová, Kristýna
dc.date.accessioned	2023-07-24T12:34:28Z
dc.date.available	2023-07-24T12:34:28Z
dc.date.issued	2023
dc.identifier.uri	http://hdl.handle.net/20.500.11956/181571
dc.description.abstract	German is known for its highly productive word formation processes, particularly in the area of compounding and derivation. In this thesis, we focus on German nominal compounds and their representation in machine translation (MT) outputs. Despite their importance in German text, commonly used metrics for MT evaluation, such as BLEU, do not adequately capture the usage of compounds. The aim of this thesis was to investigate the generation of German compounds in Transformer models and to explore the conditions that lead to their production. Our analysis revealed that MT systems tend to produce fewer compounds than humans. However, we found that due to the highly productive nature of German compounds, it is not feasible to identify them based on a fixed list. Therefore, we manually identified novel compounds, and even then, human translations still contained more compounds than MT systems. We trained our own Transformer model for English-German translation and conducted experiments to examine various factors that influence the production of compounds, in- cluding word segmentation and the frequency of compounds in the training data. Addi- tionally, we explored the use of forced decoding and the impact of providing the model with the first words of a sentence during translation. Our findings highlight the...	en_US
dc.description.abstract	Němčina je známá svou velmi produktivní slovotvorbou, zejména v oblasti kompoz- ice a derivace. V této práci se zaměřujeme na německé nominální složeniny a jejich zastoupení ve výstupech strojového překladu. Navzdory jejich důležitosti v německých textech, běžně používané metriky pro hodnocení kvality překladu, jako je BLEU, ne- dokážou použití složenin dostatečně zachytit. Cílem této práce bylo zkoumat generování německých složenin v modelech typu Transformer a prozkoumat faktory, které vedou k jejich tvorbě. Zjistili jsme, že strojové překladové systémy produkují méně složenin než lidé. Také se ukázalo, že kvůli velmi produktivní povaze německých složenin není možné je identifikovat na základě fixního seznamu. I po ručním vyhledání nových kompozit jich lidské překlady obsahovaly více než strojové. Natrénovali jsme vlastní model typu Transformer pro překlad z angličtiny do němčiny, abychom to mohli zkoumat různé faktory, které ovlivňují produkci složenin, včetně seg- mentace slov a frekvence složenin v trénovacích datech. Dále jsme experimentovali s vynuceným dekódováním (forced decoding) a zjišťovali, jak se změní výstup systému po poskytnutí prvních slov překládané věty. Naše výsledky zdůrazňují důležitost dalšího výzkumu v oblasti strojového překladu, aby se byly překladové systémy schopny lépe...	cs_CZ
dc.language	English	cs_CZ
dc.language.iso	en_US
dc.publisher	Univerzita Karlova, Matematicko-fyzikální fakulta	cs_CZ
dc.subject	Transformer\|Machine Translation\|German compounds\|Machine translation quality	en_US
dc.subject	Transformátor\|Strojový překlad\|Německá kompozita\|Kvalita strojového překladu	cs_CZ
dc.title	German Compounds in Transformer Models	en_US
dc.type	diplomová práce	cs_CZ
dcterms.created	2023
dcterms.dateAccepted	2023-06-06
dc.description.department	Ústav formální a aplikované lingvistiky	cs_CZ
dc.description.department	Institute of Formal and Applied Linguistics	en_US
dc.description.faculty	Faculty of Mathematics and Physics	en_US
dc.description.faculty	Matematicko-fyzikální fakulta	cs_CZ
dc.identifier.repId	253673
dc.title.translated	Německé složeniny v modelech typu Transformer	cs_CZ
dc.contributor.referee	Zeman, Daniel
thesis.degree.name	Mgr.
thesis.degree.level	navazující magisterské	cs_CZ
thesis.degree.discipline	Informatika - Jazykové technologie a počítačová lingvistika	cs_CZ
thesis.degree.discipline	Computer Science - Language Technologies and Computational Linguistics	en_US
thesis.degree.program	Informatika - Jazykové technologie a počítačová lingvistika	cs_CZ
thesis.degree.program	Computer Science - Language Technologies and Computational Linguistics	en_US
uk.thesis.type	diplomová práce	cs_CZ
uk.taxonomy.organization-cs	Matematicko-fyzikální fakulta::Ústav formální a aplikované lingvistiky	cs_CZ
uk.taxonomy.organization-en	Faculty of Mathematics and Physics::Institute of Formal and Applied Linguistics	en_US
uk.faculty-name.cs	Matematicko-fyzikální fakulta	cs_CZ
uk.faculty-name.en	Faculty of Mathematics and Physics	en_US
uk.faculty-abbr.cs	MFF	cs_CZ
uk.degree-discipline.cs	Informatika - Jazykové technologie a počítačová lingvistika	cs_CZ
uk.degree-discipline.en	Computer Science - Language Technologies and Computational Linguistics	en_US
uk.degree-program.cs	Informatika - Jazykové technologie a počítačová lingvistika	cs_CZ
uk.degree-program.en	Computer Science - Language Technologies and Computational Linguistics	en_US
thesis.grade.cs	Výborně	cs_CZ
thesis.grade.en	Excellent	en_US
uk.abstract.cs	Němčina je známá svou velmi produktivní slovotvorbou, zejména v oblasti kompoz- ice a derivace. V této práci se zaměřujeme na německé nominální složeniny a jejich zastoupení ve výstupech strojového překladu. Navzdory jejich důležitosti v německých textech, běžně používané metriky pro hodnocení kvality překladu, jako je BLEU, ne- dokážou použití složenin dostatečně zachytit. Cílem této práce bylo zkoumat generování německých složenin v modelech typu Transformer a prozkoumat faktory, které vedou k jejich tvorbě. Zjistili jsme, že strojové překladové systémy produkují méně složenin než lidé. Také se ukázalo, že kvůli velmi produktivní povaze německých složenin není možné je identifikovat na základě fixního seznamu. I po ručním vyhledání nových kompozit jich lidské překlady obsahovaly více než strojové. Natrénovali jsme vlastní model typu Transformer pro překlad z angličtiny do němčiny, abychom to mohli zkoumat různé faktory, které ovlivňují produkci složenin, včetně seg- mentace slov a frekvence složenin v trénovacích datech. Dále jsme experimentovali s vynuceným dekódováním (forced decoding) a zjišťovali, jak se změní výstup systému po poskytnutí prvních slov překládané věty. Naše výsledky zdůrazňují důležitost dalšího výzkumu v oblasti strojového překladu, aby se byly překladové systémy schopny lépe...	cs_CZ
uk.abstract.en	German is known for its highly productive word formation processes, particularly in the area of compounding and derivation. In this thesis, we focus on German nominal compounds and their representation in machine translation (MT) outputs. Despite their importance in German text, commonly used metrics for MT evaluation, such as BLEU, do not adequately capture the usage of compounds. The aim of this thesis was to investigate the generation of German compounds in Transformer models and to explore the conditions that lead to their production. Our analysis revealed that MT systems tend to produce fewer compounds than humans. However, we found that due to the highly productive nature of German compounds, it is not feasible to identify them based on a fixed list. Therefore, we manually identified novel compounds, and even then, human translations still contained more compounds than MT systems. We trained our own Transformer model for English-German translation and conducted experiments to examine various factors that influence the production of compounds, in- cluding word segmentation and the frequency of compounds in the training data. Addi- tionally, we explored the use of forced decoding and the impact of providing the model with the first words of a sentence during translation. Our findings highlight the...	en_US
uk.file-availability	V
uk.grantor	Univerzita Karlova, Matematicko-fyzikální fakulta, Ústav formální a aplikované lingvistiky	cs_CZ
thesis.grade.code	1
uk.publication-place	Praha	cs_CZ
uk.thesis.defenceStatus	O