LLM-Based Synthetic Data Generation for NLP Metric Validation
Generování syntetických dat pomocí LLM pro validaci evaluačních metrik v NLP
diploma thesis (DEFENDED)
View/ Open
Permanent link
http://hdl.handle.net/20.500.11956/209637Identifiers
Study Information System: 288612
Collections
- Kvalifikační práce [12377]
Author
Advisor
Referee
Kartáč, Ivan
Faculty / Institute
Faculty of Mathematics and Physics
Discipline
Computer Science - Artificial Intelligence
Department
Institute of Formal and Applied Linguistics
Date of defense
8. 6. 2026
Publisher
Univerzita Karlova, Matematicko-fyzikální fakultaLanguage
English
Grade
Excellent
Keywords (Czech)
large language models|natural language processing|automatic evaluation metrics|synthetic dataKeywords (English)
velké jazykové modely|zpracování přirozeného jazyka|automatické evaluační metriky|syntetická dataValidating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist for English datasets. We propose Meta-Judge, a scalable framework that uses LLMs to generate synthetic evaluation datasets via controlled semantic degradation of reference texts, replacing human judgment. We validate our approach using meta-correlation, measuring the alignment between metric rankings derived from synthetic data and those from human-annotated data. We experiment across Machine Translation, Question Answering, and Summarization in eight languages using 4 open-source LLMs. Large models achieve meta-correlation above 0.9 on question-answering datasets. To reduce inference cost, we finetune a 1B-parameter model using GRPO with an unsupervised ensemble of metrics, recovering most of the performance of large models.
