The Detection of Distributional Discrepancy for Text Generation

Xingyuan Chen; Ping Cai; Peng Jin; Haokun Du; Hongjun Wang; Xingyu Dai; Jiajun Chen

テキスト生成のための分布不一致の検出

ニューラル言語モデルによって生成されたテキストは、実際のテキストほど優れていません。これは、分布が異なることを意味します。それを軽減するために、Generative Adversarial Nets（GAN）が使用されます。ただし、一部の研究者は、GANバリアントはまったく機能しないと主張しています。サンプルの品質（Bleuなど）とサンプルの多様性（self-Bleuなど）の両方を考慮すると、GANバリアントは適切に調整された言語モデルよりもさらに悪化します。しかし、ブルーとセルフブルーは、この分布の不一致を正確に測定することはできません。実際、実際のテキストと生成されたテキストの分布の不一致をどのように測定するかは未解決の問題です。本論文では、実際のテキストと生成されたテキストの分布の違いを測定するための2つのメトリック関数を理論的に提案します。それに加えて、それらを推定する方法が提案されています。最初に、これら2つの関数を使用して言語モデルを評価し、違いが大きいことを確認します。次に、検出された不一致信号を使用してジェネレーターを改善するいくつかの方法を試します。ただし、その差は以前よりさらに大きくなります。 2つの既存の言語のGANを実験すると、実際のテキストと生成されたテキストの間の分布の不一致は、敵対的な学習ラウンドが増えるにつれて増加します。これらの両方の言語のGANが失敗することを示しています。

The text generated by neural language models is not as good as the real text. This means that their distributions are different. Generative Adversarial Nets (GAN) are used to alleviate it. However, some researchers argue that GAN variants do not work at all. When both sample quality (such as Bleu) and sample diversity (such as self-Bleu) are taken into account, the GAN variants even are worse than a well-adjusted language model. But, Bleu and self-Bleu can not precisely measure this distributional discrepancy. In fact, how to measure the distributional discrepancy between real text and generated text is still an open problem. In this paper, we theoretically propose two metric functions to measure the distributional difference between real text and generated text. Besides that, a method is put forward to estimate them. First, we evaluate language model with these two functions and find the difference is huge. Then, we try several methods to use the detected discrepancy signal to improve the generator. However the difference becomes even bigger than before. Experimenting on two existing language GANs, the distributional discrepancy between real text and generated text increases with more adversarial learning rounds. It demonstrates both of these language GANs fail.

updated: Sun Nov 24 2019 06:24:04 GMT+0000 (UTC)

published: Sat Sep 28 2019 07:12:34 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト