The cluster structure function

Andrew R. Cohen; Paul M. B. Vitányi

クラスター構造機能

与えられた数の部分へのデータセットの各パーティションには、すべての部分がその部分のデータの可能な限り良いモデル（「アルゴリズムの十分統計量」）になるようなパーティションがあります。これは、1からデータ数までのすべての数に対して実行できるため、結果は関数、クラスター構造関数になります。パーティションのパーツの数を、パーツによる優れたモデルであるという欠陥に関連する値にマップします。このような関数は、データセットが分割されていない場合は少なくともゼロの値で始まり、データセットがシングルトン部分に分割されている場合はゼロに下がります。最適なクラスタリングは、クラスター構造関数を最小化するために選択されたものです。この方法の背後にある理論は、アルゴリズム情報理論（コルモゴロフ複雑性）で表されます。実際には、関連するコルモゴロフの複雑さは、コンクリートコンプレッサーによって近似されます。実際のデータセットを使用した例を示します。MNISTの手書き数字と、幹細胞研究で使用される実際の細胞のセグメンテーションです。

For each partition of a data set into a given number of parts there is a partition such that every part is as much as possible a good model (an "algorithmic sufficient statistic") for the data in that part. Since this can be done for every number between one and the number of data, the result is a function, the cluster structure function. It maps the number of parts of a partition to values related to the deficiencies of being good models by the parts. Such a function starts with a value at least zero for no partition of the data set and descents to zero for the partition of the data set into singleton parts. The optimal clustering is the one chosen to minimize the cluster structure function. The theory behind the method is expressed in algorithmic information theory (Kolmogorov complexity). In practice the Kolmogorov complexities involved are approximated by a concrete compressor. We give examples using real data sets: the MNIST handwritten digits and the segmentation of real cells as used in stem cell research.

updated: Tue Jan 04 2022 16:05:59 GMT+0000 (UTC)

published: Tue Jan 04 2022 16:05:59 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト