Search CORE

773 research outputs found

동종, 이종, 그리고 나무 형태의 그래프를 위한 비지도 표현 학습

Author: 박지웅
Publication venue: 서울대학교 대학원
Publication date: 01/08/2022
Field of study

학위논문(박사) -- 서울대학교대학원 : 공과대학 전기·정보공학부, 2022. 8. 최진영.그래프 데이터에 대한 비지도 표현 학습의 목적은 그래프의 구조와 노드의 속성을 잘 반영하는 유용한 노드 단위 혹은 그래프 단위의 벡터 형태 표현을 학습하는 것이다. 최근, 그래프 데이터에 대해 강력한 표현 학습 능력을 갖춘 그래프 신경망을 활용한 비지도 그래프 표현 학습 모델의 설계가 주목을 받고 있다. 많은 방법들은 한 종류의 엣지와 한 종류의 노드가 존재하는 동종 그래프에 대한 학습에 집중을 한다. 하지만 이 세상에 수많은 종류의 관계가 존재하기 때문에, 그래프 또한 구조적, 의미론적 속성을 통해 다양한 종류로 분류할 수 있다. 그래서, 그래프로부터 유용한 표현을 학습하기 위해서는 비지도 학습 프레임워크는 입력 그래프의 특징을 제대로 고려해야만 한다. 본 학위논문에서 우리는 널리 접할 수 있는 세가지 그래프 구조인 동종 그래프, 트리 형태의 그래프, 그리고 이종 그래프에 대한 그래프 신경망을 활용하는 비지도 학습 모델들을 제안한다. 처음으로, 우리는 동종 그래프의 노드에 대하여 저차원 표현을 학습하는 그래프 컨볼루션 오토인코더 모델을 제안한다. 기존의 그래프 오토인코더는 구조의 전체가 학습이 불가능해서 제한적인 표현 학습 능력을 가질 수 있는 반면에, 제안하는 오토인코더는 노드의 피쳐를 복원하며,구조의 전체가 학습이 가능하다. 노드의 피쳐를 복원하기 위해서, 우리는 인코더 부분의 역할이 이웃한 노드끼리 유사한 표현을 가지게 하는 라플라시안 스무딩이라는 것에 주목하여 디코더 부분에서는 이웃 노드의 표현과 멀어지게 하는 라플라시안 샤프닝을 하도록 설계하였다. 또한 라플라시안 샤프닝을 그대로 적용하면 불안정성을 유발할 수 있기 때문에, 엣지의 가중치 값에 음의 값을 줄 수 있는 부호형 그래프를 활용하여 안정적인 라플라시안 샤프닝의 형태를 제안하였다. 동종 그래프에 대한 노드 클러스터링과 링크 예측 실험을 통하여 제안하는 방법이 안정적으로 우수한 성능을 보임을 확인하였다. 둘째로, 우리는 트리의 형태를 가지는 계층적인 관계를 가지고 있는 그래프의 노드 표현을 정확하게 학습하기 위하여 쌍곡선 공간에서 동작하는 오토인코더 모델을 제안한다. 유클리디언 공간은 트리를 사상하기에 부적절하다는 최근의 분석을 통하여, 쌍곡선 공간에서 그래프 신경망의 레이어를 활용하여 노드의 저차원 표현을 학습하게 된다. 이 때, 그래프 신경망이 쌍곡선 기하학에서 계층 정보를 담고 있는 거리의 값을 활용하여 노드의 이웃사이의 중요도를 활용하도록 설계하였다. 우리는 논문 인용 관계 네트워크, 계통도, 이미지 사이의 네트워크등에 대해 제안한 모델을 적용하여 노드 클러스터링과 링크 예측 실험을 하였으며, 트리의 형태를 가지는 그래프에 대해서 제안한 모델이 유클리디언 공간에서 수행하는 모델에 비해 향상된 성능을 보였다는 것을 확인하였다. 마지막으로, 우리는 여러 종류의 노드와 엣지를 가지는 이종그래프에 대한 대조 학습 모델을 제안한다. 우리는 기존의 방법들이 학습하기 이전에 충분한 도메인 지식을 사용하여 설계한 메타패스나 메타그래프에 의존한다는 단점과 많은 이종그래프의 엣지가 다른 노드 종류사이의 관계에 집중하고 있다는 점을 주목하였다. 이를 통해 우리는 사전과정이 필요없으며 다른 종류 사이의 관계에 더하여 같은 종류 사이의 관계도 동시에 효율적으로 학습하게 하는 메타노드라는 개념을 제안하였다. 또한 메타노드를 기반으로하는 그래프 신경망과 대조 학습 모델을 제안하였다. 우리는 제안한 모델을 메타패스를 사용하는 이종그래프 학습 모델과 노드 클러스터링 등의 실험 성능으로 비교해보았을 때, 비등하거나 높은 성능을 보였음을 확인하였다.The goal of unsupervised graph representation learning is extracting useful node-wise or graph-wise vector representation that is aware of the intrinsic structures of the graph and its attributes. These days, designing methodology of unsupervised graph representation learning based on graph neural networks has growing attention due to their powerful representation ability. Many methods are focused on a homogeneous graph that is a network with a single type of node and a single type of edge. However, as many types of relationships exist in this world, graphs can also be classified into various types by structural and semantic properties. For this reason, to learn useful representations from graphs, the unsupervised learning framework must consider the characteristics of the input graph. In this dissertation, we focus on designing unsupervised learning models using graph neural networks for three graph structures that are widely available: homogeneous graphs, tree-like graphs, and heterogeneous graphs. First, we propose a symmetric graph convolutional autoencoder which produces a low-dimensional latent representation from a homogeneous graph. In contrast to the existing graph autoencoders with asymmetric decoder parts, the proposed autoencoder has a newly designed decoder which builds a completely symmetric autoencoder form. For the reconstruction of node features, the decoder is designed based on Laplacian sharpening as the counterpart of Laplacian smoothing of the encoder, which allows utilizing the graph structure in the whole processes of the proposed autoencoder architecture. In order to prevent the numerical instability of the network caused by the Laplacian sharpening introduction, we further propose a new numerically stable form of the Laplacian sharpening by incorporating the signed graphs. The experimental results of clustering, link prediction and visualization tasks on homogeneous graphs strongly support that the proposed model is stable and outperforms various state-of-the-art algorithms. Second, we analyze how unsupervised tasks can benefit from learned representations in hyperbolic space. To explore how well the hierarchical structure of unlabeled data can be represented in hyperbolic spaces, we design a novel hyperbolic message passing autoencoder whose overall auto-encoding is performed in hyperbolic space. The proposed model conducts auto-encoding the networks via fully utilizing hyperbolic geometry in message passing. Through extensive quantitative and qualitative analyses, we validate the properties and benefits of the unsupervised hyperbolic representations of tree-like graphs. Third, we propose the novel concept of metanode for message passing to learn both heterogeneous and homogeneous relationships between any two nodes without meta-paths and meta-graphs. Unlike conventional methods, metanodes do not require a predetermined step to manipulate the given relations between different types to enrich relational information. Going one step further, we propose a metanode-based message passing layer and a contrastive learning model using the proposed layer. In our experiments, we show the competitive performance of the proposed metanode-based message passing method on node clustering and node classification tasks, when compared to state-of-the-art methods for message passing networks for heterogeneous graphs.1 Introduction 1 2 Representation Learning on Graph-Structured Data 4 2.1 Basic Introduction 4 2.1.1 Notations 5 2.2 Traditional Approaches 5 2.2.1 Graph Statistic 5 2.2.2 Neighborhood Overlap 7 2.2.3 Graph Kernel 9 2.2.4 Spectral Approaches 10 2.3 Node Embeddings I: Factorization and Random Walks 15 2.3.1 Factorization-based Methods 15 2.3.2 Random Walk-based Methods 16 2.4 Node Embeddings II: Graph Neural Networks 17 2.4.1 Overview of Framework 17 2.4.2 Representative Models 18 2.5 Learning in Unsupervised Environments 21 2.5.1 Predictive Coding 21 2.5.2 Contrastive Coding 22 2.6 Applications 24 2.6.1 Classifications 24 2.6.2 Link Prediction 26 3 Autoencoder Architecture for Homogeneous Graphs 27 3.1 Overview 27 3.2 Preliminaries 30 3.2.1 Spectral Convolution on Graphs 30 3.2.2 Laplacian Smoothing 32 3.3 Methodology 33 3.3.1 Laplacian Sharpening 33 3.3.2 Numerically Stable Laplacian Sharpening 34 3.3.3 Subspace Clustering Cost for Image Clustering 37 3.3.4 Training 39 3.4 Experiments 40 3.4.1 Datasets 40 3.4.2 Experimental Settings 42 3.4.3 Comparing Methods 42 3.4.4 Node Clustering 43 3.4.5 Image Clustering 45 3.4.6 Ablation Studies 46 3.4.7 Link Prediction 47 3.4.8 Visualization 47 3.5 Summary 49 4 Autoencoder Architecture for Tree-like Graphs 50 4.1 Overview 50 4.2 Preliminaries 52 4.2.1 Hyperbolic Embeddings 52 4.2.2 Hyperbolic Geometry 53 4.3 Methodology 55 4.3.1 Geometry-Aware Message Passing 56 4.3.2 Nonlinear Activation 57 4.3.3 Loss Function 58 4.4 Experiments 58 4.4.1 Datasets 59 4.4.2 Compared Methods 61 4.4.3 Experimental Details 62 4.4.4 Node Clustering and Link Prediction 64 4.4.5 Image Clustering 66 4.4.6 Structure-Aware Unsupervised Embeddings 68 4.4.7 Hyperbolic Distance to Filter Training Samples 71 4.4.8 Ablation Studies 74 4.5 Further Discussions 75 4.5.1 Connection to Contrastive Learning 75 4.5.2 Failure Cases of Hyperbolic Embedding Spaces 75 4.6 Summary 77 5 Contrastive Learning for Heterogeneous Graphs 78 5.1 Overview 78 5.2 Preliminaries 82 5.2.1 Meta-path 82 5.2.2 Representation Learning on Heterogeneous Graphs 82 5.2.3 Contrastive methods for Heterogeneous Graphs 83 5.3 Methodology 84 5.3.1 Definitions 84 5.3.2 Metanode-based Message Passing Layer 86 5.3.3 Contrastive Learning Framework 88 5.4 Experiments 89 5.4.1 Experimental Details 90 5.4.2 Node Classification 94 5.4.3 Node Clustering 96 5.4.4 Visualization 96 5.4.5 Effectiveness of Metanodes 97 5.5 Summary 99 6 Conclusions 101박

SNU Open Repository and Archive

A Graph-Based Semi-Supervised k Nearest-Neighbor Method for Nonlinear Manifold Distributed Data Classification

Author: Kasabov Nikola
Tu Enmei
Yang Jie
Zhang Yaqian
Zhu Lin
Publication venue
Publication date: 03/06/2016
Field of study

k

Nearest Neighbors (

k

NN) is one of the most widely used supervised learning algorithms to classify Gaussian distributed data, but it does not achieve good results when it is applied to nonlinear manifold distributed data, especially when a very limited amount of labeled samples are available. In this paper, we propose a new graph-based

k

NN algorithm which can effectively handle both Gaussian distributed data and nonlinear manifold distributed data. To achieve this goal, we first propose a constrained Tired Random Walk (TRW) by constructing an

R

-level nearest-neighbor strengthened tree over the graph, and then compute a TRW matrix for similarity measurement purposes. After this, the nearest neighbors are identified according to the TRW matrix and the class label of a query point is determined by the sum of all the TRW weights of its nearest neighbors. To deal with online situations, we also propose a new algorithm to handle sequential samples based a local neighborhood reconstruction. Comparison experiments are conducted on both synthetic data sets and real-world data sets to demonstrate the validity of the proposed new

k

NN algorithm and its improvements to other version of

k

NN algorithms. Given the widespread appearance of manifold structures in real-world problems and the popularity of the traditional

k

NN algorithm, the proposed manifold version

k

NN shows promising potential for classifying manifold-distributed data.Comment: 32 pages, 12 figures, 7 table

arXiv.org e-Print Archive

AUT Scholarly Commons