Search CORE

60 research outputs found

Divergent estimation error in portfolio optimization and in linear regression

Author: B. Scherer
E.J. Elton
H. Markovitz
I. Kondor
I. Kondor
I. Varga-Haszonits
I. Varga-Haszonits
R. Jagannathan
S. Ciliberti
S. Kirkpatrick
S. Pafka
S. Pafka
Z. Burda
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 09/10/2007
Field of study

The problem of estimation error in portfolio optimization is discussed, in the limit where the portfolio size N and the sample size T go to infinity such that their ratio is fixed. The estimation error strongly depends on the ratio N/T and diverges for a critical value of this parameter. This divergence is the manifestation of an algorithmic phase transition, it is accompanied by a number of critical phenomena, and displays universality. As the structure of a large number of multidimensional regression and modelling problems is very similar to portfolio optimization, the scope of the above observations extends far beyond finance, and covers a large number of problems in operations research, machine learning, bioinformatics, medical science, economics, and technology.Comment: 5 pages, 2 figures, Statphys 23 Conference Proceedin

arXiv.org e-Print Archive

Crossref

EDP Sciences OAI-PMH repository (1.2.0)

Research Papers in Economics

Regularizing Portfolio Optimization

Author: Acerbi C
Acerbi C Nordio C Sirtori C
Bengio Y
Bertsekas D P
Bordes A
Bottou L
Bouchaud J-Ph
Burda Z
Chopra V K
DeMiguel V
Elton E J
Embrechts P
Frahm G
Frahm G Memmel Ch
Gulyas N Kondor I
Imre Kondor
Jobson J D
Jorion P
Kempf A
Kondor I Varga-Haszonits I
Macrae R
Markowitz H
Morgan J P Reuters Riskmetrics
Perez-Cruz F
Potters M
Rockafellar R T
Schölkopf B
Schölkopf B
Susanne Still
Tibshirani R
Vanderbei R J
Vapnik V
Vapnik V
Vapnik V
Varga-Haszonits I
Publication venue: 'IOP Publishing'
Publication date: 09/11/2009
Field of study

The optimization of large portfolios displays an inherent instability to estimation error. This poses a fundamental problem, because solutions that are not stable under sample fluctuations may look optimal for a given sample, but are, in effect, very far from optimal with respect to the average risk. In this paper, we approach the problem from the point of view of statistical learning theory. The occurrence of the instability is intimately related to over-fitting which can be avoided using known regularization methods. We show how regularized portfolio optimization with the expected shortfall as a risk measure is related to support vector regression. The budget constraint dictates a modification. We present the resulting optimization problem and discuss the solution. The L2 norm of the weight vector is used as a regularizer, which corresponds to a diversification "pressure". This means that diversification, besides counteracting downward fluctuations in some assets by upward fluctuations in others, is also crucial because it improves the stability of the solution. The approach we provide here allows for the simultaneous treatment of optimization and diversification in one framework that enables the investor to trade-off between the two, depending on the size of the available data set

arXiv.org e-Print Archive

Crossref

ELTE Digital Institutional Repository (EDIT)

Replica theory for learning curves for Gaussian processes on random graphs

Author: Chung F K
Erdős P
Font-Clos F Massucci F A Castillo I P
Kondor R Lafferty J Sammut C Hoffmann A G
Kühn R
Kühn R
M J Urry
Malzahn D
Min R Kuang R Bonner A Zhang Z Park H Parthasarathy S Liu H Obradovic Z
Mézard M
Opper M
Opper M
P Sollich
Rasmussen C E
Rogers T
Sollich P
Sollich P
Sollich P
Sollich P
Urry M J
Urry M J Sollich P
Publication venue: 'IOP Publishing'
Publication date: 26/10/2012
Field of study

Statistical physics approaches can be used to derive accurate predictions for the performance of inference methods learning from potentially noisy data, as quantified by the learning curve defined as the average error versus number of training examples. We analyse a challenging problem in the area of non-parametric inference where an effectively infinite number of parameters has to be learned, specifically Gaussian process regression. When the inputs are vertices on a random graph and the outputs noisy function values, we show that replica techniques can be used to obtain exact performance predictions in the limit of large graphs. The covariance of the Gaussian process prior is defined by a random walk kernel, the discrete analogue of squared exponential kernels on continuous spaces. Conventionally this kernel is normalised only globally, so that the prior variance can differ between vertices; as a more principled alternative we consider local normalisation, where the prior variance is uniform

arXiv.org e-Print Archive

Crossref

King's Research Portal

Atomic-scale representation and statistical learning of tensorial properties

Author: Alred J. M.
Bartók A. P.
Bartók A. P.
Bartók A. P.
Behler J.
Bereau T.
Bereau T.
Braams B. J.
Brockherde F.
Calderon C. E.
Ceriotti M.
Chandrasekaran A.
Christensen A. S.
De S.
Glielmo A.
Glielmo A.
Grisafi A.
Grisafi A.
Hättig C.
Jain A.
Kaufmann K.
Kondor R.
Li Z.
Liang C.
Musil F.
Shapeev A.
Stone A. J.
Ward L.
Weinert U.
Wilkins D. M.
Willatt M. J.
Willatt M. J.
Williams C. K. I.
Yuan Y.
Zhang L.
Publication venue
Publication date: 02/04/2019
Field of study

This chapter discusses the importance of incorporating three-dimensional symmetries in the context of statistical learning models geared towards the interpolation of the tensorial properties of atomic-scale structures. We focus on Gaussian process regression, and in particular on the construction of structural representations, and the associated kernel functions, that are endowed with the geometric covariance properties compatible with those of the learning targets. We summarize the general formulation of such a symmetry-adapted Gaussian process regression model, and how it can be implemented based on a scheme that generalizes the popular smooth overlap of atomic positions representation. We give examples of the performance of this framework when learning the polarizability and the ground-state electron density of a molecule

arXiv.org e-Print Archive

Infoscience - École polytechnique fédérale de Lausanne

Crossref

Security analyst networks, performance and career outcomes

Author: A Shleifer
C Goldin
D Duffie
D Gromb
D Gromb
Guillaume Plantin
Igor Makarov
J Anton
J Dow
J Kondo
K Arrow
M Brunnermeier
M Brunnermeier
M Friedman
N Barberis
O Lamont
P Kondor
S Bhattacharya
S N Kaplan
T Philippon
Z He
Publication venue: 'Wiley'
Publication date: 01/01/2012
Field of study

Authors' draft. Final version to be published in The Journal of Finance. Available online at http://onlinelibrary.wiley.com/Using a sample of 42,376 board directors and 10,508 security analysts we construct a social network, mapping the connections between analysts and directors, between directors, and between analysts. We use social capital theory and techniques developed in social network analysis to measure the analyst’s level of connectedness and investigate whether these connections provide any information advantage to the analyst. We find that better-connected (better-networked) analysts make more accurate, timely, and bold forecasts. Moreover, analysts with better network positions are less likely to lose their job, suggesting that these analysts are more valuable to their brokerage houses. We do not find evidence that analyst innate forecasting ability predicts an analyst’s future network position. In contrast, past forecast optimism has a positive association with building a better network of connections

Crossref

LSE Research Online

Toulouse Capitole Publications

HAL Descartes

Open Research Exeter

Toulouse 1 Capitole Publications

Hal-Diderot

SPIRE - Sciences Po Institutional REpository

A graph-search framework for associating gene identifiers with documents

Author: A Yeh
AM Cohen
AM Cohen
AM Cohen
C Zhai
Consortium TGO
D Hanisch
E Hatcher
E Minkov
E Minkov
Einat Minkov
F Sha
J Crim
K Franzén
K Fundel
K Humphreys
L Hirschman
L Hirschman
M Collins
M Craven
R Bunescu
RI Kondor
T Rindflesch
U Leser
William W Cohen
WW Cohen
WW Cohen
WW Cohen
Y Altun
Y Freund
Z Kou
Publication venue: BioMed Central
Publication date: 01/01/2006
Field of study

BACKGROUND: One step in the model organism database curation process is to find, for each article, the identifier of every gene discussed in the article. We consider a relaxation of this problem suitable for semi-automated systems, in which each article is associated with a ranked list of possible gene identifiers, and experimentally compare methods for solving this geneId ranking problem. In addition to baseline approaches based on combining named entity recognition (NER) systems with a "soft dictionary" of gene synonyms, we evaluate a graph-based method which combines the outputs of multiple NER systems, as well as other sources of information, and a learning method for reranking the output of the graph-based method. RESULTS: We show that named entity recognition (NER) systems with similar F-measure performance can have significantly different performance when used with a soft dictionary for geneId-ranking. The graph-based approach can outperform any of its component NER systems, even without learning, and learning can further improve the performance of the graph-based ranking approach. CONCLUSION: The utility of a named entity recognition (NER) system for geneId-finding may not be accurately predicted by its entity-level F1 performance, the most common performance measure. GeneId-ranking systems are best implemented by combining several NER systems. With appropriate combination methods, usefully accurate geneId-ranking systems can be constructed based on easily-available resources, without resorting to problem-specific, engineered components

Crossref

Springer - Publisher Connector

Directory of Open Access Journals

PubMed Central

Classification of heterogeneous microarray data by maximum entropy kernel

Author: AI Su
AI Su
B Nilsson
B Rosner
B Scholköpf
B Schölkopf
DR Rhodes
GA Torunera
H Liu
H Lodhi
H Saigo
I Yanai
J Okutsu
JE Staunton
JM Boer
K Tsuda
K Tsuda
L Liu
LJ van't Veer
ME Wall
N Cristianini
O Alter
R Kondor
RK O'Donnell
RW Tothill
S Ramaswamy
SM Flechnera
T Kato
TR Golub
Tsuyoshi Kato
V Vapnik
Wataru Fujibuchi
Z Liu
Publication venue: BioMed Central
Publication date: 01/01/2007
Field of study

Abstract Background There is a large amount of microarray data accumulating in public databases, providing various data waiting to be analyzed jointly. Powerful kernel-based methods are commonly used in microarray analyses with support vector machines (SVMs) to approach a wide range of classification problems. However, the standard vectorial data kernel family (linear, RBF, etc.) that takes vectorial data as input, often fails in prediction if the data come from different platforms or laboratories, due to the low gene overlaps or consistencies between the different datasets. Results We introduce a new type of kernel called maximum entropy (ME) kernel, which has no pre-defined function but is generated by kernel entropy maximization with sample distance matrices as constraints, into the field of SVM classification of microarray data. We assessed the performance of the ME kernel with three different data: heterogeneous kidney carcinoma, noise-introduced leukemia, and heterogeneous oral cavity carcinoma metastasis data. The results clearly show that the ME kernel is very robust for heterogeneous data containing missing values and high-noise, and gives higher prediction accuracies than the standard kernels, namely, linear, polynomial and RBF. Conclusion The results demonstrate its utility in effectively analyzing promiscuous microarray data of rare specimens, e.g., minor diseases or species, that present difficulty in compiling homogeneous data in a single laboratory.</p

Crossref

Springer - Publisher Connector

Directory of Open Access Journals

PubMed Central

Bayesian Markov Random Field Analysis for Protein Function Prediction Based on Network Data

Author: A Kuzniar
A Vazquez
Aalt D. J. van Dijk
AJ Enright
C Moler
Cajo J. F. ter Braak
CJF Ter Braak
CJF Ter Braak
CM Federovitch
DJC MacKay
GD Bader
GR Lanckriet
H Lee
I Kosmidis
I Ulitsky
Iddo Friedberg
IM Cheeseman
J Besag
JA Hanley
L Milligan
L Peña Castillo
M Ashburner
M Deng
M Deng
M Punta
Marco C. A. M. Bink
N Nariai
NJ Mulder
P McCullagh
R Sharan
RI Kondor
Roeland C. H. J. van Ham
S Ferré
S Geman
S Letovsky
S Mostafavi
SF Altschul
SR Collins
SZ Li
T Gabaldon
U Karaoz
V Vethantham
XL Chen
Y Chen
Y Guan
Yiannis A. I. Kourmpetis
Z Barutcuoglu
Z Wei
Publication venue: Public Library of Science
Publication date: 01/01/2010
Field of study

Inference of protein functions is one of the most important aims of modern biology. To fully exploit the large volumes of genomic data typically produced in modern-day genomic experiments, automated computational methods for protein function prediction are urgently needed. Established methods use sequence or structure similarity to infer functions but those types of data do not suffice to determine the biological context in which proteins act. Current high-throughput biological experiments produce large amounts of data on the interactions between proteins. Such data can be used to infer interaction networks and to predict the biological process that the protein is involved in. Here, we develop a probabilistic approach for protein function prediction using network data, such as protein-protein interaction measurements. We take a Bayesian approach to an existing Markov Random Field method by performing simultaneous estimation of the model parameters and prediction of protein functions. We use an adaptive Markov Chain Monte Carlo algorithm that leads to more accurate parameter estimates and consequently to improved prediction performance compared to the standard Markov Random Fields method. We tested our method using a high quality S.cereviciae validation network with 1622 proteins against 90 Gene Ontology terms of different levels of abstraction. Compared to three other protein function prediction methods, our approach shows very good prediction performance. Our method can be directly applied to protein-protein interaction or coexpression networks, but also can be extended to use multiple data sources. We apply our method to physical protein interaction data from S. cerevisiae and provide novel predictions, using 340 Gene Ontology terms, for 1170 unannotated proteins and we evaluate the predictions using the available literature

Public Library of Science (PLOS)

Crossref

Directory of Open Access Journals

PubMed Central

Wageningen University & Research Publications

Disease-Aging Network Reveals Significant Roles of Aging Genes in Connecting Genetic Diseases

Author: A Budovsky
A Budovsky
A Friedman
A Kowald
A Kriete
A Ozgur
AL Barabasi
C Soti
D Harman
David B. Searls
DJ Watts
E Ravasz
G Jin
GRG Lanckriet
H Kitano
H Xue
HD Osiewacz
HJ Kiss
I Feldman
J Hasty
JDJ Han
Jiguang Wang
JP de Magalhaes
JP de Magalhaes
JR Managbanag
KI Goh
L Hayflick
Luonan Chen
M Wolfson
MEJ Newman
P Shannon
P Zuppan
PF Jonsson
Q Cui
R Albert
R Bell
RI Kondor
S Karni
S Maere
S Maslov
S Peri
S Vasto
Shihua Zhang
T Ideker
T Ishunina
TBL Kirkwood
U Brandes
U Stelzl
X Jiang
X Wu
Xiang-Sun Zhang
Y Li
Yong Wang
Z Spiro
Z Tu
Publication venue: Public Library of Science
Publication date: 01/09/2009
Field of study

One of the challenging problems in biology and medicine is exploring the underlying mechanisms of genetic diseases. Recent studies suggest that the relationship between genetic diseases and the aging process is important in understanding the molecular mechanisms of complex diseases. Although some intricate associations have been investigated for a long time, the studies are still in their early stages. In this paper, we construct a human disease-aging network to study the relationship among aging genes and genetic disease genes. Specifically, we integrate human protein-protein interactions (PPIs), disease-gene associations, aging-gene associations, and physiological system–based genetic disease classification information in a single graph-theoretic framework and find that (1) human disease genes are much closer to aging genes than expected by chance; and (2) diseases can be categorized into two types according to their relationships with aging. Type I diseases have their genes significantly close to aging genes, while type II diseases do not. Furthermore, we examine the topological characters of the disease-aging network from a systems perspective. Theoretical results reveal that the genes of type I diseases are in a central position of a PPI network while type II are not; (3) more importantly, we define an asymmetric closeness based on the PPI network to describe relationships between diseases, and find that aging genes make a significant contribution to associations among diseases, especially among type I diseases. In conclusion, the network-based study provides not only evidence for the intricate relationship between the aging process and genetic diseases, but also biological implications for prying into the nature of human diseases

Public Library of Science (PLOS)

Crossref

Directory of Open Access Journals

PubMed Central

Candidate gene prioritization by network analysis of differential expression using machine learning approaches

Author: A Subramanian
A Zanzoni
AJ Smola
AP Francisco
B Aranda
B Harr
Bart de Moor
C Saunders
C Stark
C von Mering
D Nitsch
D Zieker
Daniela Nitsch
F Chung
F Fouss
Fabian Ojeda
GC Cawley
GD Bader
H Yang
HY Chuang
J Chen
JA Hanley
Joana P Gonçalves
JW Park
K Lage
KR Brown
L Franke
L Gautier
L Salwinski
LC Tranchevent
M Liu
P Baldi
P Pagel
R Gupta
RA Irizarry
RI Kondor
RK Nibbe
S Aerts
S Köhler
S Mirkin
S Razick
S Vardhanabhuti
SE Choe
T Fawcett
WK Lim
Y Saad
Yves Moreau
Z Wu
Publication venue: BioMed Central
Publication date: 01/01/2010
Field of study

Abstract Background Discovering novel disease genes is still challenging for diseases for which no prior knowledge - such as known disease genes or disease-related pathways - is available. Performing genetic studies frequently results in large lists of candidate genes of which only few can be followed up for further investigation. We have recently developed a computational method for constitutional genetic disorders that identifies the most promising candidate genes by replacing prior knowledge by experimental data of differential gene expression between affected and healthy individuals. To improve the performance of our prioritization strategy, we have extended our previous work by applying different machine learning approaches that identify promising candidate genes by determining whether a gene is surrounded by highly differentially expressed genes in a functional association or protein-protein interaction network. Results We have proposed three strategies scoring disease candidate genes relying on network-based machine learning approaches, such as kernel ridge regression, heat kernel, and Arnoldi kernel approximation. For comparison purposes, a local measure based on the expression of the direct neighbors is also computed. We have benchmarked these strategies on 40 publicly available knockout experiments in mice, and performance was assessed against results obtained using a standard procedure in genetics that ranks candidate genes based solely on their differential expression levels (<it>Simple Expression Ranking</it>). Our results showed that our four strategies could outperform this standard procedure and that the best results were obtained using the <it>Heat Kernel Diffusion Ranking </it>leading to an average ranking position of 8 out of 100 genes, an AUC value of 92.3% and an error reduction of 52.8% relative to the standard procedure approach which ranked the knockout gene on average at position 17 with an AUC value of 83.7%. Conclusion In this study we could identify promising candidate genes using network based machine learning approaches even if no knowledge is available about the disease or phenotype.</p

Crossref

Springer - Publisher Connector

Directory of Open Access Journals

PubMed Central