Machine Learning
Tutorial: Safe and Reliable Machine Learning (1904.07204v1)
Suchi Saria, Adarsh Subbaswamy
2019-04-15
This document serves as a brief overview of the "Safe and Reliable Machine Learning" tutorial given at the 2019 ACM Conference on Fairness, Accountability, and Transparency (FAT* 2019). The talk slides can be found here: https://bit.ly/2Gfsukp, while a video of the talk is available here:
, and a complete list of references for the tutorial here: https://bit.ly/2GdLPme.
A Discussion on Solving Partial Differential Equations using Neural Networks (1904.07200v1)
Tim Dockhorn
2019-04-15
Can neural networks learn to solve partial differential equations (PDEs)? We investigate this question for two (systems of) PDEs, namely, the Poisson equation and the steady Navier--Stokes equations. The contributions of this paper are five-fold. (1) Numerical experiments show that small neural networks (< 500 learnable parameters) are able to accurately learn complex solutions for systems of partial differential equations. (2) It investigates the influence of random weight initialization on the quality of the neural network approximate solution and demonstrates how one can take advantage of this non-determinism using ensemble learning. (3) It investigates the suitability of the loss function used in this work. (4) It studies the benefits and drawbacks of solving (systems of) PDEs with neural networks compared to classical numerical methods. (5) It proposes an exhaustive list of possible directions of future work.
Exact Rate-Distortion in Autoencoders via Echo Noise (1904.07199v1)
Rob Brekelmans, Daniel Moyer, Aram Galstyan, Greg Ver Steeg
2019-04-15
Compression is at the heart of effective representation learning. However, lossy compression is typically achieved through simple parametric models like Gaussian noise to preserve analytic tractability, and the limitations this imposes on learning are largely unexplored. Further, the Gaussian prior assumptions in models such as variational autoencoders (VAEs) provide only an upper bound on the compression rate in general. We introduce a new noise channel, Echo noise, that admits a simple, exact expression for mutual information for arbitrary input distributions. The noise is constructed in a data-driven fashion that does not require restrictive distributional assumptions. With its complex encoding mechanism and exact rate regularization, Echo leads to improved bounds on log-likelihood and dominates
-VAEs across the achievable range of rate-distortion trade-offs. Further, we show that Echo noise can outperform state-of-the-art flow methods without the need to train complex distributional transformations
Just Jump: Dynamic Neighborhood Aggregation in Graph Neural Networks (1904.04849v2)
Matthias Fey
2019-04-09
We propose a dynamic neighborhood aggregation (DNA) procedure guided by (multi-head) attention for representation learning on graphs. In contrast to current graph neural networks which follow a simple neighborhood aggregation scheme, our DNA procedure allows for a selective and node-adaptive aggregation of neighboring embeddings of potentially differing locality. In order to avoid overfitting, we propose to control the channel-wise connections between input and output by making use of grouped linear projections. In a number of transductive node-classification experiments, we demonstrate the effectiveness of our approach.
Testing Deep Neural Networks (1803.04792v4)
Youcheng Sun, Xiaowei Huang, Daniel Kroening, James Sharp, Matthew Hill, Rob Ashmore
2018-03-10
Deep neural networks (DNNs) have a wide range of applications, and software employing them must be thoroughly tested, especially in safety-critical domains. However, traditional software test coverage metrics cannot be applied directly to DNNs. In this paper, inspired by the MC/DC coverage criterion, we propose a family of four novel test criteria that are tailored to structural features of DNNs and their semantics. We validate the criteria by demonstrating that the generated test inputs guided via our proposed coverage criteria are able to capture undesired behaviours in a DNN. Test cases are generated using a symbolic approach and a gradient-based heuristic search. By comparing them with existing methods, we show that our criteria achieve a balance between their ability to find bugs (proxied using adversarial examples) and the computational cost of test case generation. Our experiments are conducted on state-of-the-art DNNs obtained using popular open source datasets, including MNIST, CIFAR-10 and ImageNet.
The Landscape of the Planted Clique Problem: Dense subgraphs and the Overlap Gap Property (1904.07174v1)
David Gamarnik, Ilias Zadik
2019-04-15
In this paper we study the computational-statistical gap of the planted clique problem, where a clique of size
is planted in an Erdos Renyi graph resulting in a graph . The goal is to recover the planted clique vertices by observing . It is known that the clique can be recovered as long as for any , but no polynomial-time algorithm is known for this task unless . Following a statistical-physics inspired point of view as an attempt to understand this computational-statistical gap, we study the landscape of the "sufficiently dense" subgraphs of and their overlap with the planted clique. Using the first moment method, we study the densest subgraph problems for subgraphs with fixed, but arbitrary, overlap size with the planted clique, and provide evidence of a phase transition for the presence of Overlap Gap Property (OGP) at . OGP is a concept introduced originally in spin glass theory and known to suggest algorithmic hardness when it appears. We establish the presence of OGP when is a small positive power of by using a conditional second moment method. As our main technical tool, we establish the first, to the best of our knowledge, concentration results for the -densest subgraph problem for the Erdos-Renyi model when for arbitrary . Finally, to study the OGP we employ a certain form of overparametrization, which is conceptually aligned with a large body of recent work in learning theory and optimization.
PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration (1901.09290v3)
Sangkug Lym, Esha Choukse, Siavash Zangeneh, Wei Wen, Sujay Sanghavi, Mattan Erez
2019-01-26
State-of-the-art convolutional neural networks (CNNs) used in vision applications have large models with numerous weights. Training these models is very compute- and memory-resource intensive. Much research has been done on pruning or compressing these models to reduce the cost of inference, but little work has addressed the costs of training. We focus precisely on accelerating training. We propose PruneTrain, a cost-efficient mechanism that gradually reduces the training cost during training. PruneTrain uses a structured group-lasso regularization approach that drives the training optimization toward both high accuracy and small weight values. Small weights can then be periodically removed by reconfiguring the network model to a smaller one. By using a structured-pruning approach and additional reconfiguration techniques we introduce, the pruned model can still be efficiently processed on a GPU accelerator. Overall, PruneTrain achieves a reduction of 31% in the end-to-end training time of modern CNNs by reducing computation cost by 37% in FLOPs, memory accesses by 35% for memory bandwidth bound layers, and the inter-accelerator communication by 54%.
Graph-Based Method for Anomaly Detection in Functional Brain Network using Variational Autoencoder (1904.07163v1)
Jalal Mirakhorli, Mojgan Mirakhorli
2019-04-15
Functional neuroimaging techniques have accelerated progress in the study of brain disorders and dysfunction, often measured using resting-state functional MRI (rs-fMRI). Because there are slight differences between healthy and disorder brains, investigating in the complex topology of human brain functional networks is difficult and complicated task with the growth of evaluation criteria. Meanwhile, graph theory and deep learning applications have recently spread widely to understanding human cognitive functions are linked to gene expression and related distributed spatial patterns. Irregular graph analysis has been widely applied in many domain, these applications might involve both node-centric and graph-centric tasks. In this work we explore Variational Autoencoder and Graph Convolutional Networks (GCNs) for the task of region of interest (ROI) brain identification Areas which do not have normal connection. For the learning of a powerful rigid graphs among large-scale data and underlying non-Euclidean structure, here used a framework of Graph Auto-Encoder (GAE) base on graph with hyper sphere distributer for functional of brain analyzing. we also obtain the correlation and possible modes between abnormal connections.
Are Nearby Neighbors Relatives?: Diagnosing Deep Music Embedding Spaces (1904.07154v1)
Jaehun Kim, Julián Urbano, Cynthia C. S. Liem, Alan Hanjalic
2019-04-15
Deep neural networks have frequently been used to directly learn representations useful for a given task from raw input data. In terms of overall performance metrics, machine learning solutions employing deep representations frequently have been reported to greatly outperform those using hand-crafted feature representations. At the same time, they may pick up on aspects that are predominant in the data, yet not actually meaningful or interpretable. In this paper, we therefore propose a systematic way to diagnose the trustworthiness of deep music representations, considering musical semantics. The underlying assumption is that in case a deep representation is to be trusted, distance consistency between known related points should be maintained both in the input audio space and corresponding latent deep space. We generate known related points through semantically meaningful transformations, both considering imperceptible and graver transformations. Then, we examine within- and between-space distance consistencies, both considering audio space and latent embedded space, the latter either being a result of a conventional feature extractor or a deep encoder. We illustrate how our method, as a complement to task-specific performance, provides interpretable insight into what a network may have captured from training data signals.
Copula-like Variational Inference (1904.07153v1)
Marcel Hirt, Petros Dellaportas, Alain Durmus
2019-04-15
This paper considers a new family of variational distributions motivated by Sklar's theorem. This family is based on new copula-like densities on the hypercube with non-uniform marginals which can be sampled efficiently, i.e. with a complexity linear in the dimension of state space. Then, the proposed variational densities that we suggest can be seen as arising from these copula-like densities used as base distributions on the hypercube with Gaussian quantile functions and sparse rotation matrices as normalizing flows. The latter correspond to a rotation of the marginals with complexity
. We provide some empirical evidence that such a variational family can also approximate non-Gaussian posteriors and can be beneficial compared to Gaussian approximations. Our method performs largely comparably to state-of-the-art variational approximations on standard regression and classification benchmarks for Bayesian Neural Networks.