A fundamental building block towards intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless Media Access Control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labelling process. In particular, a novel cross-modality labelling algorithm is proposed that jointly optimises the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labelled data to iteratively update wireless geofence models and curate device specific thresholds. Extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve 2-fold improvement in labelling compared with conventional methods and can achieve reliable speaker recognition in the wild

Chen, Changhao

Lu, Chris Xiaoxuan

Markham, Andrew

Trigoni, Niki

Xiangli, Yuanbo

Zhao, Peijun

English

A fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild

University of Liverpool Repository

Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor Data

Trigoni, N

Oxford University Research Archive

Autonomous learning of speaker identity and WiFi geofence from noisy sensor data

Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor Data

Abstract

Similar works

Full text

Available Versions

University of Liverpool Repository

Oxford University Research Archive

Oxford University Research Archive (ORA)