307 research outputs found
Deep Learning-Based Action Recognition
The classification of human action or behavior patterns is very important for analyzing situations in the field and maintaining social safety. This book focuses on recent research findings on recognizing human action patterns. Technology for the recognition of human action pattern includes the processing technology of human behavior data for learning, technology of expressing feature values ​​of images, technology of extracting spatiotemporal information of images, technology of recognizing human posture, and technology of gesture recognition. Research on these technologies has recently been conducted using general deep learning network modeling of artificial intelligence technology, and excellent research results have been included in this edition
Toward natural interaction in the real world: real-time gesture recognition
Using a new hand tracking technology capable of tracking 3D hand postures in real-time, we developed a recognition system for continuous natural gestures. By natural gestures, we mean those encountered in spontaneous interaction, rather than a set of artificial gestures chosen to simplify recognition. To date we have achieved 95.6% accuracy on isolated gesture recognition, and 73% recognition rate on continuous gesture recognition, with data from 3 users and twelve gesture classes. We connected our gesture recognition system to Google Earth, enabling real time gestural control of a 3D map. We describe the challenges of signal accuracy and signal interpretation presented by working in a real-world environment, and detail how we overcame them.National Science Foundation (U.S.) (award IIS-1018055)Pfizer Inc.Foxconn Technolog
ModDrop: adaptive multi-modal gesture recognition
We present a method for gesture detection and localisation based on
multi-scale and multi-modal deep learning. Each visual modality captures
spatial information at a particular spatial scale (such as motion of the upper
body or a hand), and the whole system operates at three temporal scales. Key to
our technique is a training strategy which exploits: i) careful initialization
of individual modalities; and ii) gradual fusion involving random dropping of
separate channels (dubbed ModDrop) for learning cross-modality correlations
while preserving uniqueness of each modality-specific representation. We
present experiments on the ChaLearn 2014 Looking at People Challenge gesture
recognition track, in which we placed first out of 17 teams. Fusing multiple
modalities at several spatial and temporal scales leads to a significant
increase in recognition rates, allowing the model to compensate for errors of
the individual classifiers as well as noise in the separate channels.
Futhermore, the proposed ModDrop training technique ensures robustness of the
classifier to missing signals in one or several channels to produce meaningful
predictions from any number of available modalities. In addition, we
demonstrate the applicability of the proposed fusion scheme to modalities of
arbitrary nature by experiments on the same dataset augmented with audio.Comment: 14 pages, 7 figure
A Non-Anatomical Graph Structure for isolated hand gesture separation in continuous gesture sequences
Continuous Hand Gesture Recognition (CHGR) has been extensively studied by
researchers in the last few decades. Recently, one model has been presented to
deal with the challenge of the boundary detection of isolated gestures in a
continuous gesture video [17]. To enhance the model performance and also
replace the handcrafted feature extractor in the presented model in [17], we
propose a GCN model and combine it with the stacked Bi-LSTM and Attention
modules to push the temporal information in the video stream. Considering the
breakthroughs of GCN models for skeleton modality, we propose a two-layer GCN
model to empower the 3D hand skeleton features. Finally, the class
probabilities of each isolated gesture are fed to the post-processing module,
borrowed from [17]. Furthermore, we replace the anatomical graph structure with
some non-anatomical graph structures. Due to the lack of a large dataset,
including both the continuous gesture sequences and the corresponding isolated
gestures, three public datasets in Dynamic Hand Gesture Recognition (DHGR),
RKS-PERSIANSIGN, and ASLVID, are used for evaluation. Experimental results show
the superiority of the proposed model in dealing with isolated gesture
boundaries detection in continuous gesture sequence
To Draw or Not to Draw: Recognizing Stroke-Hover Intent in Gesture-Free Bare-Hand Mid-Air Drawing Tasks
Over the past several decades, technological advancements have introduced new modes of communication
with the computers, introducing a shift from traditional mouse and keyboard interfaces.
While touch based interactions are abundantly being used today, latest developments in computer
vision, body tracking stereo cameras, and augmented and virtual reality have now enabled communicating
with the computers using spatial input in the physical 3D space. These techniques are now
being integrated into several design critical tasks like sketching, modeling, etc. through sophisticated
methodologies and use of specialized instrumented devices. One of the prime challenges in
design research is to make this spatial interaction with the computer as intuitive as possible for the
users.
Drawing curves in mid-air with fingers, is a fundamental task with applications to 3D sketching,
geometric modeling, handwriting recognition, and authentication. Sketching in general, is a
crucial mode for effective idea communication between designers. Mid-air curve input is typically
accomplished through instrumented controllers, specific hand postures, or pre-defined hand gestures,
in presence of depth and motion sensing cameras. The user may use any of these modalities
to express the intention to start or stop sketching. However, apart from suffering with issues like
lack of robustness, the use of such gestures, specific postures, or the necessity of instrumented
controllers for design specific tasks further result in an additional cognitive load on the user.
To address the problems associated with different mid-air curve input modalities, the presented
research discusses the design, development, and evaluation of data driven models for intent recognition
in non-instrumented, gesture-free, bare-hand mid-air drawing tasks.
The research is motivated by a behavioral study that demonstrates the need for such an approach
due to the lack of robustness and intuitiveness while using hand postures and instrumented
devices. The main objective is to study how users move during mid-air sketching, develop qualitative
insights regarding such movements, and consequently implement a computational approach to
determine when the user intends to draw in mid-air without the use of an explicit mechanism (such
as an instrumented controller or a specified hand-posture). By recording the user’s hand trajectory,
the idea is to simply classify this point as either hover or stroke. The resulting model allows for
the classification of points on the user’s spatial trajectory.
Drawing inspiration from the way users sketch in mid-air, this research first specifies the necessity
for an alternate approach for processing bare hand mid-air curves in a continuous fashion.
Further, this research presents a novel drawing intent recognition work flow for every recorded
drawing point, using three different approaches. We begin with recording mid-air drawing data
and developing a classification model based on the extracted geometric properties of the recorded
data. The main goal behind developing this model is to identify drawing intent from critical geometric
and temporal features. In the second approach, we explore the variations in prediction
quality of the model by improving the dimensionality of data used as mid-air curve input. Finally,
in the third approach, we seek to understand the drawing intention from mid-air curves using
sophisticated dimensionality reduction neural networks such as autoencoders. Finally, the broad
level implications of this research are discussed, with potential development areas in the design
and research of mid-air interactions
Gesture Recognition Using Hidden Markov Models Augmented with Active Difference Signatures
With the recent invention of depth sensors, human gesture recognition has gained significant interest in the fields of computer vision and human computer interaction. Robust gesture recognition is a difficult problem because of the spatiotemporal variations in gesture formation, subject size, subject location, image fidelity, and subject occlusion. Gesture boundary detection, or the automatic detection of the onset and offset of a gesture in a sequence of gestures, is critical toward achieving robust gesture recognition. Existing gesture recognition methods perform the task of gesture segmentation either using resting frames in a gesture sequence or by using additional information such as audio, depth images, or RGB images. This ancillary information introduces high latency in gesture segmentation and recognition, thus making it inappropriate for real time applications. This thesis proposes a novel method to recognize time-varying human gestures from continuous video streams. The proposed method passes skeleton joint information into a Hidden Markov Model augmented with active difference signatures to achieve state-of-the-art gesture segmentation and recognition.
Active body parts are used to calculate the likelihood of previously unseen data to facilitate gesture segmentation. Active difference signatures are used to describe temporal motion as well as static differences from a canonical resting position. Geometric features, such as joint angles, and joint topological distances are used along with active difference signatures as salient feature descriptors. These feature descriptors serve as unique signatures which identify hidden states in a Hidden Markov Model. The Hidden Markov Model is able to identify gestures in a robust fashion which is tolerant to spatiotemporal and human-to-human variation in gesture articulation.
The proposed method is evaluated on both isolated and continuous datasets. An accuracy of 80.7% is achieved on the isolated MSR3D dataset and a mean Jaccard index of 0.58 is achieved on the continuous ChaLearn dataset. Results improve upon existing gesture recognition methods, which achieve a Jaccard index of 0.43 on the ChaLearn dataset. Comprehensive experiments investigate the feature selection, parameter optimization, and algorithmic methods to help understand the contributions of the proposed method
- …