Skip to main content
Back to timeline
arXivSource publication:

BlowLive combines blow-acoustic and facial biometrics into a multi-factor authenticator, reaching 100% accuracy on fused modalities from 50 participants with template and key revocation

Synopsis

The work proposes and implements BlowLive, a multi-factor biometric authentication framework that treats the acoustic signal of blowing onto a phone as a behavioral modality and a simultaneously captured face as a physiological modality, generating blow embeddings via GFCC features plus a CNN and facial embeddings via FaceNet, then deriving cryptographic keys from binarized fused embeddings through a fuzzy extractor, with an added Doppler-shift liveness detection module; on data collected from 50 participants it reports 99.56% accuracy for blow-acoustics and 100% for facial and fusion modalities, and 99.46% liveness detection accuracy under per-user thresholds.

Source-provided article image: BlowLive: Blow-Based Multi-Factor Biometrics with Liveness Detection and Revocability

Interpretation

The framework captures and fuses two modalities in a single blowing action: blow-acoustics (behavioral) and face (physiological), producing 128-dimensional blow embeddings via GFCC plus a CNN and 128-dimensional facial embeddings via FaceNet, then forming a unified representation through score-level weighted summation or feature-level concatenation. Compared with prior work by Halim et al. that also used blow acoustics and faces, this work introduces the finer-grained GFCC spectral feature pipeline and offers both score-level and feature-level fusion paths. Evaluated on data from 50 participants, each completing 10 to 12 sessions (half sitting, half standing, roughly 5 seconds each); feature-level fusion reports 100% accuracy with 0% FAR and 0% FRR at q=10, while score-level fusion reports 99.95%.

Authentication does not rely on raw biometric templates; a fuzzy extractor with BCH error-correcting codes converts the binarized fused embedding into a stable key K and a public helper string P, and authentication succeeds only if the Hamming distance falls within threshold so the same key can be reproduced. By shifting authentication from template comparison to key reproduction, the system stores no sensitive biometric features at the template level and thereby gains unlinkability, non-invertibility, and diversity properties. The paper presents unlinkability, non-invertibility, diversity, and revocability as design properties, and reports runtimes of 0.0224 ms for key generation, 1.0667 ms for key reproduction, and 0.1304 ms for key revocation.

Because blowing is a behavioral modality, users can adopt a new blowing pattern to obtain a fresh biometric identity, so the system supports revocation at both the key level and the template level. The paper notes that existing revocable or cancelable biometric schemes mainly revoke by changing keys or user-specific parameters while the underlying biometric trait stays fixed; here the behavioral modality lets the template itself be replaced. This conclusion rests on the framework design argument and the comparison table against related work rather than on a dedicated revocation attack experiment.

The liveness module has the phone continuously emit an approximately 20 kHz ultrasonic carrier and uses band-pass filtering, STFT peak-shift extraction, a Doppler envelope, and GFCC energy, combined into a weighted hybrid score to separate genuine blowing from playback. The paper argues that speech-based Doppler liveness methods rely on harmonics, lip motion, and vocal-tract dynamics and cannot be applied directly to blow signals that lack voiced components, motivating a pipeline specialized for airflow turbulence. Evaluated on 554 genuine and 558 playback samples: global threshold accuracy 90.38% (FAR 5.56%, FRR 13.72%), per-user dynamic threshold accuracy 99.46% (FAR 0%, FRR 1.08%).

Perspective

The result targets settings that use an ordinary smartphone microphone and front-facing camera in 1:1 authentication mode; the authors state explicitly that the system does not perform 1:N identification, so per-request computational complexity is independent of the number of enrolled users, and the server stores only one fixed-size helper per user. The intended users are those willing to complete acoustic and facial capture in a single blowing action, and the paper names phone unlocking, IoT device authentication, and secure access control as candidate deployments. Liveness thresholds come in two flavors: a global threshold modeling population-level genuine behavior, and a per-user threshold for cases where blowing strength varies substantially across individuals, with the measurements favoring the latter.

The evaluation is conducted offline, so end-to-end online deployment latency and stability are not yet reported. The liveness comparison covers genuine blowing versus playback, while the threat model also lists stolen-sample and AI-generated blowing attacks that do not have their own result rows in the experimental tables. The dataset comes from 50 participants with 10 to 12 sessions each, leaving cross-device, cross-noise, and broader population behavior as open questions. In addition, the body text reports liveness detection accuracy of 99.42% while Table 1 lists 99.46% for the per-user dynamic threshold, so readers citing the number should note this discrepancy.

Sources