Skip to main content
Back to timeline
arXivSource publication:

Binary neural networks beat full-precision ResNet-18 on event-camera classification with fewer operations, reaching 90.58% on N-Caltech101

Synopsis

This work adapts modern binary neural networks (BNNs) to event-camera data, introduces the Polar-wise Binary Event Volume (PBEV) binary representation and binarizes the network input layer, and shows that with knowledge distillation and RGB cross-modal pretraining BNNs match or exceed full-precision ResNet-18 on N-Caltech101 and N-ImageNet-mini classification, with a best result of 90.58% at fewer operations.

Source-provided article image: Bringing BNNs to Fast Event Processing
Figure 1 ·

Figure 1 : Classification accuracy on N-Caltech101 [ 29 ] vs. Total Operations. Markers for input representations: Best non-bin. 2-Ch (SAE / EventCount), PEV6 (non-bin. 6-Ch), BEI (bin. 2-Ch), and PBEV6 (bin. 6-Ch). Colors for model architectures: blue for ReActNet, orange for BNext, green for ResNet-18, and red for A&B BNN.

arXiv

Interpretation

Introduces the Polar-wise Binary Event Volume (PBEV) representation: events are separated by polarity and stacked into a 3D tensor of short time slices, with each channel strictly encoding presence or absence of events at a position, so event data can be fed directly to binary networks. Earlier binary event representations such as BEI and BEHI binarize the input representation, but the subsequent neural processing is not necessarily fully binarized; PBEV combined with a binary input first layer extends binarization to the input layer, letting the convolutional core rely predominantly on XNOR–popcount operations. The paper gives a formal channel definition of PBEV and compares it against multiple representations and accumulation windows on N-Caltech101 and N-ImageNet-mini; PBEV6 is equivalent to SAE and EventCount and close to PEV6 for 10–50 ms windows.

Systematically benchmarks three state-of-the-art deep convolutional BNNs (ReActNet, A&B BNN, BNext Small) against a full-precision ResNet-18 baseline on event classification, and positions results against literature CNN, SNN, transformer and GCN results. Prior BNN work on neuromorphic vision was limited to low-power small tasks such as 3-class parking lot monitoring and denoising for pedestrian detection, with no systematic comparison to classical CNNs and SNNs on neuromorphic data. Evaluation spans two datasets, multiple event representations and accumulation times; when trained from scratch, distilled binary networks already match or exceed the full-precision baseline on both datasets, with BNext leading on both.

Finds that the best event representation depends primarily on the sensor's event density and accumulation regime rather than on architecture alone: on N-Caltech101 the 6-bin PEV6 that keeps timing and density is best (84.84% at 300 ms with BNext), whereas on N-ImageNet-mini the very short samples invert this ranking and the dense two-channel EventCount and SAE are strongest (47.86% and 43.62% with BNext). The inversion of representation ranking across datasets gives a basis for choosing input encoding according to event density and accumulation regime rather than fixing one encoding. Based on controlled comparisons across representations, backbones and accumulation windows on two datasets, with concrete numbers for PEV6 collapsing on N-ImageNet-mini (ResNet-18 falling to 28.14%).

RGB cross-modal pretraining (direct weight warm-start) benefits binary networks more than the full-precision baseline, and binary networks retain and exploit that initialization: from scratch the final weight signs are uncorrelated with random initialization, while after pretraining the C2I ratio rises above the scratch floor. Combines cross-modal pretraining with an initialization-retention analysis, showing that the RGB prior is retained and exploited even under 1-bit weights. Quantified by the correlation-to-initialization (C2I) ratio across architectures on both datasets; pretrained models surpass scratch training for every time window, representation and architecture, except A&B BNN with PBEV6.

Perspective

The result targets event-camera classification, evaluated on pan-tilt recordings of static scenes (N-Caltech101 and N-ImageNet-mini), with complexity estimated theoretically rather than measured on target hardware. For embedded or FPGA deployment that needs low-latency inference at short accumulation windows under constrained compute, PBEV and the binary input first layer offer a usable accuracy/complexity trade-off; the authors point to future extension to natural event recordings, object detection, and measured on-device latency and energy.

Readers should still watch: evaluation is limited to pan-tilt recordings of static scenes, so behavior on natural event recordings remains to be validated; complexity is estimated theoretically, without measured on-device latency or energy; A&B BNN collapses with EventCount at long windows (falling to 27.63% and below from 100 ms), and how training stability interacts with pretrained initialization needs further clarification; PBEV6 stops improving at long accumulation windows and carries too little information on the very short N-ImageNet-mini samples, so matching representation choice to event density remains an open question.

Sources