TAPe+ML v3 reaches 84.7 mAP50 and 65.3 mAP50-95 on COCO detection with under 100K parameters, and lifts 30-image industrial ore-blockage accuracy from about 28% with YOLO 26s to about 65%
Synopsis
The work presents TAPe+ML v3, which replaces raw pixel tensors with a structured TAPe (Theory of Active Perception) representation and combines background, pointer, and prototype-clustering submodels under a coordinator; with fewer than 100,000 parameters it reports 84.7 mAP50 and 65.3 mAP50-95 on COCO detection, 80.7 mask mAP50 and 58.4 mask mAP50-95 on COCO instance segmentation, 88.1% ImageNet-1k Top-1 and 89.9% ImageNet-Real, 92% versus 47% on Imagenette under identical training with TAPe versus raw pixels, and an industrial ore-blockage pilot where backbone adaptation outperforms head-only adaptation.
Interpretation
The system builds recognition on a structured TAPe representation rather than raw pixel tensors: an image is mapped to a set of T-bits and a relational graph, with the number of T-bits orders of magnitude smaller than the number of pixels and the index occupying orders of magnitude less memory. Unlike learned tokenizers such as VQ-VAE and BEiT, TAPe is positioned as a perception layer before recognition rather than a component inside a large pretraining pipeline; unlike pixel-input detectors such as YOLO and RF-DETR, it shifts part of the modeling burden from network parameters to the input representation. The paper gives a formal description of the TAPe mapping (I↦G(I)=({t1,…,tk},E)) and representational properties (invariance to local transformations, semantic stability, conciseness and comparability), while stating that the full algorithm is proprietary and can only be described at the level of input, output, and representational properties.
On COCO detection, TAPe+ML v3 reports mAP50=84.7% and mAP50-95=65.3% with fewer than 100,000 total parameters; on COCO instance segmentation it reports mask mAP50=80.7 and mask mAP50-95=58.4. The paper places these results alongside public baselines including RF-DETR 2XL (about 126.9M parameters, mAP50-95 about 60.1) and YOLO11-M (about 20.1M parameters, mAP50-95 48.6), noting that external detector results come from official benchmark tables and may differ in training and inference protocol. Detection is evaluated on standard COCO train2017/val2017 with 80 classes using mAP50 and mAP50-95 computed with pycocotools; latency is measured on an NVIDIA GTX 1070 Ti 8GB at batch size 1 with a 1024-pixel long-side cap, giving 10.8 ms per frame for the detection stage.
Classification experiments show a large effect from the representation itself: on Imagenette, with the same 3-layer CNN (about 516K parameters), the same 10% of data, and no augmentation, TAPe input reaches about 92% validation accuracy versus about 47% for the raw-pixel baseline; ImageNet-1k Top-1 is 88.1% and ImageNet-Real is 89.9%. The paper emphasizes that architecture, parameter count, data volume, and training regime are identical in this comparison, with input representation as the only difference; it also reports ImageNet-1k training from scratch and distribution-shift results including ObjectNet 78.6, ImageNet v2 80.2, ImageNet-R 90.3, and ImageNet-Sketch 75.8. ImageNet-1k follows the standard train/validation split and preprocessing; distribution-shift benchmarks are listed alongside public numbers for ResNet-50, ConvNeXt-B, DINOv2 ViT-B/14, and DINOv3 ViT-7B/16, with TAPe+ML v3 at about 0.1M parameters and about 0.1 GFLOPs.
In the industrial ore-blockage pilot, with 30 images the YOLO 26s baseline reaches about 28% accuracy, TAPe+ML head-only about 45%, and backbone adaptation about 65%; with 85 images backbone adaptation reaches about 85%, and with 500 images plus NMS it reaches 85.7–96%. The paper attributes this gap to the combination of structured representation and backbone adaptation, and notes that the gain from head to backbone adaptation is larger under distribution shift on RF100-VL (45.6→52.7→58.5) than in the standard setting (64.1→68.2→72.4), indicating that TAPe+ML v3 delivers its largest benefit under distribution shift. The pilot is a two-class industrial task of detecting ore blockages on conveyor-belt images, with accuracy defined as a binary detection designation of detections above IoU≥0.5 with correct Top-1 classification; the backbone-adaptation protocol uses 15 epochs, batch size 64, OneCycleLR with peak learning rate 1e-4, and early stopping after 3 consecutive epochs without improvement after epoch 5.
Perspective
The work is aimed at readers who need classification, detection, and segmentation under limited data, memory, and compute, especially in industrial quality control and edge deployment; the adaptation modes (head-only, backbone adaptation, from-scratch) and the auto-labeling workflow offer an operational path to building a system in a new domain with only dozens of labeled images. The video scene-detection experiment reports indexing one hour of video in about 10–11 seconds with an index under 1 MB and clustering in about 1 second, suggesting the representation can also serve as an input layer for video retrieval and scene organization.
The paper itself notes three scope limits: tight box localization for small objects remains the primary operational limitation, since the strictest IoU range requires very precise alignment between predicted and ground-truth boxes; the backbone carries a COCO bias when transferring to out-of-distribution tasks, and full backbone adaptation requires an additional dataset and a separate compute budget; and the full mathematical and algorithmic specification of TAPe is proprietary and not publicly disclosed. In addition, the paper does not address 3D vision or full video detection including tracking and temporal consistency, so conclusions should not be automatically extended to settings where inter-frame consistency is decisive. Segmentation coverage and AP@75 use the LVIS methodology, which is not the same metric as the COCO Mask AP used for RF-DETR-Seg, Mask DINO, and Co-DETR, so direct comparison requires care.
