AI Core
927 items
Cohere proposes RCP-nDCG@10: a calibrated AI judge replaces fixed answer keys, matching human preference in 77% of 289 blind contests versus 52% for conventional nDCG
Cohere introduces RCP-nDCG@10, a retrieval evaluation methodology in which a calibrated AI judge grades every retrieved document against the same explicit relevance rubric and combines rubric answers with group-wise comparisons and calibration, so relevant documents absent from the answer key can still earn credit; in a blind study with 46 annotators, 273 queries and 289 head-to-head contests, reviewers sided with RCP-nDCG 70% of the time where the two metrics named different winners, and across all contests RCP-nDCG matched reviewer preference 77% of the time versus 52% for conventional nDCG, with the gain driven mainly by crediting relevant documents the answer key missed (19 percentage points).
Cohere launches Embed 5 with Pro at 85.8 average on ViDoRe V3 and Fast at 2.4x throughput sharing one embedding space
Cohere released the Embed 5 family of embedding models in two tiers, Pro and Fast, which share a single embedding space and support 128K context, text/image/fused inputs, 100+ languages, and Matryoshka outputs from 2048 down to 256 dimensions; Pro averages 85.8 on ViDoRe V3 (an 8.8 gain over Embed 4), scores 80.1 on FinanceBench and 84.8 across the parsed-document suite, while Fast averages 84.5 on ViDoRe V3 with about 2.4x the document throughput of Pro, priced at $0.12 and $0.08 per million tokens respectively.
Survey maps high-level synthesis for approximate computing around error estimation, approximation techniques, and design space exploration, and flags research gaps
Addressing the lack of a systematic survey and in-depth analysis of the latest methodologies in high-level synthesis for approximate computing (AHLS), this survey summarizes recent technologies in the field with particular focus on error estimation, approximation techniques, and design space exploration (DSE), and analyzes current research gaps, aiming to give researchers, engineers, and scholars a theoretical and practical framework for AHLS.
FedTrust-GNN reaches 94.2% accuracy at 10,000-100,000 participants and cuts label-flipping attack success from 34% to 6.1%
The work proposes FedTrust-GNN, a decentralized user-modeling framework that combines differentially private federated learning with secure multi-party computation, a permissioned blockchain using PBFT consensus, and a heterogeneous graph attention network (HGAT) that infers dynamic trust scores, with trust-weighted robust aggregation (TWRA, combining norm clipping and coordinate-wise median aggregation) providing Byzantine fault tolerance; on Federated EMNIST, Stack Overflow, and synthetic datasets with 10,000-100,000 participants it reports 94.2% accuracy (within 1.3% of centralized models), a reduction of label-flipping attack success from 34% to 6.1% (an 82% reduction), a 41% improvement in convergence stability, and blockchain performance of 1,200 TPS with 2.3-second finality.
DCONVNET splits 2D direction-of-arrival estimation into two 1D problems and estimates azimuth and elevation via the alternating direction method of multipliers
The work presents a fast two-dimensional direction-of-arrival (DOA) estimation approach for low-elevation targets of very-high-frequency array radar: it uses the azimuth and pitch angle uncoupling properties of a uniform planar array to turn the 2D angle estimation problem into two 1D DOA estimation problems, retrieves target information in the azimuth and elevation dimensions with digital beamforming, and then estimates azimuth and pitch angles using the alternating direction method of multipliers, thereby reducing complexity and eliminating the need for eigenvalue decomposition during operation.
1,003 people with depression rated psychotherapy with different levels of AI involvement: they preferred human therapists, willing to pay 31.6% less for assistive or collaborative AI and 57.2% less for fully autonomous AI
The study had 1,003 participants with depression read vignettes describing psychotherapy options with different levels of AI involvement (a human therapist without AI, assistive AI, collaborative AI, and fully autonomous AI) and rate them; participants consistently evaluated human therapists more favorably, reporting greater likelihood of seeking treatment, less hesitancy, and greater treatment acceptability, and compared with a human therapist they were willing to pay 31.6% less for therapists using assistive or collaborative AI and 57.
SCITFS folds adaptive redundancy penalization and bootstrap stability regularization into one objective, lifting SVM accuracy by 3.7% and Random Forest by 4.2% on eight benchmarks while cutting 92.3% of features.
The work proposes Stability-Constrained Information-Theoretic Feature Selection (SCITFS), which integrates conditional entropy, normalized mutual information maximization, and a stability term penalizing feature-ranking variance across bootstrap samples into a single formally defined objective, implemented via a greedy forward-selection strategy with proven monotonicity guarantees at O(B n^2 m d^4 + k n^2); across eight benchmark datasets and five classifiers (SVM, Random Forest, k-NN, XGBoost, Logistic Regression), SCITFS outperforms Information Gain, Mutual Information, mRMR, ReliefF, Fisher Score, and JMI, achieving a 3.7% average accuracy improvement on SVM and 4.2% on Random Forest with 92.3% feature reduction while maintaining performance; Friedman test (χ² = 127.4, p < 0.
Evo 2 shows in-context learning on five binary classification tasks with F1 up to 0.902 on short sequences, but collapses at kilobase scale and the 7B model beats the 40B
This study maps the in-context learning operating regime of Evo 2, a nucleotide-level foundation genomic language model, across five binary classification tasks spanning biological and artificial sequences, finding robust performance on shorter natural sequences (F1=0.902 for miRNA, 0.785 for Toxins), degradation with sequence length and collapse at kilobase scale, no benefit from model scaling (the 7B model systematically outperforms the 40B variant), poor prediction of accuracy by perplexity, and mechanistic interpretability via logit-lens and Jacobian Scope suggesting a prediction-generalisation trade-off and that models might track prompt structure rather than signal-carrying content.
Ridezy derives driver credibility in real time from edge AI and IoT sensors and anchors hashes and reputation updates on Polygon, outperforming rating-based, AI-only, and blockchain-only baselines in behavioural fidelity and trust guarantees
The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
Page 13 · showing 10