Skip to main content

AI Core

970 items

  1. arXiv

    DeepMind's SynthIDBio watermarks AI-designed proteins while binding viral, vascular and immune targets, but another design tool can scrub the tag

    A Google DeepMind team developed SynthIDBio, which weaves a statistical watermark into both the amino-acid sequence and the 3D shape of AI-designed proteins to mark their machine-generated origin without noticeably compromising function; the team reports that watermarked proteins bound targets involved in viral infection, blood-vessel formation and immune regulation as efficiently as unwatermarked ones, but the tag can in many cases be scrubbed by running a watermarked protein through another design tool, so it is framed as one layer in a layered biosecurity framework rather than a standalone solution.
  2. arXiv

    ByteDance team finds chunked KV-cache compression makes long-context retrieval periodically weak at the compression stride, with up to 40 percentage points between phases in DeepSeek-V4

    The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
  3. arXiv

    Fudan team trains Chinese-Jev on 10 million Chinese decisions, reaching 69.20% general accuracy and roughly 20x faster than Jev

    The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
  4. arXiv

    TabFM-Auto pairs an LLM agent with frozen TabFM to evolve data pipelines, lifting Elo from 1785 to 2013 across all 51 TabArena datasets

    TabFM-Auto pairs the frozen tabular foundation model TabFM with a language-model coding agent that iteratively rewrites four pipeline stages—data cleaning, feature engineering, context selection, and post-processing—guided by dataset metadata and validation feedback; across all 51 TabArena datasets five configurations take the top five overall positions, the best (Codex with Opus 5) raises TabFM from 1785.3 to 2013.0 Elo, the discovered pipelines transfer without further search to TabPFN-3, TabICLv2, and EXAONE-Tabular (+69 to +143 Elo), and TabFM-Auto ranks first overall among MLE agents on the 8 tabular competitions of MLE-Bench.
  5. arXiv

    EmoRES splits emotion vectors into shared and residual parts, lifting emotion hit rate by up to 12.95 points on IndexTTS-2 and CosyVoice2

    The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
  6. arXiv

    Google's 400M-parameter TabFM tops all 51 TabArena datasets zero-shot, and TabFM-Auto adds Elo with frozen weights

    TabFM is a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning, is pretrained entirely on synthetic tables generated from structural causal models, and produces calibrated zero-shot predictions in a single forward pass; across all 51 TabArena datasets (38 classification, 13 regression) zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines, while its two extensions, TabFM+ (multi-view feature expansion, ensembling, and post-hoc calibration) and TabFM-Auto (LLM-guided, dataset-specific data processing and feature engineering), improve both tracks over the same frozen weights.
  7. arXiv

    Tsinghua team turns the SIS proposal into a learning problem with GFlowNets: one network zero-shot matches or beats the post-hoc best of 31 analytic proposals on 1,190 unseen margins

    The work shows that the zero-variance sequential importance sampling (SIS) proposal for binary matrices with fixed margins is exactly the policy of a unit-reward GFlowNet, and proposes MarginFlow, a set transformer that reads the remaining margins and, trained on 1,904 margins, runs zero-shot on 1,190 held-out margins, matching or beating the post-hoc best of 31 analytically designed configurations on 1,187 of them with a median effective sample fraction of 99.8%.
  8. arXiv

    TGRL turns temperature differences into a training signal: +1.6% math average, +196.7 CodeForces rating, with no extra rollout budget

    The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
  9. arXiv

    SAKI routes teacher supervision through maximal-coupling accept/correct events, lifting Mean@8 and Pass@8 for both 1.7B and 0.6B students across seven math reasoning benchmarks

    SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
  10. arXiv

    AutoDataBench makes agents deliver training tasks one at a time: five frontier agents all score below 20 out of 100 at 45 minutes, with difficulty calibration rather than mode coverage as the binding constraint

    The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.

Page 20 · showing 10