Skip to main content
Back to timeline
arXivSource publication:

Yandex team's AsyncLLM lets Qwen 3.x models watch, think and act concurrently without fine-tuning, speeding up streaming video, games and system monitoring

Synopsis

The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.

AI-generated editorial illustration: LLMs are General Asynchronous Agents

Interpretation

It proposes AsyncLLM, which places concurrency inside inference rather than around API calls: an agent is a set of Python async/await coroutines, each writing to its own CacheBlock and choosing which blocks to attend to through cache views, so streams can read each other's evolving, possibly incomplete state. Previously, real-time voice, streaming video, VLA robot control and asynchronous tool calling each handled concurrency with task-specific architectures or training; this work unifies them under one programming model and distills five recurring concurrency patterns: probes and monitors, parallel inference streams, subroutines and delegation, interruptions and incremental inputs, and shared evolving context. The paper provides a framework definition, a formal description of CacheBlocks and cache views, an open-source reference implementation built on Mini-SGLang, and validation across three task domains with the same Qwen 3.x model family, making this a framework-level design with multi-domain empirical support.

It describes parallel GPU inference algorithms supporting hybrid and multimodal LLMs: full-attention layers extend query-rotation (RoPE) manipulation, linear attention and Gated DeltaNet layers are reformulated from per-token to per-block affine transitions composed in block order, and MRoPE maintains per-block position spans instead of rotating by token count. Prior concurrent-attention memory sharing targeted mainly full-attention RoPE models; this work generalizes block composition to Gated DeltaNet and KDA-style linear attention variants and to multimodal rotary embeddings, so modern hybrid MLLMs can use arbitrary cache views. The paper derives block composition and its order-independence in equations 1 through 10, discusses compatibility with sliding-window, GQA/MQA, MLA, sparse attention, ALiBi and NoPE variants in Appendix D, and reports throughput and latency measurements under synthetic workloads in Appendix E.

For streaming video understanding, the AsyncLLM agent comprises five concurrent components (event probe, background thinking thread, output probe, description writer and a non-LLM output compiler) and uses training-free probes that run the base LLM as-is with a pre-filled prompt in a single forward pass; on SoccerNet-Caption and ProactiveVideoQA, Qwen3.5+ models outperform the purpose-trained Mage-VL. Prior streaming video approaches typically train lightweight probes or specialized models; this work shows a generalist MLLM inside AsyncLLM can handle event detection and description generation without task-specific training, reusing the base model itself as the probe. The paper reports event-probe ROC AUC, TriggerAcc and TimVal on SoccerNet and Proactive AUC on ProactiveVideoQA, along with a hyperparameter sweep for Mage-VL; the authors note the official SoccerNet protocol does not measure description quality, with the description advantage appearing on ProactiveVideoQA.

On ViZDoom games and DevOps-Gym system monitoring, asynchronous agents respond much faster while preserving reasoning gains: in monitoring, AsyncLLM reaches 55.88% accuracy using about 3837 forward passes versus 61.76% with 8452 forward passes for the sequential baseline, and in games the AsyncLLM agent reacts much faster than a sequential agent based on the same model. Prior system-monitoring and game agents processed logs or frames sequentially; this work transfers the streaming-video agent architecture to log streams and interactive games and measures response timeliness in addition to accuracy. The paper reports accuracy and forward-pass counts on 34 DevOps-Gym monitoring tasks and compares against sequential and probe-only baselines on two ViZDoom scenarios (HealthGathering and DeadlyCorridor, frame skip 4); the authors state the game results demonstrate basic capability rather than state-of-the-art performance for those environments.

Perspective

The framework targets deployments that must accept new inputs while inference is in progress: real-time voice and video assistants, embodied and interactive agents, and system or log monitoring. It suits developers who want to build asynchronous agents quickly from existing open models rather than collecting data and fine-tuning for each concurrency scenario; the paper also suggests passing agent definitions through a modified realtime API or letting the model write its own agent definition. Experiments cover Qwen 3.x models and three task domains, and the authors position the results as a feasibility demonstration of general asynchrony rather than best-in-class performance on each task.

The paper states that letting agents define their own coroutines is 'not yet reliable', working well on HealthGathering but worse on DeadlyCorridor, so this direction needs further study. Concurrent execution may let an agent emit an incorrect or harmful action before slower reasoning or monitoring coroutines can intervene, and the authors recommend sandboxing, output validation and monitoring in sensitive deployments. The reference implementation is described as minimal and does not include every possible code optimization, and the experiments do not use speculative decoding or compression, so the efficiency numbers reflect this implementation rather than an upper bound. In addition, this evidence bundle is a full-text parse in which some table and figure values appear as text; exact reproduction of experimental settings would still require the original appendices.

Sources