Skip to main content

Engineering Sciences

140 items

  1. arXiv

    CUA-SWE tested 105 executable tasks and found that giving agents the running interface lifted GPT-6-Astra's four-domain mean from 11.3% to 59.9%

    The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
  2. arXiv

    DyRAD renders full range-azimuth-Doppler radar tensors of dynamic driving scenes through a fixed point-spread function, raising radar detection recovery on RADIal from 26.9% to 90.7% of reference-detected objects

    DyRAD represents dynamic driving scenes as static background point reflectors plus motion-tracked dynamic point reflectors and renders complete range-azimuth-Doppler (RAD) tensors through a fixed analytic point-spread function derived from the radar signal-processing chain, with Doppler serving both as a rendered output and as supervision for object tracks, evaluated on RADIal, Boreas, and a synthetic benchmark at both on-path poses and displaced viewpoints, raising radar detection recovery on RADIal from 26.9% for the strongest baseline to 90.7% of reference-detected objects.
  3. arXiv

    RL post-training of vision-language-action models concentrates parameter updates in Timestep Modules that hold only 14.83%–27.58% of the action expert, and their low-rank directions predict and improve task success

    This work systematically analyzes how reinforcement learning (RL) reshapes flow-based vision-language-action (VLA) models, finding that across π0.5 and GR00T N1.5/N1.6 on LIBERO, ManiSkill, MetaWorld, and CALVIN, RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules, which hold 14.83%–27.58% of the parameters yet capture a disproportionate share of the RL gain, with shift-vector update directions predicting task success at up to 99.6% ROC-AUC and steering along them improving policies without additional RL training.
  4. arXiv

    GGSD turns five buttons into playable skills via 1v1 self-play: humans clear Maze and CubePush on Ant, Franka and G1 with no extra training

    The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
  5. arXiv

    RoboCoach turns imagined failures into targeted demos: 150 extra subtask demonstrations lift Franka success from 13.3% to 75.0%

    The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
  6. arXiv

    EvolvingNav predicts where targets go with a time-indexed 4D belief, lifting first-inspection success from 45.33% to 61.32% on EvoWorld-Bench

    The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
  7. Journal of Computer Science and Technology

    Survey maps high-level synthesis for approximate computing around error estimation, approximation techniques, and design space exploration, and flags research gaps

    Addressing the lack of a systematic survey and in-depth analysis of the latest methodologies in high-level synthesis for approximate computing (AHLS), this survey summarizes recent technologies in the field with particular focus on error estimation, approximation techniques, and design space exploration (DSE), and analyzes current research gaps, aiming to give researchers, engineers, and scholars a theoretical and practical framework for AHLS.
  8. Natural Sciences and Applied Technology

    DCONVNET splits 2D direction-of-arrival estimation into two 1D problems and estimates azimuth and elevation via the alternating direction method of multipliers

    The work presents a fast two-dimensional direction-of-arrival (DOA) estimation approach for low-elevation targets of very-high-frequency array radar: it uses the azimuth and pitch angle uncoupling properties of a uniform planar array to turn the 2D angle estimation problem into two 1D DOA estimation problems, retrieves target information in the azimuth and elevation dimensions with digital beamforming, and then estimates azimuth and pitch angles using the alternating direction method of multipliers, thereby reducing complexity and eliminating the need for eigenvalue decomposition during operation.
  9. Natural Sciences and Applied Technology

    Ridezy derives driver credibility in real time from edge AI and IoT sensors and anchors hashes and reputation updates on Polygon, outperforming rating-based, AI-only, and blockchain-only baselines in behavioural fidelity and trust guarantees

    The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
  10. Journal of Mining Institute

    Russian team used fuzzy clustering to sort seven high-temperature slags into three groups, finding steelmaking slag resource-valuable but moderately hazardous while copper and incinerator slags are high-hazard

    The study measured the chemical composition and physical properties of seven types of high-temperature process waste from the Ural industrial region (granulated blast-furnace slag, lump blast-furnace slag, steelmaking slag, electric steelmaking slag, copper-smelting slag, ferrochrome production slag, and waste-incinerator slag), selected resource indicators (mass fractions of metallic iron, Cu, Zn, Ni; basicity modulus; crystallinity; particle-size distribution) and environmental indicators (mass fractions of Pb, Cr, S, P; dust fraction proportion; leachability), and after normalization and multicollinearity checks applied fuzzy clustering in Statistica (c=3, m=2, ε=0.

Page 2 · showing 10