Skip to main content

Engineering Sciences

140 items

  1. Research Square

    Pure Hamiltonian mechanics plus active braking for swarm formation control: 9 UAVs hold 2.5 m buffers and 0.9 m spacing in undulating-terrain simulation

    The work proposes a multi-UAV formation control framework called Pure Hamiltonian 3D RK Swarm, embedding an anisotropic vertically-scaled Rimon-Koditschek navigation potential into a pseudo-Hamiltonian dynamical framework, adding an active kinematic preview and deflection layer on the virtual target's trajectory, projecting rigid spatial offsets via a dynamic SO(3) rotation matrix, and using a spatial decay braking force modulated by the normal gradient of the workspace topology to curb overshoot; in numerical simulations with N=9 agents crossing an undulating sinusoidal terrain cluster, the formation satisfies hard safety constraints of 2.5 m buffer, 2.0 m altitude buffer, and 0.9 m inter-drone spacing.
  2. Frontiers in Pharmacology

    Without touching a line of SAS source, a metadata layer turns a 558-macro clinical reporting library into LLM-readable JSON, with 11 of 14 real reports at 80%+ cell-level parity

    The study presents a non-destructive metadata-layer framework (bridge map, typed Intermediate Representation, orchestrator) that re-exposes a legacy clinical reporting library's outputs as machine-readable JSON without modifying validated SAS source, validated on a 558-component, 372,698-line industrial SAS macro library: immediate AI readiness under coexistence mode, an optional 92% reduction in proprietary code, cell-level parity of 80% or above on 11 of 14 report types from internal Phase III study PROT008-SR1 (mean 82.7%, best 99.2%), 100% parity across 5 reports and 4,764 cells on the public CDISC CDISCPilot01 benchmark, and LLM experiments covering table summarization, adverse event anomaly detection, and trial configuration generation.
  3. arXiv

    Tactile-JEPA pretrains on the taxel connectivity graph and cuts force-estimation error by 6.3% and in-hand pose error by 20.8%

    The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
  4. arXiv

    VLA-Precision reaches 98.3% mean success across nine precision chemistry tasks with 45.8 minutes of online training per task

    The work presents VLA-Precision, a framework for real-world online reinforcement learning of large vision-language-action (VLA) models, combining the Asymmetric Co-Bootstrapping (ACoB) algorithm with the ACoB-Stream training architecture, and reports 98.3% mean success across nine high-precision chemistry tasks in four categories and four robot platforms, with 45.8 minutes of online training per task, 27.6-second episodes, and up to 10.9x improvements in throughput and computational efficiency.
  5. arXiv

    Berkeley team's Morphometric Imitation retargets human hand-object interaction to three-, four-, and five-fingered robot hands, reaching 89.3% zero-shot success in 300 real-world trials

    The work presents Morphometric Imitation, a three-stage framework that first kinematically retargets human hand-object interaction to robot hands of different morphologies while preserving demonstrated contacts via morphometric optimization (MMO), then uses residual reinforcement learning with object pose and contact information to produce dynamically feasible demonstrations, and finally distills them into visuomotor policies; across three robot hands and ten GRAB trajectories, MMO improves contact F1 over the strongest of five baselines by at least 8 points, downstream dynamic retargeting success by as much as 35 points, and the distilled policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects.
  6. arXiv

    InternW0-Δ pretrains a world action model on 20K+ hours of heterogeneous data, reaching 92.8% on LIBERO-Plus and 71.9% on RoboTwin 2.0 Clean2Random

    InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
  7. arXiv

    D-JEPA supervises relations among candidate futures with executed outcomes, reaching 87.89% on PushT and a 17-point gain on physical robots

    D-JEPA introduces a decision-aligned latent world model: it first quantifies a decision-local prediction gap in which, among the few futures competing for execution, a candidate predicted closer to the goal can fail while an available alternative succeeds; it then uses executed outcomes to supervise a bounded, permutation-equivariant set operator that learns decision-relevant relations among candidate futures, and realizes the learned decision structure in JEPA-compatible future representations so aligned actions can be read out through native latent-distance planning; it reaches 87.89% success on PushT, a 15.04-point average gain on RoboTwin, a 17-point gain on physical robot tasks, and raises mean driving PDMS from 57.36 to 95.34.
  8. arXiv

    CodeGraph annotates 145 million source files into a knowledge graph with about 1 billion typed edges and grounds its concepts in Wikidata

    The work presents a pipeline that uses a code-specialised LLM (based on Qwen3-Coder-30B-A3B-Instruct) to annotate source files under an open taxonomy, extracting algorithms, paradigms, design patterns, and application domains, then grounds these in Wikidata through a three-stage procedure (deterministic SPARQL, a Deep Research Agent for the long tail, and parent-of hierarchy rollup), with a calibrated quality-assurance protocol combining a human gold set and an LLM-as-a-judge filter; applied to the Stack-Edu corpus it yields CodeGraph with roughly 158 million nodes (about 145 million file nodes, about 63,000 extracted concept entities, and roughly 19,800 grounded Wikidata entities) and about 1 billion typed edges across 14 programming languages.
  9. arXiv

    MorphIK solves inverse kinematics for unseen 6-to-9-DoF robots with morphology-conditioned flow matching, reaching about 5 cm and sub-millimeter accuracy after optimization

    MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
  10. arXiv

    Review maps how large language models enter nanophotonic design as surrogate models and agentic systems

    This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.

Page 6 · showing 10