Engineering Sciences
140 items
Learning-based framework enables continuous autonomous excavation on a scaled hydraulic excavator, averaging 6.52 kg payload per cycle versus 2.68 kg for Fixed Dig
The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
Does per-frame early exit pay off? A dynamic-depth speech enhancer lands on the same latency-quality frontier as static models on STM32N6, with a 26 microsecond per-frame policy
The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
Full development history of a wholly AI-authored codebase released: 14.3% of AI code-generation events in a 21,000-line Python tool contained real errors, and roughly 1 in 4-5 interactive responses contained factual errors
The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
ART turns VLA models into tool-calling robot agents, reporting a 20% higher success rate than mainstream baselines in simulation and real-world tasks
The work proposes Agentic Robot with Tool-use (ART), a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement; the authors built a dataset of 30K tool-use trajectories and action demonstrations and designed a training regimen for long-trajectory tool-use reasoning in challenging environments, with experiments showing ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks such as pick-and-place in the dark at novel viewpoints.
Modeling pedestrian crossing as a POMDP with perceptual, cognitive, and motor constraints and training it with deep RL reproduces the broadest range of empirical crossing behaviors
The work defines pedestrian-vehicle interaction as a partially observable Markov decision process (POMDP) with theory-grounded perceptual, cognitive, and motor constraints, trains it in a simulator via deep reinforcement learning with domain randomization, and shows that the learned policies produce human-like behavior in complex scenarios including multiple lanes, heavy traffic, and dangerous driving styles, reproducing the broadest range of empirical findings on human crossing behavior so far, including gap acceptance, yielding acceptance, hesitation, and evasive speed adjustment; the learned policies also transfer to unseen traffic environments and can be adapted to local traffic norms with finetuning.
A growth-inspired graph-generation framework uses dot-matrix database augmentation and a GCNN for inverse design of mechanical lattices, with a design targeting 1000 MPa validated by finite element analysis at 1027.49 MPa
This work introduces a morphogenetic graph-generation framework in which a discrete dot matrix supplies candidate nodes and the final architecture is built by sequential cross-layer and intra-layer growth; a dataset of distinct three-dimensional lattices on a 3x3x3 nodal matrix with 27 candidate nodes is evaluated by beam-based finite element analysis and represented directly as graphs, a graph convolutional neural network with three graph-convolution layers and dual global pooling learns the topology-property mapping and predicts effective compressive stiffness, and coupling this surrogate with rapid structural sampling enables inverse design: for a target stiffness of 1000 MPa the selected design was predicted at 1042.43 MPa and validated by finite element analysis at 1027.
Calibrating the torque constant with a dynamometer and aligning torque-difference observations lets a direct-drive multifingered gripper reach 100% zero-shot sim-to-real grasp success on nine in-distribution objects
The work proposes a simple torque-observation alignment method for direct-drive (DD) actuators: dynamometer calibration identifies the motor torque constant K_tau* to correct the scale mismatch between simulated and real torque, torque differences delta_tau(t) = tau(t) - tau(t-1) are used as the observation in both domains to remove the domain-dependent constant offset, and Gaussian noise derived from dynamometer measurement data is injected during learning; the authors train a teacher-student grasping policy entirely in simulation, deploy the distilled student on a multifingered DD gripper for proprioceptive grasping using only joint positions and torque differences, and the proposed method achieves 100% grasp success in an ablation study on nine in-distribution (ID) objects.
A tendon-driven robotic jellyfish uses discrete constraints to reach 150-degree bending and reinforcement learning for closed-loop depth regulation
This work presents a tendon-driven robotic jellyfish with constrained soft actuation: each actuator combines a flexible substrate with discrete constraints, enabling bending up to 150 degrees with an approximately linear tendon displacement-bending relationship; eight actuators driven by four servos perform stable swimming, attitude adjustment, and self-righting; and a reinforcement-learning controller built on that linear actuation achieves closed-loop depth regulation in both simulation and physical experiments.
HarnessPAI's evolving code harness lifts LIBERO-PRO by 61.6 points without retraining the underlying model
The work introduces HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface organizing the underlying action primitive: within a rollout it executes open-loop with a fixed program, and across rollouts it evolves closed-loop by using execution feedback to revise the program and distill failures into reusable skills; across desktop robot arms, household robots, a robot vacuum, and a legged walking agent it improves on both pure action models and code-as-policy baselines without retraining the underlying model, with a 61.6-point gain over π0.5 on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks, and the converged program also serves as an expert-data collector whose data lifts π0.
Page 7 · showing 10