A point-cloud grasp policy cleared log piles on a hydraulic forestry crane, with BC and BCRL recovering 93.8% and 88.9% of inventory versus 80.4% for a geometric heuristic
Synopsis
The work trains one network that classifies which observed point to grasp and predicts depth and yaw there from unsegmented point clouds, learns to clear 200-log packed piles in simulation via behavior cloning (BC), its reinforcement-learning fine-tuning (BCRL), and RL from scratch, and runs twelve field trials on a trailer-mounted hydraulic forestry crane: BC and BCRL deposit 93.8% and 88.9% of pooled inventory against 80.4% for a geometric heuristic, and BCRL succeeds on 83.6% of cycles versus 79.6% for the heuristic and 65.7% for BC, while fine-tuning's simulated stability gain does not carry over to the crane testbed.
Fig. 1: Trailer-mounted hydraulic forestry crane, suspended grapple, source rack, and destination trailer. The policy selects the grasp; a common controller transports and deposits the logs.
arXivInterpretation
A policy maps unsegmented point clouds directly to grasp decisions: the network outputs a selection score, vertical offset, and doubled-angle yaw encoding at every observed point, and the highest-scoring unmasked point sets grasp position and orientation, with one network serving behavior cloning, RL fine-tuning, and deployment. Unlike approaches that rely on geometric filters or a separate segmentation network, this policy learns material selection and grasp prediction together; unlike a regression head that directly predicts grasp coordinates, classifying over points preserves separate graspable regions and avoids squared-error averaging of alternative targets into empty space. Across three training seeds the scoring head empties 98, 98, and 99 of 100 simulated racks, while the regression head empties 33, 0, and 0; among 1784 recorded regression decisions, 14.6–24.7% have no observed point within 0.5 m horizontally, whereas scoring BC shows neither diagnostic failure over 2140 decisions.
On the physical crane testbed, BC and BCRL, trained entirely in simulation, run with unchanged weights, and fed unfiltered clouds that still contain the storage rack's rails and poles, deposit more inventory than the deployed geometric heuristic. Prior forestry grasping work plans motion for a given target or plans a single grasp on a static scene; it does not address sequential clearing of a pile whose state evolves under the policy's own actions, nor report pile-scale hardware trials. This work provides twelve complete grasp-transport-deposit field trials. Pooled over three trials, BC deposits 93.8% and BCRL 88.9% of initial inventory against 80.4% for the heuristic; the learned policies lead on two of three pile shapes and overall, most clearly on the flat rack (both empty it while the heuristic stops at 62.0% Clear), while the double mound reverses the ordering (the heuristic deposits everything).
Fine-tuning's load-stability gain does not transfer to the crane testbed, while its clearing advantage does; field failures are classified by cause and trial phase. The result localizes the sim-to-real gap in handling rather than in clearing completion, and motivates evaluating task completion and handling separately. The fine-tuned policy's stability is below BC and the heuristic on every shape, by 0.02 to 0.06, about the size of its simulated gain; empty-cycle causes are grouped into structure or noise targeting, chokes, depth faults, and control faults, and all six heuristic structure/noise cycles occur in the final third of its trials.
RL from scratch does not improve completion over BC under the tested reward and budget, and shows a systematic depth bias that requires manual depth offsets to grasp on the crane. Exploration does depart from the demonstrated top-of-pile strategy—consecutive targets in the first five cycles are a median 0.4 m apart versus 1.5 m for the expert and fine-tuned policy—but this regional strategy does not yield better clearing completion. RL from scratch has the lowest stability on all three shapes and the lowest pooled alignment and logs per successful cycle; with applied offsets removed, its mound targets lie a median 0.17 m above the observed surface, whereas BC and BCRL learn depth from demonstrations relabeled using crane commissioning.
Perspective
The policy targets clearing packed log piles inside a rack with a trailer-mounted hydraulic forestry crane, with grasp decisions restricted to a predefined action box and horizontal targets confined to measured surfaces; training is entirely in simulation and deployment uses unchanged weights on unfiltered clouds that include rack rails, poles, and noise. For readers who want to reuse it, this offers a path without a segmentation network or geometric support filter, and shows that a point-classification grasp representation can serve behavior cloning, RL fine-tuning, and deployment with one network. The authors suggest next steps: correcting remaining depth faults with small amounts of real-trial data, extending the rack-pole penalty to the base rails, and randomizing contact properties during training.
Each policy and shape has one rebuilt pile, and rebuilt piles share shape categories rather than identical paired states, so cross-platform performance and the cause of the stability result remain open; the authors also note that simulation misses physical chokes and the field stability ordering, with possible sources including hydraulic and suspension dynamics, unmeasured friction, and the bark and knots of real logs. In addition, RL from scratch results include manual depth offsets, so its performance is not directly equivalent to fully autonomous depth prediction; the scoring-versus-regression comparison and the categorical-versus-Gaussian exploration comparison rest on limited seeds, and learning from scratch varies widely by seed.
