LOPAL teaches robots from imperfect human demonstrations by combining the best parts of each demonstration and asking the user for corrections only in regions where high-quality data is missing.
Learning from Demonstration (LfD) enables intuitive robot skill acquisition by allowing robots to learn directly from human task demonstrations. However, current methods often fail to address the fact that due to suboptimal and inconsistent human behavior, the quality of the demonstration can vary within each demonstration. Therefore, we introduce LOPAL (LOcal Performance-aware Active Learning), an active learning approach that leverages this local demonstration quality information. Our approach consists of two synergistic components. First, a local performance-driven LfD method uses a Gaussian Mixture Model (GMM) to encode both the demonstrated trajectories and their associated local quality assessments. This enables the generation of trajectories that outperform the imperfect demonstrations by utilizing complementary local data of high performance. Second, active data acquisition allows to improve beyond the imperfect demonstrations by collecting additional informative samples. In areas missing good data, the user is actively requested to provide corrections through a shared autonomy (SA) mechanism, while the robot autonomously executes the learned behavior. The efficacy of LOPAL was validated in both a simulation and a real-world experiment. The results from a real-world pipe inspection task showed that the proposed approach can achieve up to 27.31% improvement in task performance while also reducing the effort required to collect the demonstrations.
The starting point is simple. Whenever a person demonstrates a task, the local performance can be assessed at each point along the demonstrated motion. This local quality signal can be binary — for example, did the car stay on the track here, or was the probe in contact with the pipe? — or continuous, reflecting properties such as motion smoothness, energy efficiency, distance to a target, or any other task-relevant criterion. LOPAL records this local quality information alongside the demonstrated motion and uses it in two coordinated ways.
The robot encodes every demonstrated motion together with its local performance information in a single statistical model. When generating its own behavior, the model conditions the motion on a high local-performance target, so that, for each part of the task, it follows the locally best available demonstrated behavior. Two demonstrations with imperfections in different regions can therefore be combined into a single motion that outperforms either one.
In regions where every demonstration was of low quality, the robot interpolates between adjacent high-quality data points to estimate a better nominal behavior — producing a plausible motion even where good demonstration data is missing.
The robot then executes the task autonomously. In regions where high-quality demonstration data is missing, it visibly slows down and lights up a red LED, indicating to the user that better data is needed in this region.
The user can take over control at any time by applying wrench to the robot. The robot then dynamically adjusts its control compliance — becoming compliant under the user's touch and stiff again when released. Each correction is recorded as a new demonstration and added to the dataset, so the next iteration is better — and the user only has to demonstrate the parts where the robot is still missing high-quality data.
End-to-end view of the framework. Blue: demonstrations and their local performance information are encoded into a motion model that selectively reproduces the locally best demonstrated behavior. Orange: the robot executes the generated reference motion, slowing down in regions of low local performance; the user provides force-based corrections, which are recorded and incrementally added to the demonstration dataset.
We evaluated LOPAL on two tasks: a car-racing simulation and a real-world pipe-inspection task on a 7-DoF robot arm.
Drive a car around a 2D track from start to finish without leaving the road. A penalty was incurred whenever the vehicle's centre left the road. We collected 15 keyboard-driven human demonstrations — all of which contained local imperfections.
Trace a copper pipe with a probe held by the robot, maintaining continuous contact between probe and pipe — representative of contact-based inspection tasks, where sensors typically require direct contact for proper functioning. Loss of contact was counted as a penalty. Eight participants taught the task incrementally on a Franka Research 3 robot.
When all demonstrations contained imperfections, the baseline that does not leverage performance information reproduced those imperfections through its averaging behavior. A quality-aware alternative that maximises distance from failed demonstrations performed even worse. LOPAL, by contrast, combined the locally best parts of each demonstration and generated a motion that outperformed any single demonstration.
When only one or two demonstrations were available, interpolating regions lacking high-quality demonstration data provided a further benefit over the other performance-aware variants. In practice, this means fewer demonstrations are needed to reach a given level of performance.
On the pipe-inspection task, LOPAL achieved an average task-performance improvement of up to 27.31% over the baseline that does not leverage local performance information, consistently across the user study.
Participants injected significantly less energy into the robot when teaching with the performance-aware variants than with gravity-compensation teaching or with a shared-autonomy variant that ignored local performance. Participants also rated LOPAL lower on mental demand than the other shared-autonomy variants — consistent with the active feedback (slow-down and visual cue) helping them focus their corrections where they were most needed.
@article{heidersberger2026lopal,
title = {{LOPAL}: Local Performance-Aware Active Learning from Imperfect Demonstrations},
author = {Heidersberger, Johannes and Jadav, Shail and Lee, Dongheui},
journal = {IEEE Robotics and Automation Letters},
year = {2026},
volume = {11},
number = {7},
pages = {9040--9047},
doi = {10.1109/LRA.2026.3698364}
}