LOPAL: Local Performance-Aware Active Learning
from Imperfect Demonstrations

IEEE Robotics and Automation Letters (RA-L), 2026

LOPAL teaches robots from imperfect human demonstrations by combining the best parts of each demonstration and asking the user for corrections only in regions where high-quality data is missing.

Conceptual illustration of the LOPAL framework.

Each demonstration (black dashed) contains local imperfections — here, the parts where the car left the track (red). LOPAL generates a motion that follows the locally best available demonstrated behavior (blue) instead of averaging the imperfections. In regions where every demonstration is of low quality, it interpolates between adjacent high-quality data points (orange) to estimate a better nominal behavior. During execution, the user is actively requested to correct this nominal behavior (green) only in those regions.

Abstract

Learning from Demonstration (LfD) enables intuitive robot skill acquisition by allowing robots to learn directly from human task demonstrations. However, current methods often fail to address the fact that due to suboptimal and inconsistent human behavior, the quality of the demonstration can vary within each demonstration. Therefore, we introduce LOPAL (LOcal Performance-aware Active Learning), an active learning approach that leverages this local demonstration quality information. Our approach consists of two synergistic components. First, a local performance-driven LfD method uses a Gaussian Mixture Model (GMM) to encode both the demonstrated trajectories and their associated local quality assessments. This enables the generation of trajectories that outperform the imperfect demonstrations by utilizing complementary local data of high performance. Second, active data acquisition allows to improve beyond the imperfect demonstrations by collecting additional informative samples. In areas missing good data, the user is actively requested to provide corrections through a shared autonomy (SA) mechanism, while the robot autonomously executes the learned behavior. The efficacy of LOPAL was validated in both a simulation and a real-world experiment. The results from a real-world pipe inspection task showed that the proposed approach can achieve up to 27.31% improvement in task performance while also reducing the effort required to collect the demonstrations.

Video

Teaser video

Full video

How LOPAL Works

The starting point is simple. Whenever a person demonstrates a task, the local performance can be assessed at each point along the demonstrated motion. This local quality signal can be binary — for example, did the car stay on the track here, or was the probe in contact with the pipe? — or continuous, reflecting properties such as motion smoothness, energy efficiency, distance to a target, or any other task-relevant criterion. LOPAL records this local quality information alongside the demonstrated motion and uses it in two coordinated ways.

1

Combine the locally best parts of each demonstration

The robot encodes every demonstrated motion together with its local performance information in a single statistical model. When generating its own behavior, the model conditions the motion on a high local-performance target, so that, for each part of the task, it follows the locally best available demonstrated behavior. Two demonstrations with imperfections in different regions can therefore be combined into a single motion that outperforms either one.

In regions where every demonstration was of low quality, the robot interpolates between adjacent high-quality data points to estimate a better nominal behavior — producing a plausible motion even where good demonstration data is missing.

2

Actively request corrections only where they are needed

The robot then executes the task autonomously. In regions where high-quality demonstration data is missing, it visibly slows down and lights up a red LED, indicating to the user that better data is needed in this region.

The user can take over control at any time by applying wrench to the robot. The robot then dynamically adjusts its control compliance — becoming compliant under the user's touch and stiff again when released. Each correction is recorded as a new demonstration and added to the dataset, so the next iteration is better — and the user only has to demonstrate the parts where the robot is still missing high-quality data.

LOPAL framework overview.

End-to-end view of the framework. Blue: demonstrations and their local performance information are encoded into a motion model that selectively reproduces the locally best demonstrated behavior. Orange: the robot executes the generated reference motion, slowing down in regions of low local performance; the user provides force-based corrections, which are recorded and incrementally added to the demonstration dataset.

Experiments & Results

We evaluated LOPAL on two tasks: a car-racing simulation and a real-world pipe-inspection task on a 7-DoF robot arm.

Task A · Simulation

Car racing on a 2D track

Car racing simulation track with example demonstrations.

Drive a car around a 2D track from start to finish without leaving the road. A penalty was incurred whenever the vehicle's centre left the road. We collected 15 keyboard-driven human demonstrations — all of which contained local imperfections.

Task B · Real-world

Pipe inspection with a robot arm

Real-world pipe inspection experiment setup.

Trace a copper pipe with a probe held by the robot, maintaining continuous contact between probe and pipe — representative of contact-based inspection tasks, where sensors typically require direct contact for proper functioning. Loss of contact was counted as a penalty. Eight participants taught the task incrementally on a Franka Research 3 robot.

What we found

Better motions from imperfect demonstrations

When all demonstrations contained imperfections, the baseline that does not leverage performance information reproduced those imperfections through its averaging behavior. A quality-aware alternative that maximises distance from failed demonstrations performed even worse. LOPAL, by contrast, combined the locally best parts of each demonstration and generated a motion that outperformed any single demonstration.

More efficient learning from fewer demonstrations

When only one or two demonstrations were available, interpolating regions lacking high-quality demonstration data provided a further benefit over the other performance-aware variants. In practice, this means fewer demonstrations are needed to reach a given level of performance.

Up to 27% higher task performance on the real robot

On the pipe-inspection task, LOPAL achieved an average task-performance improvement of up to 27.31% over the baseline that does not leverage local performance information, consistently across the user study.

Reduced teaching effort for the user

Participants injected significantly less energy into the robot when teaching with the performance-aware variants than with gravity-compensation teaching or with a shared-autonomy variant that ignored local performance. Participants also rated LOPAL lower on mental demand than the other shared-autonomy variants — consistent with the active feedback (slow-down and visual cue) helping them focus their corrections where they were most needed.

BibTeX

@article{heidersberger2026lopal,
  title   = {{LOPAL}: Local Performance-Aware Active Learning from Imperfect Demonstrations},
  author  = {Heidersberger, Johannes and Jadav, Shail and Lee, Dongheui},
  journal = {IEEE Robotics and Automation Letters},
  year    = {2026},
  volume  = {11},
  number  = {7},
  pages   = {9040--9047},
  doi     = {10.1109/LRA.2026.3698364}
}