Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

Chronological Source Flow
Back

AI Fusion Summary

Graphical User Interface agents utilize on-policy self-distillation (OPSD) to fulfill complex instructions via multi-turn interactions. While OPSD provides dense token-level supervision through privilege-conditioned self-teachers, extending it to multi-turn agents is difficult due to limited privilege-following abilities. Research indicates that this paradigm can lead agents to act with confidence despite lacking necessary information, sometimes underperforming compared to RL or base models. Consequently, Privileged Self-Practice (PSP) is proposed to better integrate privileged information for improved agent performance.
Community Comments
Loading updates...
0