Detection-assisted closed-loop visual pursuit

Cross-Scene Visual UAV-to-UAV Pursuit: A Large-Scale Benchmark Built with DAgger-Based Data Aggregation

基于 DAgger 学习器状态聚合的跨场景视觉无人机追逐

Research project page

From first-person RGB observations and target bounding boxes to continuous body-frame control—across diverse UE4/AirSim environments.

闭环追逐任务、Teacher–Student 数据生成、DAgger 状态聚合和跨场景数据组成的整体流程图
Overview. Expert demonstrations provide nominal pursuit behavior, while learner-controlled rollouts and shadow-expert relabeling extend the corpus toward policy-induced long-tail states.

Abstract

快速运动的小尺度 UAV 目标、相对运动以及感知与通信延迟,使视觉空中追逐面临持续的观测—动作耦合。现有数据集多面向预录制检测或跟踪,难以支持闭环追逐学习,也较少覆盖学习器自身误差产生的长尾状态。

本工作在 UE4/AirSim 中构建检测辅助的跨场景视觉 UAV-to-UAV 追逐数据与评估流程。除在线 Expert 示范外,采用 DAgger learner rollout 与 shadow-expert relabeling 聚合 target loss、reacquisition 和 recovery 等困难状态;学习策略部署时仅接收第一视角 RGB 与目标 bbox。

5Base maps
15Scene configurations
13,521Episodes
1,401,278Recorded frames

Contributions

01

Closed-loop visual pursuit

保留 Pursuer 动作、后续视觉观测与任务结果之间的回合级闭环关系。

5 maps · 15 configurations
02

Learner-state-augmented corpus

结合在线 Expert 示范与七种策略的 DAgger rollout,覆盖策略自身诱发的困难状态。

7,955 Expert + 5,566 DAgger episodes
03

Privileged teacher · visual student

Expert 提供特权监督,部署阶段学生仅使用第一视角 RGB 与目标 bbox。

Privileged labels · restricted observations
04

Paired closed-loop evaluation

使用冻结场景清单和统一执行协议,比较相同初始条件下的闭环结果。

Frozen manifests · no expert fallback

Task definition

Region-entry-triggered visual pursuit

Target 从区域外出发;进入保护区域后触发 CHASE。Pursuer 必须在 Target 到达目标点、发生碰撞或超时之前完成捕获。

Target 进入保护区域后触发 Pursuer 追逐的任务示意
Region entry activates pursuit; capture requires a 3D pursuer–target distance of at most 1.5 m.
  1. 1
    Target starts outside

    沿场景约束的路径朝区域另一侧目标点运动。

  2. 2
    Region entry

    首次进入保护区域时激活 Pursuer 与计时。

  3. 3
    Closed-loop pursuit

    每次动作都会改变后续视角、相对几何和目标可见性。

  4. 4
    Terminal outcome

    Capture、Collision、Target Goal 或 120 s Timeout。

1.5 mcapture threshold
120 smaximum CHASE
3-axisbody-frame velocity
Capture 终止事件的 AirSim 回合画面
CapturePursuer enters the 1.5 m capture range.
Collision 终止事件的 AirSim 回合画面
CollisionPursuer collides with the environment.
Target goal 终止事件的 AirSim 回合画面
Target goalTarget reaches its goal before capture.
120seconds

No other terminal event

Privileged teacher · visual student

Separate supervision from deployment observations

Expert 可以访问仿真特权状态并生成一致动作标签;Visual Student 在部署时遵循更严格的 RGB + bbox 观测接口。

TRAINING ONLY

Privileged Expert

  • Simulator poses
  • Current target velocity
  • Online depth

Short-horizon interception (≤ 2 s) with depth-aware collision avoidance.

Body-frame expert action label a*
supervision→
DEPLOYMENT

Visual Student

  • First-person monocular RGB
  • AirSim target bounding box

When detection is unavailable: [-1, -1, -1, -1].

Predict [vx, vy, vz] in the Pursuer body frame
Visual Student 的第一视角 RGB 输入示例
Student RGB. First-person monocular observation.
包含目标红色 bbox 的 Visual Student 输入示例
Target cue. AirSim detection-interface bounding box.
仅供 Privileged Expert 使用的在线深度图
Expert depth. Privileged training-time input only.
No target 3D positionNo target velocityNo future trajectoryNo depth or global map

DAgger state aggregation

Label the states the learner actually visits

Learner 独立执行控制,shadow expert 在相同状态提供纠正动作;Expert 不接管 rollout。

01Expert demonstrationsMainly nominal states
Expert demonstration frame 020Expert demonstration frame 025Expert demonstration frame 030Expert demonstration frame 035Expert demonstration frame 040
closed-loop learner execution ↓
02Learner rolloutTarget loss · deviation · recovery
Learner rollout frame 070Learner rollout frame 096Learner rollout frame 147Learner rollout frame 195Learner rollout frame 209
same visited states ↓
03Shadow-expert relabelingCorrective labels, no takeover
相同 learner 状态上的 Expert 纠正标签 070相同 learner 状态上的 Expert 纠正标签 096相同 learner 状态上的 Expert 纠正标签 147相同 learner 状态上的 Expert 纠正标签 195相同 learner 状态上的 Expert 纠正标签 209
Seven heterogeneous rollout policies
ConvNetLSTMNetViTViT–LSTMU-Net–LSTMMamba–LSTMMamba

Dataset composition

Expert demonstrations + learner-visited states

数据按 episode 组织,保留观测、动作、任务事件、下一观测与终止结果之间的时间关系。

104.4 hrecorded interaction
74.3 hclosed-loop pursuit
12,072training episodes
1,449validation episodes
7,955 Online Expert5,566 Seven-policy DAgger
Data sourceEpisodesRecorded framesCHASE contextValid supervision
Online Expert7,955547,123306,075271,304
Seven-policy DAgger5,566854,155685,671551,315
Total13,5211,401,278991,746822,619
RecordedSTANDBY 与 CHASE 阶段保存的全部 RGB 帧。
CHASE context按时间顺序提供给模型的主动追逐图像。
Valid supervision通过学习掩码、实际参与动作回归损失的帧。

Simulation environments

Five independent base maps

每个地图均设置 Easy、Medium 与 Difficult,难度由区域位置、局部地形/障碍以及 Pursuer–Target 速度共同构成。

AirSimNH 的 UE4/AirSim 场景
AirSimNHNeighborhood and urban geometry
R = 60 m
CoastCliffs 的 UE4/AirSim 场景
CoastCliffsCliffs, islands and maritime obstacles
R = 95 m
FloatingIslands 的 UE4/AirSim 场景
FloatingIslandsVertically fragmented island terrain
R = 115 m
Forest Map01 的 UE4/AirSim 场景
Forest Map01Forested terrain with elevation variation
R = 85 m
Forest Map02 的 UE4/AirSim 场景
Forest Map02Independently configured dense forest
R = 100 m
Difficulty profilesPursuer / Target nominal speed
Easy8 / 4 m·s⁻¹
Medium10 / 4 m·s⁻¹
Difficult10 / 5 m·s⁻¹

Evaluation protocol

Paired closed-loop evaluation

比较策略使用相同冻结场景清单,评估期间无 Expert fallback;SR 是主要指标。

Q1

Same-map unseen trajectories

训练地图相同,但初始位置、目标点、相对几何与目标轨迹未见。

Q2

Nested cross-scene expansion

使用固定训练地图扩展序列,并只在仍然未见的地图上比较。

Q3

Learner-state data ablation

Expert Only、Matched 与 Full 分离状态覆盖和语料规模的作用。

Q4

Paired closed-loop behavior

在相同冻结场景上分析终止结果、捕获时序和轨迹几何。

Execution

640×480camera RGB
90×60policy input
0.1 sperception interval
0.15 scontrol period
200episodes / difficulty / map
0expert fallback

Reported metrics

SRCapture successCRCollision rateGRTarget-goal rateBABBox availabilityCTCapture timePEPath efficiencyLATInference latency

Wilson 95% confidence intervals · exact two-sided McNemar tests · multiple-comparison correction when required

Nested expansion:L1 CoastCliffs→L2 + FloatingIslands→L3 + Forest Map01→L4 + Forest Map02

AirSimNH remains unseen at every level. The sequence changes both data volume and map identity, so it does not isolate either factor causally.

Design takeaways

01

Action changes observation

闭环追逐中的每个控制动作都会改变后续视角、目标尺度、相对几何与可见性,因此不能只用固定离线图像误差描述策略。

02

Learner states matter

DAgger rollout 将 target loss、reacquisition、large deviation 和 recovery 等由策略自身引发的状态带回训练语料。

03

Supervision is not deployment input

Privileged Expert 提供一致监督;Visual Student 的部署观测仍限制为第一视角 RGB 与目标 bbox。

这里只总结任务与数据设计,不提前陈述尚未由完整消融和跨场景实验支持的性能结论。

Scope & limitations

What this work does—and does not—claim

清楚标注证据边界,让网页陈述不强于论文。

01

Simulator-provided bbox

目标框来自 AirSim detection interface,当前研究未覆盖真实检测误差和感知延迟。

02

Sim-to-real gap

数据全部来自 UE4/AirSim,尚需在实体无人机平台上验证迁移能力。

03

Single pursuer and target

当前任务为单 Pursuer 与单个目标导向 Target,尚未覆盖多智能体交互。

Resources

资源尚未正式开放;完成匿名性和公开权限核对后接入真实链接。

PaperManuscript PDFComing soon
SupplementaryAdditional materialComing soon
CodeImplementationComing soon
DatasetEpisode-level corpusComing soon
Demo videoClosed-loop pursuitComing soon
BibTeXCitation recordComing soon