Are Context and Physical Feedback Complementary?
On novel objects in Flip Box, the combined policy succeeds where either input alone is insufficient.
Loading the task highlights…
A closer look at interaction, with and without FP2.
Keep task-level action generation in the foundation policy.
Add lightweight, high-frequency force regulation downstream.
Robotic foundation models (RFMs) are increasingly capable of general-purpose manipulation, yet reliable physical interaction remains challenging in contact-rich settings. We present FP2, a lightweight downstream interface that equips task-adapted RFMs with explicit force control while preserving their action-generation capability. FP2 adopts an action-regulation decomposition: the task-adapted RFM serves as a foundation policy responsible for task-level action generation, while a high-frequency force control policy focuses solely on interaction regulation. To condition force regulation on the ongoing manipulation, FP2 compresses foundation-policy contextual representations and combines them with wrench and proprioceptive histories to predict structured force-control parameters. We evaluate FP2 with four RFM backbones across four real-world contact-rich manipulation tasks. FP2 consistently improves task performance and force regulation quality over the corresponding foundation policies, while comparing favorably with representative force-aware and force-control baselines. Ablations further show that foundation-policy context and physical feedback are complementary for effective force regulation, while preserving foundation-policy action generation improves both efficiency and novel-object generalization.
A robot can generate a plausible task-level motion and still lose contact, apply insufficient force, or jam during execution.
Robotic foundation models (RFMs) bring broad task understanding and strong motion priors. But contact-rich manipulation also requires regulating how the robot interacts with the environment.
FP2 separates these responsibilities. The task-adapted foundation policy keeps generating actions. A lightweight downstream policy uses physical feedback to regulate their execution.
Force control does not need to be learned inside the foundation model. FP2 preserves task-level action generation and introduces force only into downstream interaction regulation.
Preserve action generation. Compress context. Regulate physical interaction.
Task-adapted. Then frozen.
↓ Wrench + proprioceptive history
Lightweight. Physically reactive.
From foundation policyOriginal action chunks 15 Hz waypoints → 50 Hz references
Structured force-control parameters No action re-generation
The task-adapted foundation policy predicts an action chunk and contextual tokens. Force is not an input to this model.
Illustrative data flow, slowed down for explanation—not measured timing or force data. In deployment, both policies run concurrently; the force-control policy reuses cached context between foundation-policy updates.
Adapt the foundation model. Use its native observations, without force input.
Compress its context. Learn a compact token through self-supervised reconstruction.
Train force regulation. Keep the foundation policy and compressor frozen.
The full framework is in Figure 2 of the paper (coming soon).
We evaluate four foundation policies and five force-aware or force-control baselines on four real-world tasks, using 25 test configurations per task. Select a task and method below to inspect the results and evaluation videos.
Average score is the arithmetic mean of the four task metrics (three success rates and Wipe Curve completion rate), rounded as in Table I. Normalized Force Error (NFE) measures departure from demonstrated force ranges; lower is better. NFE uses valid successful executions only, and successful continuous-wiping executions for Wipe Curve. “—” means no valid successful execution is available.
FP2 improves physical interaction while retaining foundation-policy action generation. With π₀ LoRA, the average task score increases from 15% to 80%, a gain of 65 percentage points (Table I). The average combines three success rates and Wipe Curve completion rate.
Loading verified paper results…
Each foundation policy is compared with its own FP2 version.
Each grid contains multiple trials, not one rollout. Playback is independently edited and accelerated; grids are not synchronized trial-by-trial.
FP2 primarily reduces interaction-related failures. The remaining failures are more often associated with position mismatch and geometric alignment. Force regulation improves execution of viable task-level motion; it does not replace action generation or geometric reasoning.
Table II · Flip Box with π₀.₅ · 40 novel-object trials across color, texture, stiffness, and geometry. Percentages are rounded as in the paper; latency is for the force-control policy on an RTX 5090, not the entire system.
On novel objects in Flip Box, the combined policy succeeds where either input alone is insufficient.
Re-generating actions downstream reduces novel-object performance and adds inference latency.
Keeping action generation in the foundation policy and removing the extra wrist-camera stream from the force-control policy improves novel-object generalization.
FP2 assumes the foundation policy provides viable task-level motion. It cannot fully recover from incorrect action generation or geometric reasoning. Force regulation is currently learned at the task level; shared, multi-task interaction skills remain future work.
Manuscript citation.
@article{fang2026fp2,
title = {FP2: Equipping Robotic Foundation
Models with Force Control},
author = {Fang, Hongjie and Tang, Shirun and
Hu, Junjian and Zhang, Shidong and
Zhang, Zhongze and Chen, Linhao and
Li, Dehai and Mei, Mingyu and
Liu, Wanxi and Lu, Cewu and
Wang, Shiquan},
journal = {arXiv preprint arXiv:},
year = {2026}
}See the PDF for the full paper.