FP2: Equipping Robotic Foundation Models
with Force Control

  • 1FORTE Lab
  • 2Noematrix
  • 3Flexiv
  • 4SJTU
  • 5UPenn
  • 6FDU
  • 7UIUC
  • 8ZJU
  • 9SII

* Equal contribution. † Corresponding author.

Does force control need to be learned inside the foundation model?

See the difference

Loading the task highlights…

A closer look at interaction, with and without FP2.

Keep task-level action generation in the foundation policy.
Add lightweight, high-frequency force regulation downstream.

Abstract

Robotic foundation models (RFMs) are increasingly capable of general-purpose manipulation, yet reliable physical interaction remains challenging in contact-rich settings. We present FP2, a lightweight downstream interface that equips task-adapted RFMs with explicit force control while preserving their action-generation capability. FP2 adopts an action-regulation decomposition: the task-adapted RFM serves as a foundation policy responsible for task-level action generation, while a high-frequency force control policy focuses solely on interaction regulation. To condition force regulation on the ongoing manipulation, FP2 compresses foundation-policy contextual representations and combines them with wrench and proprioceptive histories to predict structured force-control parameters. We evaluate FP2 with four RFM backbones across four real-world contact-rich manipulation tasks. FP2 consistently improves task performance and force regulation quality over the corresponding foundation policies, while comparing favorably with representative force-aware and force-control baselines. Ablations further show that foundation-policy context and physical feedback are complementary for effective force regulation, while preserving foundation-policy action generation improves both efficiency and novel-object generalization.

Separating Action Generation from Interaction Regulation

A robot can generate a plausible task-level motion and still lose contact, apply insufficient force, or jam during execution.

Robotic foundation models (RFMs) bring broad task understanding and strong motion priors. But contact-rich manipulation also requires regulating how the robot interacts with the environment.

FP2 separates these responsibilities. The task-adapted foundation policy keeps generating actions. A lightweight downstream policy uses physical feedback to regulate their execution.

Force control does not need to be learned inside the foundation model. FP2 preserves task-level action generation and introduces force only into downstream interaction regulation.

  1. Preserve the motion prior. Keep task-level action generation in the foundation policy. Freeze it after task adaptation.
  2. React to physical feedback. Combine the model’s task context with recent force, torque, and proprioceptive histories.
  3. Regulate, rather than re-predict. Predict the interaction frame, force–position selection, and desired wrench—not another set of actions.

Equipping Foundation Policies with Force Control

Preserve action generation. Compress context. Regulate physical interaction.

Vision + languageWrench + proprioceptive history
Action generation2 Hz

Foundation policy

Task-adapted. Then frozen.

π₀π₀.₅GR00T N1.7LaWAM
ContextCompress

↓ Wrench + proprioceptive history

Interaction regulation50 Hz

Force-control policy

Lightweight. Physically reactive.

Frame ΣSelection SWrench Ŵ

From foundation policyOriginal action chunks 15 Hz waypoints → 50 Hz references

Structured force-control parameters No action re-generation

Hybrid force–position controller1 kHz
01 / 04

Keep task-level action generation.

The task-adapted foundation policy predicts an action chunk and contextual tokens. Force is not an input to this model.

Illustrative data flow, slowed down for explanation—not measured timing or force data. In deployment, both policies run concurrently; the force-control policy reuses cached context between foundation-policy updates.

01

Adapt the foundation model. Use its native observations, without force input.

02

Compress its context. Learn a compact token through self-supervised reconstruction.

03

Train force regulation. Keep the foundation policy and compressor frozen.

See the full architecture from the paper

The full framework is in Figure 2 of the paper (coming soon).

Experiments

Contact-Rich Manipulation across Foundation Policies

We evaluate four foundation policies and five force-aware or force-control baselines on four real-world tasks, using 25 test configurations per task. Select a task and method below to inspect the results and evaluation videos.

All Backbones and Baselines

Average score is the arithmetic mean of the four task metrics (three success rates and Wipe Curve completion rate), rounded as in Table I. Normalized Force Error (NFE) measures departure from demonstrated force ranges; lower is better. NFE uses valid successful executions only, and successful continuous-wiping executions for Wipe Curve. “—” means no valid successful execution is available.

FP2 improves physical interaction while retaining foundation-policy action generation. With π₀ LoRA, the average task score increases from 15% to 80%, a gain of 65 percentage points (Table I). The average combines three success rates and Wipe Curve completion rate.

Loading verified paper results…

Each foundation policy is compared with its own FP2 version.

Evaluation Videos

25 trials per method · edited playback
Loading evaluation videos…

Each grid contains multiple trials, not one rollout. Playback is independently edited and accelerated; grids are not synchronized trial-by-trial.

Failure Analysis

FP2 primarily reduces interaction-related failures. The remaining failures are more often associated with position mismatch and geometric alignment. Force regulation improves execution of viable task-level motion; it does not replace action generation or geometric reasoning.

Ablations and Analysis

Table II · Flip Box with π₀.₅ · 40 novel-object trials across color, texture, stiffness, and geometry. Percentages are rounded as in the paper; latency is for the force-control policy on an RTX 5090, not the entire system.

Are Context and Physical Feedback Complementary?

On novel objects in Flip Box, the combined policy succeeds where either input alone is insufficient.

Why Preserve Foundation-Policy Action Generation?

Re-generating actions downstream reduces novel-object performance and adds inference latency.

Generalization to Novel Objects

Keeping action generation in the foundation policy and removing the extra wrist-camera stream from the force-control policy improves novel-object generalization.

Loading the generalization comparison…

Limitations and Future Work

FP2 assumes the foundation policy provides viable task-level motion. It cannot fully recover from incorrect action generation or geometric reasoning. Force regulation is currently learned at the task level; shared, multi-task interaction skills remain future work.

BibTeX

Manuscript citation.

@article{fang2026fp2,
  title   = {FP2: Equipping Robotic Foundation
             Models with Force Control},
  author  = {Fang, Hongjie and Tang, Shirun and
             Hu, Junjian and Zhang, Shidong and
             Zhang, Zhongze and Chen, Linhao and
             Li, Dehai and Mei, Mingyu and
             Liu, Wanxi and Lu, Cewu and
             Wang, Shiquan},
  journal = {arXiv preprint arXiv:},
  year    = {2026}
}

See the PDF for the full paper.