A Survey · TF-ART

Learning Physical Interaction

A Survey of Tactile- and Force-aware Robot Learning

Shilin Shan1,*,‡, Chuhao Zhou1,*, Ruize Wang1,*, Xinyan Chen1,*, Xiangyu Chen1,*, Xinyu Zhou1,*, Boyu Ma1,*, Iris Yuxuan Hu1, Jingliang Li1, Celeste Yuxuan Hu1, Geng Li1, Guohao Chen1, Tianrui Zhu1, Zhe Li1, Yanjie Ze2, Haoran Geng3, Zhiyang Dou4, Jianxin Bi5, Yuejiang Liu2, Jianshu Zhou5, Jiachen Li6, Paul Liang4, Tatsuya Harada7, Robert Katzschmann8, Harold Soh5, Na Li9, Edward Johns10, Danica Kragic11, Jan Peters12, Wojciech Matusik4, Masayoshi Tomizuka3, Jitendra Malik3, Jianfei Yang1,†

1Nanyang Technological University 2Stanford University 3UC Berkeley 4MIT 5National University of Singapore 6Georgia Tech 7The University of Tokyo 8ETH Zurich 9Harvard University 10Imperial College London 11KTH Royal Institute of Technology 12TU Darmstadt

*Equal contribution Project lead Corresponding author

§ 01 — Abstract

Beyond seeing and moving: feeling and regulating.

Learning-based

Learned policies generalize from multimodal demonstrations, but offer no guarantee of stable contact.

Control-oriented

Model-based motion generation and compliance control stabilize contact, but cannot decide what to do.

TF-ART: one taxonomy mapping how modern systems combine both.

Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on force regulation and adaptive control. In this context, recent robot learning methods have made substantial progress by integrating force, tactile, vision, language, and proprioceptive sensing into learned manipulation policies. In parallel, many systems adopt multi-phase architectures that combine high-level policies, action-refinement modules, and low-level controllers to bridge semantic task understanding with reactive physical execution.

Despite these advances, existing surveys have not explicitly reviewed force- and tactile-aware robot learning from a unified perspective that jointly captures multimodal sensing and multi-phase system design. This survey addresses this gap by proposing TF-ART, a Tactile/Force-Aware Robot learning Taxonomy for multimodal and multi-phase frameworks, which maps individual methods into a unified hierarchical structure. The framework characterizes how recent works organize observation modalities, encode and fuse heterogeneous sensory inputs, generate and refine actions across multiple phases, and connect learned policies to reactive robot-end control. Building on this methodological view, we further examine the task settings and infrastructure requirements of physical interaction, thereby integrating both algorithmic and practical perspectives on force- and tactile-aware robot learning.

Overview of the unified multi-modal multi-phase framework: multimodal perception, encoding and fusion, primary action generation, action refinement, and low-level robot-end control.
FIG. 01 · Framework Overview

The unified architecture: perception, fusion, action generation, refinement, and robot-end control, with auxiliary reconstruction and prediction branches.

§ 02 — Framework Explorer

One architecture. Every method, mapped.

The full pipeline, from raw sensing to reactive control — with all 53 corpus papers placed on it. Hover any module to surface its papers  ·  Select a paper to trace its full pipeline  ·  Esc to clear

§ 03 — Inside the Survey

Each phase, dissected.

Multimodal fusion and encoding framework: physical interpretation of modalities and model structures of representation-learning and fusion methods.
FIG. 03 · Phase 0 — Fusion & Encoding

From heterogeneous senses to one representation.

Vision, language, force/torque, tactile and state observations — encoders trained by reconstruction, masked autoencoding, vector quantization or contrastive alignment, then fused through FiLM, attention, mixture-of-experts, or gated filtering.

Phase 1 primary action policies: VLA, diffusion and flow matching, action chunking transformers, reinforcement learning, and regression.
FIG. 04 · Phase 1 — Primary Action Policy

From fused observations to actions.

Five policy families — VLA, diffusion / flow matching, ACT, RL, regression — and what they hand downstream: action chunks, control stiffness, reference forces, plus predictions that stand in for sensors the robot lacks.

Phase 2 refinement policies refining primary actions with forwarded and predicted modalities.
FIG. 05 · Phase 2 — Refinement Policy

A second, faster policy fixes the draft.

A lightweight refiner runs at higher frequency than the primary policy, correcting its action plan with the force and tactile feedback a large model is too slow to use.

Phase 3 robot-end control: impedance, admittance, hybrid force-position, and PID control closing the loop at the robot end.
FIG. 06 · Phase 3 — Robot-end Control

Where intent meets contact.

PD, impedance, admittance and hybrid force–position control close the loop at the robot end — on arms from Franka to Flexiv and UR, grippers and dexterous hands.

0 5 10 15 20 25 PAPERS · OF 53 SURVEYED A1 Insertion — 26 papers 26 A1 Insertion A2 Cap Rotation — 5 papers 5 A2 Cap Rotation A3 Assembly Plug — 11 papers 11 A3 Assembly Plug A4 Bi. Insertion — 1 paper 1 A4 Bi. Insertion A5 Rotate Handle/Box — 2 papers 2 A5 Rotate Handle/Box B1 Peeling — 8 papers 8 B1 Peeling B2 Wiping — 18 papers 18 B2 Wiping B3 Cutting — 1 paper 1 B3 Cutting B4 Slide Object — 1 paper 1 B4 Slide Object B5 Soft/Fragile Grasp — 9 papers 9 B5 Soft/Fragile Grasp C1 Mobile Catch — 1 paper 1 C1 Mobile Catch C2 Sliding — 2 papers 2 C2 Sliding C3 Reorientation — 6 papers 6 C3 Reorientation C4 In-hand Manip. — 3 papers 3 C4 In-hand Manip. C5 Bi. Lifting — 3 papers 3 C5 Bi. Lifting C6 Bi. Wiping — 1 paper 1 C6 Bi. Wiping D1 Open/Close Door — 9 papers 9 D1 Open/Close Door D2 Pick-and-place — 13 papers 13 D2 Pick-and-place D3 Pouring — 3 papers 3 D3 Pouring D4 Weight Pulling — 1 paper 1 D4 Weight Pulling D5 Push Object — 1 paper 1 D5 Push Object D Household C Dynamic & Multi-contact B Deformable & Surface-contact A Precision Contact
FIG. 07 · Task Landscape

What the field manipulates.

21 contact-rich tasks in four families. Insertion, wiping and pick-and-place account for nearly half of all task–method pairings.

§ 04 — Cite

BibTeX

@misc{shan2026learningphysicalinteractionsurvey,
      title={Learning Physical Interaction: A Survey of Tactile- and
             Force-aware Robot Learning},
      author={Shilin Shan and Chuhao Zhou and Ruize Wang and Xinyan Chen and
              Xiangyu Chen and Xinyu Zhou and Boyu Ma and Iris Yuxuan Hu and
              Jingliang Li and Celeste Yuxuan Hu and Geng Li and Guohao Chen and
              Tianrui Zhu and Zhe Li and Yanjie Ze and Haoran Geng and
              Zhiyang Dou and Jianxin Bi and Yuejiang Liu and Jianshu Zhou and
              Jiachen Li and Paul Liang and Tatsuya Harada and
              Robert Katzschmann and Harold Soh and Na Li and Edward Johns and
              Danica Kragic and Jan Peters and Wojciech Matusik and
              Masayoshi Tomizuka and Jitendra Malik and Jianfei Yang},
      year={2026},
      eprint={2608.07558},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2608.07558}
}