via Transformer-Based Fusion of Visuotactile Data, Image Data, and Vibration Signals
Intelligent Autonomous Systems
11.09.2026
ToDo
Can slip velocity be estimated from visuotactile and vibration data?

Chen, Prepscius, Lee & Lee, IEEE RA-L 6(2) (2021)

Waltersson & Karayiannidis, arXiv:2606.11952 (2026)








Hanaor, Gan & Einav, Tribology International 93 (2016)
Giovani Luís Rech, youtu.be/1fHGTir78Hk

Visuotactile

raw image

marker flow

depth map
Vibration

Slip velocity

Appendix
Hybrid at 2.13 mm/s slip-MAE

parametric term in the output head to model slip spike
Event and twin Event worse than Hybrid
fast causal vibration core, conditioned by slow visuotactile context
Rate-core is the only one that meets the 1.5ms budget from Chen et al.
Hybrid trained on time-shifted labels,


Böhm, A., Schneider, T., Belousov, B., Kshirsagar, A., Lin, L., Doerschner, K., … & Peters, J., What matters for active texture recognition with vision-based tactile sensors, in 2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 15099–15105, IEEE (2024).
youtu.be/KWCL-jCnivI, arx-x.com
More variance


Combinations of architecture

Slip velocity control
Thank you for your attention, self-, cross-, and otherwise.I hope you’ll let that pun slide and stick around for the questions.
Apendix
Cross-modal self-supervised pre-training of encoders

Strip no image stream at all, physically motivated scalars instead
depth added to Fusion (L), Transformer, Hybrid

Hybrid S beats Fusion L by 10 % with 13.5× fewer parametersHybrid and Transformer lose accuracy at the top rung: overfitting
Evidence





