Learning Joint Representation of Human Motion and Language
Contrastive learning for shared motion–language representations.
MoLang learns a shared representation of human motion and language from paired and unpaired data. The contrastive motion–language model supports action recognition and motion retrieval within a single representation space. (Kim et al., 2022)
@article{kim2022learning,title={Learning Joint Representation of Human Motion and Language},author={Kim, Jihoon and Yu, Youngjae and Shin, Seungyoun and Byun, Taehyun and Choi, Sungjoon},journal={arXiv preprint arXiv:2210.15187},month=oct,year={2022},doi={10.48550/arXiv.2210.15187},}